llmstxt. link. blocked-by-robots
An llms.txt must not list URLs that robots.txt forbids fetching
Why it matters
The file offers URLs to agents; robots.txt forbids agents to fetch them. Anthropic documents that its bots, Claude-User among them, honour robots.txt, so each such entry is unreadable by exactly the reader it was written for. Read for User-agent: *, the file itself included, and left to robots.blocks-site when the whole origin is disallowed.
What the finding looks like
The message goflag prints, with example values substituted.
How to fix it
The catalogue carries no snippet for a site rule: a cross-page rule builds its remedy from what the crawl found, so goflag attaches the fix to the finding itself rather than to the rule. Run it and read the finding.
Says who?
This rule is guideline, and it cites the documents below. Nothing here is goflag’s opinion of what a good page looks like — if you disagree with a finding, this is what you are disagreeing with.
- The /llms.txt file (a proposal)
llmstxt.org · guideline
- RFC 9309 — Robots Exclusion Protocol
IETF · normative
- Does Anthropic crawl data from the web, and how can site owners block the crawler?
Anthropic · vendor-spec