Skip to content
Documentation
Rule catalogue

llmstxt.link.blocked-by-robots

An llms.txt must not list URLs that robots.txt forbids fetching

warningsite-wideguideline

Why it matters

The file offers URLs to agents; robots.txt forbids agents to fetch them. Anthropic documents that its bots, Claude-User among them, honour robots.txt, so each such entry is unreadable by exactly the reader it was written for. Read for User-agent: *, the file itself included, and left to robots.blocks-site when the whole origin is disallowed.

What the finding looks like

The message goflag prints, with example values substituted.

warn llmstxt.link.blocked-by-robots 2 URLs /llms.txt points agents to are disallowed by robots.txt for User-agent: *: https://example.com/raw/guide.md (line 5: Disallow: /raw/), https://example.com/raw/faq.md (line 5: Disallow: /raw/). The file offers them to agents, and an agent that honours robots.txt — Anthropic documents that Claude-User does — never fetches them.

How to fix it

The catalogue carries no snippet for a site rule: a cross-page rule builds its remedy from what the crawl found, so goflag attaches the fix to the finding itself rather than to the rule. Run it and read the finding.

Says who?

This rule is guideline, and it cites the documents below. Nothing here is goflag’s opinion of what a good page looks like — if you disagree with a finding, this is what you are disagreeing with.