sitemap. orphans
Indexable pages the crawl found, and the canonical URLs they name, should be listed in the sitemap
Why it matters
One finding with a count and a sample rather than one per page: the omission belongs to the sitemap, not to each page it forgot. A consumer that reads the sitemap instead of following links never sees them, and link-only discovery is the part of a site nobody audits. It counts only what a sitemap should list: not a page that asks for noindex — by robots or googlebot meta tag or by X-Robots-Tag, none included, to every crawler or to one — nor one that names another URL as canonical, though the URL it names is counted when nothing audited it and no sitemap lists it; not a page the sitemap reaches through a redirect; and nothing at all unless goflag read every sitemap the site declares — past a cap, an unreadable child or an unread Sitemap: line, a listed page and an unlisted one look the same. When some entries were never fetched, the count is given as a ceiling: one of them may redirect to a page counted here.
What the finding looks like
The message goflag prints, with example values substituted.
How to fix it
The catalogue carries no snippet for a site rule: a cross-page rule builds its remedy from what the crawl found, so goflag attaches the fix to the finding itself rather than to the rule. Run it and read the finding.
Says who?
This rule is guideline, and it cites the documents below. Nothing here is goflag’s opinion of what a good page looks like — if you disagree with a finding, this is what you are disagreeing with.
- Sitemaps XML protocol
sitemaps.org · normative
- Build and submit a sitemap
Google · guideline