Stored homepage words, last 12 crawls
The persistent core versus the churning shell: how consistently the homepage lands in each monthly snapshot, and at what depth. An × marks a crawl with no homepage capture.
reading 12 monthly indexes, this can take a minute…
Stored homepage words per crawl
Every AI crawler, all three gates
allowed / served
blocked
no response / partial
not applicable
Permission is the effective robots.txt policy for the homepage; "named"
means the bot appears by name, "default" means it inherits the wildcard
rules. Edge and rendering come from a live homepage fetch sending each
crawler's own User-Agent from this server, which catches UA-based
blocking and cloaking. These probes come from this server's IP, not the
crawlers' verified ranges, so an edge that checks IPs may 403 a probe
while admitting the real bot, or admit a probe while blocking the real
bot. The browser control row helps you read the difference: bot blocked
while the control passes means the edge is UA-selective; both blocked is
marked inconclusive. Only CCBot has ground truth here, and it is in the
gates above, straight from the corpus.
probing crawlers…