Access rules should reflect a visibility decision
Why AI crawler categories change the access decision
A request from an AI company is not automatically a request to train a model. It may be a crawl that builds a training corpus, a live fetch made while answering a user, or a search crawl that can later supply an answer surface. Treating all three alike creates an avoidable choice: a broad block may protect material from one use while also removing a useful page from the answers where prospective customers look for it.
That distinction sits at the centre of a practical generative engine optimization programme. The objective is not to open every URL to every bot. It is to set an intentional access policy by content type, platform and business outcome. Our Google AI Overviews playbook explains why ordinary search accessibility still matters when Google generates an answer from web results.
Training crawlers collect material for future models
Training crawlers visit at scale and may collect eligible public pages for model improvement or model training, subject to each provider’s policies. A publisher may reasonably limit this access for paid research, proprietary methodologies, customer portals or material where reuse is not aligned with the commercial model. The result of a training block is mostly a decision about future model knowledge, not a promise that the material will disappear from every AI product.
Training policies can also carry licensing, contractual and jurisdictional questions that robots.txt cannot settle. Robots directives are a crawl preference, not an access-control system. If a document must not be available to the public, put it behind authentication or a proper entitlement layer. A public URL that is merely disallowed remains public to people and to systems that do not honour the directive.
Live retrieval fetchers support a current answer
Live retrieval fetchers are used when an answer product searches or opens web pages near the time of a user query. Their value is freshness: an updated specification, a current policy or a new comparison can be read rather than recalled from historical training data. A citation, link or attributed passage may depend on that retrieval path remaining available.
Blocking a live fetcher can therefore reduce the chance of being cited in an answer you wanted to influence. It does not reliably stop all model use, nor does allowing it guarantee a mention. The sensible question is narrower: for this URL, is the potential referral, authority and answer visibility worth the permitted access? That is the same strategic tension examined in our guide to brand visibility in AI search.
Set robots directives by agent and purpose
Begin with a simple inventory: public editorial pages, product and service pages, gated resources, customer areas, internal search, preview URLs and assets. Then map the bots you observe or expect to the purpose their provider documents. Do not copy a large robots.txt block from another site. Names, purposes and documentation change, and the consequence of one disallow line can be wider than it appears.
Use documented user agents, not a generic AI block
Examples commonly discussed in an access policy include OpenAI’s GPTBot for training-related collection and OAI-SearchBot for search, Anthropic’s ClaudeBot and Claude-SearchBot, PerplexityBot, and Microsoft’s bingbot. Google deserves special care: Googlebot is the crawler for Google Search, while Google-Extended is a separate control concerning Gemini and Vertex AI model improvement. Google-Extended is not a switch for appearing in Google Search or its AI answer experiences.
For each documented agent, decide whether to allow the whole site, block selected directories, or block it entirely. Keep the rule specific and add a short internal note stating the owner, rationale and review date. When a provider has a distinct search or retrieval agent, do not assume a training rule applies to it. Conversely, a provider may use infrastructure or partners that are not represented by the bot name you expected.
Training-sensitive areas: consider disallowing proprietary research libraries, paid templates and non-public operational documentation, while protecting truly private material with authentication.
Answer-facing pages: preserve access to current service pages, expert explainers, policies and comparison pages when citation and referral value outweigh the exposure.
Operational paths: disallow duplicate filters, staging environments and internal search where crawling adds cost without improving user answers.
A robots.txt rule requests compliant crawlers not to fetch matching paths. It does not make a URL confidential, deindex a page, remove content already collected or prove which requests reached your origin. Use the provider’s current documentation alongside HTTP controls, authentication and legal terms where those mechanisms are relevant.
What llms.txt proposes, and what it cannot promise
An llms.txt file is a proposed convention for publishing a concise, machine-readable map of a site. Usually placed at the root, it uses Markdown-like headings, short descriptions and links to the pages a model or agent may find most useful. Some publishers also offer an expanded llms-full.txt version for longer source material. The proposal aims to reduce discovery friction, not to replace the web.
Think of it as a navigation hint
A well-maintained file can point an agent toward canonical service pages, product documentation, editorial hubs and concise reference material. It can also make the scope of a site easier to understand, particularly where navigation is complex. That makes it a low-risk publishing experiment for teams already maintaining clear information architecture.
Its support remains uneven and is changing. llms.txt is not a universal standard, a mandatory crawler instruction, a licensing agreement or a ranking signal. No major answer engine is obliged to fetch it, follow its links, cite its pages or prefer it over normal crawling and search indexes. It cannot allow or block a bot, and it cannot overrule robots.txt, an X-Robots-Tag, a login wall or a provider’s own retrieval system.
Publish llms.txt only with expectations that match this status: keep it accurate, list canonical URLs, avoid unsupported claims, and measure whether anything uses it. Do not let a supplementary file distract from crawlable pages, clear internal links and reliable canonicalisation. If it becomes stale, it can send agents toward retired pages and dilute rather than clarify the site’s best evidence.
Structure pages so a retrieved passage can stand alone
Access is only the first threshold. A retrieval system still needs to find a passage that answers a question with enough context to use or cite. Pages built for human scanning often help here too, but the test is stricter: if a paragraph is separated from the page, does it identify the subject, make a precise claim, define its limits and remain useful without nearby sales copy?
Use an answer-first passage pattern
Start a section with the direct answer, then add the condition, method, date or scope that qualifies it. Name the entity instead of relying on a vague pronoun. Keep one idea per paragraph where practical, use descriptive headings, and put definitions next to the terms they define. Tables and lists can be useful when they genuinely reduce ambiguity, but the surrounding text should explain what the reader should conclude.
For service content, state who the work is for, what is included, what depends on discovery and what a visitor should do next. Avoid copying a generic answer across multiple pages. Clear page ownership, updated dates where material facts change, first-party evidence and consistent terminology all help a passage remain credible when it travels outside its original page. This is a core application of generative engine optimization, not an argument for writing solely for bots.
Technical foundations still matter. A bot cannot retrieve a useful passage from a page that returns an error, renders essential content only after a fragile script sequence, loops through redirects or is blocked by a challenge page. The technical SEO checks that move rankings are also a useful baseline for making important content reachable and intelligible.
Verify access with server logs, not assumptions
A robots.txt edit is a policy declaration, not evidence of crawler behaviour. Search interfaces and analytics platforms rarely show the complete picture for AI agents. Your CDN, reverse proxy or origin logs are the closer record of whether a request arrived, which path it requested, what response it received and whether an intervention stopped it.
Build a small observation routine
Review a defined period before and after an access change. Segment requests by user agent, requested URL, timestamp, status code, response size, referring host and edge action where available. Confirm that robots.txt itself is reachable, that desired public pages return a clean 200 response, and that disallowed paths are not being served by an accidental alternate route. A 403, 429 or JavaScript challenge may affect a desired retrieval fetcher even when the robots file allows it.
Record the current directive, targeted paths and the reason for the change before deployment.
Monitor the relevant logs after deployment and compare request patterns, status codes and edge decisions.
Validate significant bots using the verification method published by the provider, such as reverse DNS or documented IP ranges, because a user-agent string alone can be spoofed.
Review citations, referral quality and conversion signals alongside crawl data before expanding a rule sitewide.
Logs show requests, not the full reasoning inside an answer engine. Absence of a visible request does not prove a platform lacks your content, and a successful fetch does not prove it will cite you. Still, this evidence is stronger than guessing from a robots file or relying on a single test prompt.
Choose a policy that protects content without hiding the brand
The practical goal is not maximum openness or maximum restriction. It is a policy that distinguishes material with long-term proprietary value from material designed to earn discovery. A public methodology overview may be valuable as a citable proof of expertise; the paid implementation detail behind it may need stronger access control. Both choices can be correct when they are deliberate.
Make the trade-off reviewable
Assign a content owner and answer four questions for each meaningful content zone: What access is being considered? Which named agents and purposes does it affect? What benefit or exposure is expected? What evidence will trigger a review? Then pair the rule with an operational check. Revisit the policy when a provider changes documentation, a new bot appears in logs, a commercial offer changes, or answer visibility becomes material to acquisition.
For many organisations, a balanced starting point is to protect private and paid material with real access controls, apply focused training restrictions where appropriate, and keep high-value public pages available for compliant search and retrieval systems. Context Root can turn that policy into an auditable crawl map, page plan and measurement routine through generative engine optimization services. The important discipline is to test the choice, document its consequences and adjust as platform support evolves.



