A practical guide to deciding which automated agents can discover, render, and reuse your content

Rod Stockebrand
Co-founder, Brandleap.ai

Key Takeaways
Short on time? Here are the top things to know.
Article framework
Does robots.txt control every AI system?
What is the difference between GPTBot and ChatGPT-User?
When should I use noindex or nofollow?
Can I block AI training while allowing search visibility?
How should teams test crawler controls?
Teams often arrive with a one-line request: block AI crawlers. It sounds precise until you ask what “AI crawler” means. A publisher may want search engines to discover an article, answer features to cite it, training crawlers to exclude it, user-triggered fetchers to access it, and private customer material to remain unavailable to everyone. Those are different audiences with different controls.
The web has several layers for expressing those decisions. robots.txt is a site-level crawl preference. A meta robots tag or X-Robots-Tag is a document-level indexing and link directive. Authentication, WAF rules, rate limits, and network controls are enforcement mechanisms. Mixing them creates false confidence: a disallow rule is not a password, while a noindex header cannot help if a crawler is blocked before it can read the response.
Write a crawler policy as a matrix of purpose, user agent, URL scope, and enforcement. “AI” is a product category, not a reliable access-control rule.
The Robots Exclusion Protocol, standardized as RFC 9309, defines a convention for publishing rules at /robots.txt. A crawler identifies itself with a user-agent token, finds the most specific matching group, and evaluates Allow and Disallow paths. The file can also point to sitemap locations. It is intentionally simple and intentionally voluntary: compliant crawlers use it to decide what to request.
User-agent: GPTBot Disallow: / User-agent: OAI-SearchBot Allow: / User-agent: * Disallow: /account/ Disallow: /admin/ Sitemap: https://example.com/sitemap.xml
The example is illustrative, not a universal OpenAI recommendation. Provider documentation is the authority for current crawler names and purposes. Notice the separate rules: one user agent is denied, another is allowed, and sensitive application paths are excluded for everyone. Keep rules readable because the person debugging a visibility incident may not be the person who wrote them.
Robots rules are path-based and host-specific. A policy on https://www.example.com/robots.txt does not automatically control api.example.com, a different scheme, or an image CDN. Redirect chains and cached robots files also matter. Serve a plain-text 200 response from the exact origin and hostname that crawlers use, and avoid generating a different policy for every request unless there is a compelling operational reason.
A user-agent string is a label sent by a client. OpenAI, Google, Anthropic, and other providers publish crawler documentation so site owners can make informed policies, often including reverse-DNS verification guidance. A request that merely claims to be GPTBot is not proof that it came from OpenAI. If preventing unauthorized access matters, use authentication or an allowlist at the network layer; do not rely on the string alone.
This distinction prevents a common mistake: blocking every provider bot because one training policy is unwanted, then wondering why an answer engine cannot cite the public documentation. Decide whether your business objective is search discovery, model training, direct retrieval, or all three.
A page can carry directives in HTML, such as <meta name="robots" content="noindex, nofollow">, or in the HTTP response through X-Robots-Tag. Google documents directives including noindex, nofollow, nosnippet, noarchive, and max-snippet. These are signals a crawler reads after fetching the response. They therefore have a different job from robots.txt: they describe how an accessible resource should be indexed, represented, or followed.
HTTP/1.1 200 OK X-Robots-Tag: noindex, nofollow Content-Type: text/html
Use the header for PDFs, images, feeds, and other non-HTML resources where a meta tag cannot be placed. Apply it at the response that matters, not merely on a redirect. For a page that must remain discoverable but should expose only a short extract, a supported snippet directive may be more appropriate than noindex. Always check the target product’s support: directives are not a universal command language for every answer engine.
Choose the control for the failure you need to prevent
✗ Un-optimized
robots.txt Disallow: “Please do not crawl this path.” It does not authenticate, erase an existing index, or govern a copied URL.
✓ Triple-rich rewrite
noindex / X-Robots-Tag: “If you fetch this response, do not index it.” Authentication: “Only an authorized client can fetch this resource.”
Public marketing pages and private application pages should not share a blanket policy. Allow the crawler access needed to discover stable, useful documentation. Disallow account pages, internal search results, cart states, admin routes, and infinite calendar or filter combinations that create crawl traps. Protect customer records and paid research with real authorization. A robots rule for /private/ is not a data-protection boundary.
For AI search, ask whether a page contains a fact you want represented accurately. If yes, blocking a retrieval crawler removes your opportunity to supply the canonical wording. If no, do not assume that hiding one URL removes the underlying fact from the web. Partner sites, public filings, cached pages, and user-provided documents may still supply it. Access policy reduces one acquisition path; it does not control an ecosystem.
Treat changes like code. Review them, test a staging copy with representative agents, and record who owns the policy. A one-character path mistake can expose an admin route or hide an entire documentation tree. A provider can also introduce a new crawler or change how it identifies one, so schedule periodic review instead of considering the file finished forever.
No robots file guarantees citation, and no crawler rule can guarantee that information will never appear in an AI answer. The responsible goal is narrower: make intentional access decisions, enforce sensitive content with real security controls, and keep the public evidence you want reused available, current, and attributable. That is a much more durable AEO practice than chasing a universal bot blacklist.
A useful policy review asks two questions instead of one. First, may this agent fetch the resource? Second, if it fetches the resource, what kind of downstream use does the publisher permit or expect? The first question is where robots.txt and network controls matter. The second may involve a provider’s terms, licensing agreement, copyright policy, or a contract with a customer. A crawler rule cannot encode every downstream legal or commercial condition, and a legal notice cannot replace a technical access boundary.
This separation also clarifies why “allow search, block training” is a meaningful but incomplete request. A search crawler may create an index used to answer a query, while a training crawler may collect examples for a later model. A provider may expose separate user agents and policies, but another provider may not. Record the provider’s documented purpose, the business reason for your choice, and the residual risk if the agent ignores the convention. That makes the policy reviewable instead of turning it into a permanent argument about labels.
Large sites often have several teams editing access behavior. The SEO team owns robots.txt, security owns the WAF, platform engineering owns CDN rules, and a product team adds a noindex header to a route. Each change may look reasonable in isolation while the combination blocks an important documentation path. A crawler can receive a 403 from the edge even though robots.txt allows the URL, or it can receive a 200 page whose X-Robots-Tag says noindex. Build one inventory that shows all layers and test from outside the corporate network.
The best control is the one a future maintainer can understand. Prefer a few explicit rules over a dense collection of bot-specific exceptions. When an exception is necessary, explain the purpose in version control and link to the provider documentation. Then verify the resulting behavior with logs and a representative fetch. This is ordinary change management, but it is especially valuable when the consequence is that an answer engine can no longer see the evidence your content team worked to publish.
Primary references: RFC 9309, Robots Exclusion Protocol; Google Search Central documentation on robots.txt and robots meta tags; and OpenAI documentation on GPTBot, OAI-SearchBot, and ChatGPT-User.