The practical protocol, markup, and content decisions that let machines fetch, understand, and quote your pages

Rod Stockebrand
Co-founder, Brandleap.ai

Key Takeaways
Short on time? Here are the top things to know.
Article framework
What is the first requirement for an AI crawler to read a web page?
Does robots.txt control whether a page is used to train an AI model?
How should a page expose content to retrieval systems?
What makes a passage easy for an answer engine to retrieve?
How can a team validate AI-crawler readability?
“Make it easy for AI to read” sounds like a copywriting brief. It is not. Before an answer engine can quote your explanation, a crawler must resolve the URL, obtain the response, decide what it is allowed to fetch, and extract useful content from it. A retrieval system then has to split that content into passages without destroying the relationships between facts.
Those are separate failure points. A page can be excellent for a human and invisible to a fetcher because it returns a challenge page. It can be crawlable but empty in the initial HTML because the answer arrives after JavaScript runs. It can render beautifully but produce chunks such as “it starts at $49” with no product name attached. And a page can be perfectly readable while its canonical and structured-data signals point somewhere else.
Think in layers: access determines whether a system can fetch the page; rendering determines what it can see; semantics determine what the content means; chunk boundaries determine whether that meaning survives retrieval. None of these layers replaces the others.
Start with the URL a user and a crawler are supposed to use. It should be public when the content is public, served over HTTPS, and stable enough to be linked. Return the correct status code: a successful document should not masquerade as a 200 response when it is actually a not-found page. Redirect obsolete URLs deliberately, and do not send every unknown path to the homepage. A retrieval system cannot build reliable evidence from a route that lies about what it contains.
Inspect the response as a client that does not have your browser cookies, your logged-in session, or your JavaScript runtime. Check the final URL after redirects, Content-Type, compression, caching headers, and whether a CDN or WAF serves a verification page to an unfamiliar user agent. Rate limiting is reasonable; indiscriminate blocking of legitimate fetches is a visibility decision.
curl -I https://example.com/guides/product-comparison curl -L --compressed https://example.com/guides/product-comparison | head -80
This is not a request to whitelist every bot. Identify the crawlers you choose to support, monitor their requests, and apply sensible limits. The point is to know what an unauthenticated client actually receives, rather than assuming the browser view is the network view.
The Robots Exclusion Protocol, standardized as RFC 9309, defines how crawlers discover rules in /robots.txt. It is a convention that tells a matching user agent which paths a site operator requests it not to fetch. It is not authentication, encryption, or a universal legal license. A disallow rule cannot protect a private document that is reachable by other means.
Be precise about the audience of a rule. Search crawlers, user-triggered retrieval crawlers, and model-training crawlers may use different user-agent names and policies. Allowing Googlebot does not answer what you want OpenAI’s training crawler, Anthropic’s crawler, or another provider to do. Provider documentation and terms may offer separate controls; review those controls with whoever owns your data and licensing policy.
Two decisions that are often confused
✗ Un-optimized
Crawl access: “May this named crawler fetch /docs/?”
✓ Triple-rich rewrite
Training or reuse permission: “May this operator use fetched content for model training or another purpose?”
Do not use robots.txt to hide confidential information, and do not infer training permission from search visibility. Document your policy, configure the relevant provider controls, and verify behavior from server logs. Retrieval access is a technical path; permission is a governance decision.
Modern browsers can turn a small HTML shell into a rich application. That is useful for interaction, but it is a fragile place to store the only copy of a definition, price, specification, or policy. Search engines may render JavaScript in a later stage; not every retrieval fetcher will, and rendering can fail because of blocked scripts, timeouts, consent gates, or an API that requires a session.
Server-side rendering, static generation, or progressive enhancement gives important content a dependable baseline. Send the title, introductory answer, headings, body copy, links, and key data in the document response. Use JavaScript to enhance filtering, calculators, and navigation, not to make the page’s only meaningful sentence appear after a click.
Semantic HTML is a compact map of the document. Use one clear h1, a logical h2 and h3 hierarchy, p for prose, ol for ordered procedures, ul for unordered collections, table elements for tabular relationships, and descriptive a or nav elements for links. This helps assistive technology and gives extraction software stronger boundaries than a page made from anonymous div elements.
Write the subject into the sentence. “It supports SSO” is a poor standalone passage; “Acme Cloud supports SSO on its Enterprise plan” carries the entity, capability, and scope. Put units, currency, geography, version, and effective date next to the value they qualify. A model should not have to reconstruct whether $49 means monthly or annual pricing from a distant toggle.
<article>
<h1>Acme Cloud SSO</h1>
<p>Acme Cloud supports SAML SSO on the Enterprise plan for organizations
with 100 or more seats.</p>
<h2>How to enable SSO</h2>
<ol>
<li>Open Admin settings and select Identity.</li>
<li>Upload your SAML metadata and save the connection.</li>
</ol>
</article>Notice what this avoids: a heading that says “Enterprise features,” a screenshot instead of instructions, and a sentence that depends on a previous paragraph to identify “it.” The markup is not decoration. It tells a retriever where a fact starts, what kind of relationship it has, and which steps belong together.
One fact often has several URLs: tracking parameters, print views, language variants, filters, and trailing-slash alternatives. Choose the preferred URL and expose it with a rel="canonical" link. Link to that URL internally, include it in the sitemap, and make sure redirects and hreflang decisions do not contradict it. A canonical hint is a signal, not a magic command, so consistency across the site matters.
Structured data provides another machine-readable description of the page. Use Schema.org JSON-LD for the type and properties that genuinely describe visible page content: Article, Product, Organization, FAQPage where appropriate, and so on. Keep the markup current with the prose. Do not add invented reviews, hidden prices, or properties that are not supported by the page merely because a rich-result guide lists them.
Structured data does not guarantee a search feature or an AI citation. Its job is to reduce ambiguity about entities, dates, authorship, products, and relationships. Validate the JSON-LD with the Schema Markup Validator, and use Google’s Rich Results Test when you are checking eligibility for Google features. Passing a validator proves syntax and interpretation; it does not prove that a crawler can reach the page or that the prose is useful.
An XML sitemap is a discovery aid. List the canonical, indexable URLs that matter, keep the file valid, and reference it from robots.txt or your search-console property. A sitemap is not a guarantee of crawling, a ranking boost, or a replacement for internal links. It is especially helpful for a new site, a large site, or pages that are not reached easily through navigation.
Change signals help a crawler decide when a fetch is worth repeating. Return a truthful Last-Modified value and support conditional requests with ETag where your infrastructure can do so. Update those values when the representation materially changes; do not touch every timestamp on every deploy. In sitemaps, lastmod should describe the last significant page update, not the time your build ran. Fake freshness creates noise and makes real changes harder to prioritize.
Retrieval systems rarely hand an entire page to a model. They select passages. You cannot control every chunking algorithm, but you can make every likely boundary less destructive. Give a section a descriptive heading, answer its implied question early, and keep the definition, qualifiers, example, and exception in a coherent neighborhood.
Avoid “as mentioned above,” “the table below,” and “this option” when the referent matters. Repeat a product or organization name when a new section could be retrieved alone. Put a single topic in a paragraph instead of mixing eligibility, pricing, setup, and a sales pitch. Use lists for steps and comparison tables when row and column headers are essential to interpretation.
A useful editorial test: copy one paragraph without its heading or surrounding design. If a reader cannot identify the subject, claim, scope, and time frame, the passage is not ready to be retrieved on its own.
Run this as a release check, not as a one-time SEO project. First request the production URL with curl and save the response. Confirm status, final URL, content type, title, canonical, visible text, links, and structured data. Then fetch robots.txt and the sitemap independently. Test the page with JavaScript disabled and compare the visible facts with the rendered version.
Next, test retrieval rather than merely crawling. Build a small set of real customer questions: definition, comparison, price, eligibility, troubleshooting, and “best for” questions. Retrieve your own pages with the same parser or indexing stack your product uses, inspect the selected passages, and ask whether each answer preserves its qualifiers. Record the URL and passage so a content change can be evaluated against a baseline.
The goal is not to trick a model into citing you. It is to publish an honest, stable representation of your knowledge that a machine can fetch and a person can verify. Better access makes the page available; better structure makes it interpretable; better writing makes it worth retrieving.
Primary references: RFC 9309 (Robots Exclusion Protocol), RFC 9110 (HTTP Semantics), Google Search Central documentation, Schema.org, and the Sitemap protocol.