The proposed markdown convention, its useful applications, and the limits marketers should understand

Rod Stockebrand
Co-founder, Brandleap.ai

Key Takeaways
Short on time? Here are the top things to know.
Article framework
What is llms.txt?
Is llms.txt an official Google or OpenAI standard?
Does llms.txt control crawling or training?
What belongs in a useful file?
Should every company create one?
The idea behind llms.txt is straightforward. Websites contain navigation, scripts, cookie notices, repeated chrome, and thousands of pages. A language-model tool may benefit from a short, human-readable map that names the site and points to the pages an author considers authoritative. The proposal places that map at /llms.txt and uses Markdown so both people and programs can read it.
The excitement is understandable, but the implementation claim must remain modest. llms.txt is not an instruction that every model must follow. It is not a replacement for a sitemap, robots.txt, canonical tags, or documentation architecture. The important distinction is between a useful published resource and a magical visibility switch.
Treat llms.txt as a curated index card for willing tools—not as a control plane for the web or a guarantee of AI citations.
The proposal describes a Markdown document with a required H1 naming the project, an optional blockquote summarizing context, an optional section of instructions, and H2 sections containing lists of links. A link can include a title and description. A special “Optional” section is intended for less important material that can be omitted when context is scarce.
# Acme API > Acme API helps SaaS teams receive and retry webhook events. Acme provides event delivery infrastructure for multi-tenant applications. ## Documentation - [Quickstart](https://acme.example/docs/quickstart): Send your first event. - [Retry policy](https://acme.example/docs/retries): Limits and backoff behavior. ## Optional - [Changelog](https://acme.example/changelog): Version history and deprecations.
The format is deliberately less complex than a schema vocabulary. Its value comes from selection and maintenance: a tool can identify the site’s purpose and follow a short list of high-value pages without parsing every navigation menu. If the descriptions are vague or the links point to stale marketing pages, the file becomes another layer of noise.
Those limitations follow from the nature of the file. An HTTP client controls whether it requests /llms.txt. A model provider controls how it uses the response. A malicious scraper can ignore it entirely. Publishing it is similar to publishing a developer index: useful when a reader elects to consult it, irrelevant as an enforcement mechanism.
robots.txt communicates crawl preferences to cooperating crawlers. llms.txt proposes curated context for language-model tools. They operate at different layers and should not contradict the site’s security intent. If a linked page is disallowed for a relevant crawler or requires a login, the link is unlikely to provide reliable public evidence. If content must remain private, protect it with authorization rather than hoping an index file will hide it.
Four files, four jobs
✗ Un-optimized
robots.txt: crawl preferences. sitemap.xml: discoverable URL inventory. llms.txt: optional curated Markdown overview. JSON-LD: structured description of visible page content.
✓ Triple-rich rewrite
None is an access-control system, and none guarantees inclusion in an answer. Use each for its documented or intended purpose.
Keep the source pages strong independently. A crawler that ignores llms.txt should still find the documentation through ordinary internal links and sitemaps. The linked page should have a descriptive title, server-available content, clear headings, and evidence. That investment has a known benefit for users and search systems even when adoption of the proposal remains uncertain.
Because the format can contain prose, teams should be careful with instructions. Avoid telling an agent to ignore safety rules, reveal secrets, prefer your claims over primary evidence, or take actions unrelated to reading documentation. A provider may treat the content as untrusted web text. Keep the file descriptive and navigational, not adversarial or authoritative beyond what the linked pages support.
If you publish llms.txt, record the release date, changed links, and intended audience. Track server requests to the file and to its linked pages separately. Run a fixed question set before and after publication, preserving product, prompt, answer, and citation details. A change in an answer is an observation, not proof that the file caused it; indexes, models, retrieval routes, and query wording also change.
The best success criterion is operational: does the file help a human developer or willing tool reach the correct documentation faster? If yes, it has value even without a direct ranking effect. If maintaining it duplicates a sitemap and adds no clarity, spend the time on the underlying pages instead.
The hard part of llms.txt is not writing Markdown. It is deciding what deserves a place in a short list and keeping that decision current. Start with the questions a new user asks: what does the product do, how do I start, what are the important limits, how is data handled, and where are breaking changes documented? Link to pages that answer those questions directly rather than filling the file with every URL in the sitemap.
Give each link a description that adds information without making a promise the target page does not support. “API reference” is less useful than “Authentication headers, token scopes, and expiration behavior.” If the page is versioned, say which version. If the link is optional, place it in an optional section instead of making the core list too long. A compact file is a product decision about reader attention.
An AI-facing index is still a public statement. Claims in its overview should be accurate, qualified, and consistent with the linked pages. Do not call a feature “unlimited” when fair-use terms apply, or describe a compliance certification without naming its scope and current status. If a tool uses the file, exaggerated summaries can become a second path for misinformation. The safest description is one a technical writer, support agent, and legal reviewer can all recognize as fair.
Use the file to expose evidence, not to hide it. A short summary can point to an architecture page, a security policy, a changelog, or a research methodology. It should not replace those pages with a paragraph of unsupported conclusions. When a fact matters, let the canonical page carry the detail, date, owner, and source.
A simple script can catch broken links and duplicate destinations, while a content owner can review meaning. Do not automatically generate the file from every navigation item unless the result is manually curated. Navigation often contains account routes, utility pages, duplicate localization paths, and pages designed for humans already familiar with the product. The proposal’s usefulness comes from editorial compression, not exhaustive enumeration.
Adoption may remain uneven. Some providers may fetch the file, some may ignore it, and others may build a different convention. Keep the investment reversible and make the underlying information architecture the priority. A well-written llms.txt can be removed without damaging discovery because every important page remains linked, indexed, and useful on its own. That is a good test for whether the experiment is healthy.
The broader lesson is valuable even if the filename changes. Answer systems benefit when organizations publish concise descriptions, stable canonical pages, explicit relationships, and evidence that can be inspected. Those practices improve developer experience and conventional search today. The file is merely one possible interface to that work.
Primary references: the llmstxt.org proposal; RFC 9309 for the distinct purpose of robots.txt; Google Search Central documentation on AI features and sitemaps; and Schema.org structured-data guidance.