Back to Blog
AI SearchSeptember 9, 202610 min read

What llms.txt Can and Cannot Do for AI Visibility

The proposed markdown convention, its useful applications, and the limits marketers should understand

Rod Stockebrand

Rod Stockebrand

Co-founder, Brandleap.ai

What llms.txt Can and Cannot Do for AI Visibility

Key Takeaways

Short on time? Here are the top things to know.

Article framework

How the key ideas connect

1

What is llms.txt?

2

Is llms.txt an official Google or OpenAI standard?

3

Does llms.txt control crawling or training?

4

What belongs in a useful file?

5

Should every company create one?

A visual map of the five concepts developed in this article. Read from left to right.

A small file carrying a large promise

The idea behind llms.txt is straightforward. Websites contain navigation, scripts, cookie notices, repeated chrome, and thousands of pages. A language-model tool may benefit from a short, human-readable map that names the site and points to the pages an author considers authoritative. The proposal places that map at /llms.txt and uses Markdown so both people and programs can read it.

The excitement is understandable, but the implementation claim must remain modest. llms.txt is not an instruction that every model must follow. It is not a replacement for a sitemap, robots.txt, canonical tags, or documentation architecture. The important distinction is between a useful published resource and a magical visibility switch.

Treat llms.txt as a curated index card for willing tools—not as a control plane for the web or a guarantee of AI citations.

What the proposal contains

The proposal describes a Markdown document with a required H1 naming the project, an optional blockquote summarizing context, an optional section of instructions, and H2 sections containing lists of links. A link can include a title and description. A special “Optional” section is intended for less important material that can be omitted when context is scarce.

markdown
# Acme API

> Acme API helps SaaS teams receive and retry webhook events.

Acme provides event delivery infrastructure for multi-tenant applications.

## Documentation

- [Quickstart](https://acme.example/docs/quickstart): Send your first event.
- [Retry policy](https://acme.example/docs/retries): Limits and backoff behavior.

## Optional

- [Changelog](https://acme.example/changelog): Version history and deprecations.

The format is deliberately less complex than a schema vocabulary. Its value comes from selection and maintenance: a tool can identify the site’s purpose and follow a short list of high-value pages without parsing every navigation menu. If the descriptions are vague or the links point to stale marketing pages, the file becomes another layer of noise.

What it cannot do

  • It cannot force a model, crawler, or answer engine to fetch or cite a URL.
  • It cannot grant access to content protected by authentication or bypass a robots policy.
  • It cannot prevent training, scraping, indexing, or reuse of content.
  • It cannot make inaccurate claims authoritative or turn a non-canonical page into the source of truth.
  • It cannot replace an XML sitemap, internal links, HTML headings, or structured data.

Those limitations follow from the nature of the file. An HTTP client controls whether it requests /llms.txt. A model provider controls how it uses the response. A malicious scraper can ignore it entirely. Publishing it is similar to publishing a developer index: useful when a reader elects to consult it, irrelevant as an enforcement mechanism.

The relationship with robots.txt

robots.txt communicates crawl preferences to cooperating crawlers. llms.txt proposes curated context for language-model tools. They operate at different layers and should not contradict the site’s security intent. If a linked page is disallowed for a relevant crawler or requires a login, the link is unlikely to provide reliable public evidence. If content must remain private, protect it with authorization rather than hoping an index file will hide it.

Four files, four jobs

✗ Un-optimized

robots.txt: crawl preferences. sitemap.xml: discoverable URL inventory. llms.txt: optional curated Markdown overview. JSON-LD: structured description of visible page content.

✓ Triple-rich rewrite

None is an access-control system, and none guarantees inclusion in an answer. Use each for its documented or intended purpose.

How to publish a responsible experiment

  • Create the file at the exact root origin that hosts the public content.
  • Describe the organization or product in one accurate paragraph.
  • Link only to canonical, public, stable pages that you would be comfortable recommending to a customer.
  • Put definitions, API references, policies, and primary research before promotional summaries.
  • Keep link descriptions concrete and update the file when URLs or product behavior change.
  • Monitor requests and referrals, but do not infer causation from an appearance in one answer.

Keep the source pages strong independently. A crawler that ignores llms.txt should still find the documentation through ordinary internal links and sitemaps. The linked page should have a descriptive title, server-available content, clear headings, and evidence. That investment has a known benefit for users and search systems even when adoption of the proposal remains uncertain.

Do not use it as a prompt injection surface

Because the format can contain prose, teams should be careful with instructions. Avoid telling an agent to ignore safety rules, reveal secrets, prefer your claims over primary evidence, or take actions unrelated to reading documentation. A provider may treat the content as untrusted web text. Keep the file descriptive and navigational, not adversarial or authoritative beyond what the linked pages support.

Measure the opportunity honestly

If you publish llms.txt, record the release date, changed links, and intended audience. Track server requests to the file and to its linked pages separately. Run a fixed question set before and after publication, preserving product, prompt, answer, and citation details. A change in an answer is an observation, not proof that the file caused it; indexes, models, retrieval routes, and query wording also change.

The best success criterion is operational: does the file help a human developer or willing tool reach the correct documentation faster? If yes, it has value even without a direct ranking effect. If maintaining it duplicates a sitemap and adds no clarity, spend the time on the underlying pages instead.

A curated file needs a curation process

The hard part of llms.txt is not writing Markdown. It is deciding what deserves a place in a short list and keeping that decision current. Start with the questions a new user asks: what does the product do, how do I start, what are the important limits, how is data handled, and where are breaking changes documented? Link to pages that answer those questions directly rather than filling the file with every URL in the sitemap.

Give each link a description that adds information without making a promise the target page does not support. “API reference” is less useful than “Authentication headers, token scopes, and expiration behavior.” If the page is versioned, say which version. If the link is optional, place it in an optional section instead of making the core list too long. A compact file is a product decision about reader attention.

Avoid turning documentation into marketing copy

An AI-facing index is still a public statement. Claims in its overview should be accurate, qualified, and consistent with the linked pages. Do not call a feature “unlimited” when fair-use terms apply, or describe a compliance certification without naming its scope and current status. If a tool uses the file, exaggerated summaries can become a second path for misinformation. The safest description is one a technical writer, support agent, and legal reviewer can all recognize as fair.

Use the file to expose evidence, not to hide it. A short summary can point to an architecture page, a security policy, a changelog, or a research methodology. It should not replace those pages with a paragraph of unsupported conclusions. When a fact matters, let the canonical page carry the detail, date, owner, and source.

Operational checks before publication

  • Confirm /llms.txt returns a UTF-8 text response with a stable 200 status.
  • Resolve every link and check that it does not redirect to a login, tracking URL, or obsolete host.
  • Compare the overview with the current product name, policies, and documented capabilities.
  • Check that linked pages are crawlable and contain the facts their descriptions promise.
  • Add the file to the release checklist for URL migrations and major documentation changes.
  • Log requests, but do not treat a request as proof that an answer system used the content.

A simple script can catch broken links and duplicate destinations, while a content owner can review meaning. Do not automatically generate the file from every navigation item unless the result is manually curated. Navigation often contains account routes, utility pages, duplicate localization paths, and pages designed for humans already familiar with the product. The proposal’s usefulness comes from editorial compression, not exhaustive enumeration.

Design for an uncertain ecosystem

Adoption may remain uneven. Some providers may fetch the file, some may ignore it, and others may build a different convention. Keep the investment reversible and make the underlying information architecture the priority. A well-written llms.txt can be removed without damaging discovery because every important page remains linked, indexed, and useful on its own. That is a good test for whether the experiment is healthy.

The broader lesson is valuable even if the filename changes. Answer systems benefit when organizations publish concise descriptions, stable canonical pages, explicit relationships, and evidence that can be inspected. Those practices improve developer experience and conventional search today. The file is merely one possible interface to that work.

Primary references: the llmstxt.org proposal; RFC 9309 for the distinct purpose of robots.txt; Google Search Central documentation on AI features and sitemaps; and Schema.org structured-data guidance.

Want an AI visibility plan without the hype?

Brandleap separates documented behavior from speculation, then prioritizes the crawl, content, and evidence changes that can be tested.