Back to Blog
AI SearchSeptember 11, 202611 min read

Passage-Level Retrieval: How to Structure Content for Chunking

Write pages whose useful answers survive the trip from document to retrieved context

Rod Stockebrand

Rod Stockebrand

Co-founder, Brandleap.ai

Passage-Level Retrieval: How to Structure Content for Chunking

Key Takeaways

Short on time? Here are the top things to know.

Article framework

How the key ideas connect

1

What is passage-level retrieval?

2

What makes a passage retrievable?

3

How large should content chunks be?

4

Do headings affect retrieval?

5

Should every paragraph target a keyword?

A visual map of the five concepts developed in this article. Read from left to right.

The page may not be the retrieval unit

A search result can point to a page, but an answer system often needs smaller evidence. It may split a document into passages, embed those passages, retrieve a handful of candidates, rerank them, and give only the winners to a language model. If your page is excellent as a complete essay but every paragraph depends on a table, sidebar, or prior definition, it can lose meaning at the moment it is selected.

Chunking is not one standardized operation. A pipeline may split by headings, paragraphs, sentences, HTML nodes, token windows, or a hybrid of these. Some systems overlap windows; others preserve document metadata; others use a model to identify semantic units. Because the implementation is usually private, the durable authoring principle is not “write exactly 512 tokens.” It is “make each important section carry its own context.”

Write for extraction: if a paragraph were pasted into an answer with its heading, URL, and author, would a careful reader know what it means and when it applies?

A chunk needs an anchor

Pronouns and promotional abstractions are cheap to write and expensive to retrieve. “It supports them out of the box” leaves three unresolved questions: what is it, who are they, and what does supports mean? Replace the sentence with the entity, capability, audience, and boundary. This improves human scanning as well as machine matching.

A passage before and after extraction

✗ Un-optimized

“It handles retries automatically and works for growing teams. Learn more below.”

✓ Triple-rich rewrite

“Webhookly retries failed webhook deliveries up to five times with exponential backoff. The policy applies to paid workspaces and can be changed per endpoint.”

The second passage is not stuffed with synonyms. It is simply portable. A retriever can match “webhook retries,” “five times,” “exponential backoff,” or “paid workspaces,” while a reader can evaluate the scope without opening another panel.

Headings are metadata in plain sight

A heading tells both a reader and a parser what follows. “Limits,” “Details,” and “More information” are weak labels. “Webhook retry limits by plan” is a useful label because it names the object and relationship. Use one H1 for the page’s main subject, then a logical H2/H3 hierarchy. Do not choose headings only for visual size; CSS can make a semantic H2 look however the design requires.

  • Use question-shaped headings when the section answers a real question.
  • Name the entity in the heading or the first sentence.
  • Keep one primary concept per section and split unrelated caveats into their own section.
  • Place a definition before specialized terminology that depends on it.
  • Avoid headings that promise a list while the body supplies an unstructured essay.

Boundaries should preserve meaning

A paragraph is not an island if the missing context changes its truth. Put scope beside the claim: date, geography, plan, version, audience, denominator, or exception. If a table contains the only relationship between a product and a price, repeat the critical relationship in a sentence or ensure the table has proper row and column headers. HTML structure can help, but a chunker may separate a cell from the table caption.

html
<section>
  <h2>Webhook retry limits by plan</h2>
  <p>Webhookly retries failed deliveries five times on every paid plan.</p>
  <table>
    <caption>Maximum automatic retries</caption>
    <tr><th scope="row">Free</th><td>2</td></tr>
    <tr><th scope="row">Paid</th><td>5</td></tr>
  </table>
</section>

The HTML gives browsers and assistive technology more structure than a styled grid of divs. The direct sentence gives a text retriever a robust fallback. Accessibility and AEO frequently point in the same direction: state relationships clearly and do not make visual proximity carry the entire meaning.

Write answers, then evidence

For high-value sections, use a consistent order: direct answer, explanation, qualification, evidence, and next step. The first sentence satisfies a quick question; later sentences provide the material a reranker or evaluator needs to judge it. Cite primary documentation or research near the claim. Put the publication date and methodology beside changing data rather than in a distant footer.

This does not mean every section should sound like a FAQ. A technical guide can have narrative, examples, and argument. It means the factual spine should be explicit. Contextual prose should add understanding, not force a retriever to infer the answer from metaphors.

Chunking and internal links

Internal links provide a second route to context. Link “retry policy documentation” rather than “click here,” and make the destination’s title and opening passage agree with the anchor. A chunk may be selected without the link, so do not outsource the definition to the destination when the source passage makes a standalone claim. Conversely, do not repeat an entire manual on every page; use a concise answer and a canonical deeper resource.

A writer’s passage audit

  • Copy the page title, one heading, and each important paragraph into a plain-text note.
  • Read every passage without images, navigation, cards, or the paragraph before it.
  • Underline undefined pronouns, unnamed products, relative dates, and unsupported numbers.
  • Move qualifications next to the claims they limit.
  • Check that a question can be answered from one or two adjacent passages.
  • Link the original source and state what was measured, when, and for whom.

Engineering teams can complement this editorial audit with a local experiment. Parse the rendered HTML into candidate sections, prepend heading paths, create embeddings with the same or a comparable model used by your application, and test real questions. Inspect both false negatives and false positives. A passage that retrieves for the wrong question may need a clearer scope, not more keywords.

Do not confuse chunk quality with ranking certainty

A beautifully self-contained passage can still lose to a more authoritative, fresher, or better-linked source. Retrieval also depends on the query, index coverage, filters, reranking, and product policy. Passage design increases the chance that your evidence is understandable when selected; it cannot force selection or guarantee a citation.

The practical payoff is broader than AI answers. Clear sections improve search snippets, accessibility, support handoffs, documentation maintenance, and human trust. A page becomes a set of durable claims rather than a visual composition that works only when consumed from top to bottom.

Preserve the path from claim to qualification

A common editorial failure is separating a confident claim from the condition that makes it true. The headline says “supports real-time sync,” a later paragraph explains that sync is limited to the Enterprise plan, and a footnote says the feature is in beta. A passage retriever may select the headline and miss both qualifications. Put the decisive limitation in the same section and repeat it in the answer sentence when it changes the buying decision.

The same principle applies to statistics. “Teams reduce handling time by 42%” is not a portable claim unless the passage says which teams, compared with what baseline, during which period, and how the number was measured. A compact methodology sentence can prevent an answer system from turning an internal observation into a universal promise. Specificity is useful metadata, not needless verbosity.

Use document structure as a chunking hint

HTML gives a parser more options than a flat text export. Use section and article elements where appropriate, real list elements for lists, figure and figcaption for explanatory visuals, and code blocks for code. Keep captions descriptive. When a section has a heading, a definition, and an example, place them in a stable order rather than relying on CSS grid placement to communicate sequence. If a scraper flattens the page, the reading order should still make sense.

html
<article>
  <h2>How long does a failed delivery remain retryable?</h2>
  <p>Webhookly retries a failed delivery for up to 24 hours.</p>
  <p>The 24-hour window starts when the first delivery attempt fails;
  changing the endpoint does not restart the window.</p>
  <figure>
    <img src="/retry-window.svg" alt="Retry attempts spread across a 24-hour window">
    <figcaption>The retry window ends 24 hours after the first failed attempt.</figcaption>
  </figure>
</article>

Not every parser preserves every element, so the text still carries the core answer. The semantic structure adds useful signals for browsers, assistive technology, and systems that retain heading paths or captions during ingestion. It also makes later transformations safer: a documentation platform can generate navigation and excerpts without guessing which sentence is the title or the caveat.

Avoid the giant page and the atomized page

A single enormous page can bury related answers among unrelated topics, while hundreds of tiny pages can remove context and create duplicate, thin routes. Choose boundaries around user tasks and stable concepts. A page about retry behavior can include configuration, limits, and failure handling when those concepts are needed together; split a separate authentication protocol when it has a different audience, lifecycle, and question set. Link the pages explicitly so a reader and a crawler can traverse the concept graph.

Review boundaries when the same passage is copied into many pages. Repetition may be useful for a short definition, but conflicting copies are a maintenance hazard. Keep a canonical explanation, summarize it in related pages, and link to the source. When the claim changes, one owner should be able to identify every derivative that needs review.

Evaluate with questions, not abstract scores

  • Collect real questions from support, sales, documentation search, and onboarding.
  • For each question, identify the passage that should answer it and the adjacent qualification it requires.
  • Test exact wording and paraphrases; note whether the right passage is retrieved, a wrong passage is retrieved, or no answer is found.
  • Inspect the winning and losing passages for entity names, headings, scope, and evidence—not just keyword counts.
  • Revise one boundary at a time and keep the question set stable so improvements are attributable.

This evaluation also reveals when a retrieval problem is not an editorial problem. If the correct passage is clear, current, and accessible but absent from the index, investigate crawl and ingestion. If it is retrieved but the generated answer ignores a nearby caveat, investigate reranking or context assembly. Content structure is one lever in a pipeline; disciplined diagnosis keeps teams from endlessly rewriting a page that is already clear.

Primary references: Lewis et al., “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks”; Microsoft’s guidance on chunking in Azure AI Search; and W3C WAI guidance on headings and tables.

Would your best answer survive extraction?

Brandleap maps customer questions to the passages answer systems can retrieve, then turns weak boundaries into clear, citable evidence.