Back to Blog
AI SearchSeptember 6, 20269 min read

Your RAG Pipeline Is Throwing Away the Good Stuff

What the HtmlRAG paper taught me about feeding HTML — not plain text — to LLMs

Rod Stockebrand

Rod Stockebrand

Co-founder, Brandleap.ai

Your RAG Pipeline Is Throwing Away the Good Stuff

Key Takeaways

Short on time? Here are the top things to know.

Article framework

How the key ideas connect

1

What is HtmlRAG?

2

Why is HTML better than plain text for RAG?

3

How does HtmlRAG fit huge HTML pages into an LLM's context window?

4

Does HtmlRAG actually improve answer accuracy?

5

What does HtmlRAG cost to run?

A visual map of the five concepts developed in this article. Read from left to right.

The habit nobody questions

Here's a thing almost every RAG pipeline does without thinking twice: it retrieves a web page, strips out all the HTML, and hands the LLM a wall of plain text. It feels like cleanup. It feels responsible, even — who wants to waste context window on angle brackets?

But a paper out of Renmin University and Baidu — "HtmlRAG: HTML is Better Than Plain Text for Modeling Retrieved Knowledge in RAG Systems" (WWW 2025) — makes a case that stopped me mid-scroll: that "cleanup" step is quietly destroying knowledge. Not the words themselves, but the relationships between them.

Think about a pricing table. In HTML, the model can see that the A100 row holds "$3/hr" and the H100 row holds "$5/hr." Flatten it to text and you get "A100 $3/hr H100 $5/hr" — and now the model is guessing which price goes with which GPU. Same words, less knowledge. Headings do this too. So do lists, and code blocks, and every other piece of structure your text converter cheerfully deletes.

The same content as HTML and as converted plain text
Figure 1 — The same content as HTML and as converted plain text. The words survive; the structure doesn't.

Okay, but raw HTML is a monster

If keeping HTML were free, everyone would already do it. It isn't. The paper's running example is 20 retrieved web pages for a single query — about 1.6 million tokens of raw HTML. Most of that is CSS, JavaScript, tracking scripts, and attribute noise that no model needs and no context window can hold.

This is the actual contribution of HtmlRAG. It's not the observation that structure matters — it's the engineering that makes HTML practical. The pipeline squeezes 1.6M tokens down to about 4K, a roughly 400× reduction, while keeping the markup that carries meaning.

It happens in three moves. Move one is boring and brilliant: pure rule-based cleaning. Delete the CSS, the JavaScript, the comments, the bloated tag attributes; merge redundant nested tags. No model involved, and over 94% of the tokens are already gone.

Move two gets smarter. The cleaned HTML is parsed into a DOM tree, then merged into a coarser "block tree" where each block holds up to ~256 words. A small embedding model — BGE, about 200M parameters — scores every block's similarity to the query, and low scorers get greedily pruned. Down to roughly 8K tokens.

Move three is my favorite part. A small fine-tuned LLM (Phi-3.5-mini) looks at a finer block tree — ~128-word blocks — and scores each block by the generation probability of its HTML tag path, something like <html>→<body>→<div>→<p>. The clever bit: the model never generates long text, just short tag paths, so it stays fast while being far more precise than embeddings alone. That takes you to ~4K tokens of dense, relevant, still-valid HTML.

The five-stage HtmlRAG pipeline with token counts at each step
Figure 2 — The five-stage HtmlRAG pipeline, with the paper's token counts at each step.

Why pruning a tree beats chopping text

Notice what pruning a block tree buys you that chopping text into chunks never could: the output is still coherent HTML. When you cut the <nav> subtree, the footer links, and the ad <div>, what remains is a well-formed document — the heading still sits above its paragraph, the table still has its rows. Chunk-based text splitting, by contrast, happily slices a table in half and calls it a day.

Block-tree pruning: boilerplate subtrees are cut; answer-bearing blocks survive as valid HTML
Figure 3 — Block-tree pruning: boilerplate subtrees are cut; answer-bearing blocks survive as valid HTML.

The part that made me sit up: it wins everywhere

Papers with a cute idea and cherry-picked results are a dime a dozen, so here's the test that matters: six QA benchmarks, a serious reader model (Llama-3.1-70B-Instruct, 4K context), and strong plain-text baselines including BGE reranking, E5-Mistral, LongLLMLingua, and JinaAI Reader.

HtmlRAG won on all six — ASQA, NQ, TriviaQA, HotpotQA, MuSiQue, and ELI5 — with statistically significant gains. The headline number is ASQA: 68.5 Hit@1 against 62.5 for the best baseline. Six points from changing the input format. And the whole pruning pipeline costs a small fraction of what the 70B model spends generating the answer, so this isn't accuracy bought with latency.

HtmlRAG vs. the best plain-text baseline on six benchmarks
Figure 4 — HtmlRAG vs. the best plain-text baseline on each of the six benchmarks (higher is better).
  • ASQA: 68.5 vs 62.5 Hit@1 (HtmlRAG vs. best plain-text baseline)
  • NQ, TriviaQA, HotpotQA, MuSiQue, ELI5: wins across the board
  • Reader model: Llama-3.1-70B-Instruct at 4K context
  • Baselines beaten: BGE, E5-Mistral, LongLLMLingua, JinaAI Reader

Your challenge

I'm not going to wrap this up with a tidy conclusion, because the paper's real message isn't a summary — it's a dare.

So here's mine to you. Take one pipeline you own that retrieves web content, and go look at what your HTML-to-text converter actually threw away this week. Diff the raw page against what your LLM received. Count the tables that got flattened, the headings that vanished, the lists that turned to mush. Then run ten of your hardest real queries twice — once with your current plain-text input, once with cleaned HTML — and compare the answers side by side.

If the HTML version doesn't win a single one, tell me and I'll be surprised. But if it does — and the numbers above say it will — you've just found accuracy that was sitting in your pipeline all along, waiting for you to stop deleting it.

Source: Tan et al., "HtmlRAG: HTML is Better Than Plain Text for Modeling Retrieved Knowledge in RAG Systems," WWW 2025. arXiv:2411.02959

Want your brand to show up in AI answers?

We audit how AI models perceive and cite your brand today — and build the strategy to change it.