What the HtmlRAG paper taught me about feeding HTML — not plain text — to LLMs

Rod Stockebrand
Co-founder, Brandleap.ai

Key Takeaways
Short on time? Here are the top things to know.
Article framework
What is HtmlRAG?
Why is HTML better than plain text for RAG?
How does HtmlRAG fit huge HTML pages into an LLM's context window?
Does HtmlRAG actually improve answer accuracy?
What does HtmlRAG cost to run?
Here's a thing almost every RAG pipeline does without thinking twice: it retrieves a web page, strips out all the HTML, and hands the LLM a wall of plain text. It feels like cleanup. It feels responsible, even — who wants to waste context window on angle brackets?
But a paper out of Renmin University and Baidu — "HtmlRAG: HTML is Better Than Plain Text for Modeling Retrieved Knowledge in RAG Systems" (WWW 2025) — makes a case that stopped me mid-scroll: that "cleanup" step is quietly destroying knowledge. Not the words themselves, but the relationships between them.
Think about a pricing table. In HTML, the model can see that the A100 row holds "$3/hr" and the H100 row holds "$5/hr." Flatten it to text and you get "A100 $3/hr H100 $5/hr" — and now the model is guessing which price goes with which GPU. Same words, less knowledge. Headings do this too. So do lists, and code blocks, and every other piece of structure your text converter cheerfully deletes.

If keeping HTML were free, everyone would already do it. It isn't. The paper's running example is 20 retrieved web pages for a single query — about 1.6 million tokens of raw HTML. Most of that is CSS, JavaScript, tracking scripts, and attribute noise that no model needs and no context window can hold.
This is the actual contribution of HtmlRAG. It's not the observation that structure matters — it's the engineering that makes HTML practical. The pipeline squeezes 1.6M tokens down to about 4K, a roughly 400× reduction, while keeping the markup that carries meaning.
It happens in three moves. Move one is boring and brilliant: pure rule-based cleaning. Delete the CSS, the JavaScript, the comments, the bloated tag attributes; merge redundant nested tags. No model involved, and over 94% of the tokens are already gone.
Move two gets smarter. The cleaned HTML is parsed into a DOM tree, then merged into a coarser "block tree" where each block holds up to ~256 words. A small embedding model — BGE, about 200M parameters — scores every block's similarity to the query, and low scorers get greedily pruned. Down to roughly 8K tokens.
Move three is my favorite part. A small fine-tuned LLM (Phi-3.5-mini) looks at a finer block tree — ~128-word blocks — and scores each block by the generation probability of its HTML tag path, something like <html>→<body>→<div>→<p>. The clever bit: the model never generates long text, just short tag paths, so it stays fast while being far more precise than embeddings alone. That takes you to ~4K tokens of dense, relevant, still-valid HTML.

Notice what pruning a block tree buys you that chopping text into chunks never could: the output is still coherent HTML. When you cut the <nav> subtree, the footer links, and the ad <div>, what remains is a well-formed document — the heading still sits above its paragraph, the table still has its rows. Chunk-based text splitting, by contrast, happily slices a table in half and calls it a day.

Papers with a cute idea and cherry-picked results are a dime a dozen, so here's the test that matters: six QA benchmarks, a serious reader model (Llama-3.1-70B-Instruct, 4K context), and strong plain-text baselines including BGE reranking, E5-Mistral, LongLLMLingua, and JinaAI Reader.
HtmlRAG won on all six — ASQA, NQ, TriviaQA, HotpotQA, MuSiQue, and ELI5 — with statistically significant gains. The headline number is ASQA: 68.5 Hit@1 against 62.5 for the best baseline. Six points from changing the input format. And the whole pruning pipeline costs a small fraction of what the 70B model spends generating the answer, so this isn't accuracy bought with latency.

I'm not going to wrap this up with a tidy conclusion, because the paper's real message isn't a summary — it's a dare.
So here's mine to you. Take one pipeline you own that retrieves web content, and go look at what your HTML-to-text converter actually threw away this week. Diff the raw page against what your LLM received. Count the tables that got flattened, the headings that vanished, the lists that turned to mush. Then run ten of your hardest real queries twice — once with your current plain-text input, once with cleaned HTML — and compare the answers side by side.
If the HTML version doesn't win a single one, tell me and I'll be surprised. But if it does — and the numbers above say it will — you've just found accuracy that was sitting in your pipeline all along, waiting for you to stop deleting it.
Source: Tan et al., "HtmlRAG: HTML is Better Than Plain Text for Modeling Retrieved Knowledge in RAG Systems," WWW 2025. arXiv:2411.02959