RAG Systems

Course Content

RAG Systems

12 sections · 66 lessons

Why is customized HTML parsing critical for a RAG system?


What survives from a 1,800-token bank FAQ pageMenus, banner, footer: 1,100 tokens cutMain content kept: about 700 tokensRate table rewritten as one row per lineHeading path prepended to every chunk
The same menu text on 900 pages made the home-loan FAQ look relevant to a credit-card question.

What you need to know

Three ways bad HTML parsing hurts retrieval

  1. Near-duplicate chunks. If 40% of every page's text is the same menu and footer, every chunk's embedding is pulled toward that shared text. Similarity scores between unrelated pages rise, and ranking becomes noisy.
  2. Lost structure. A table flattened into "1 year 6.60% 7.10% 3 years 6.75% 7.25%" loses which number belongs to which column. The model then guesses the mapping.
  3. Lost context. A paragraph that says "This is waived for senior citizens" means little without its heading, "Processing fee — Personal loans".

What a custom parser does

  • Selects the main content element for each site template.
  • Removes script, style, nav, footer, cookie banners and "related articles" blocks.
  • Converts tables to Markdown or row-per-line text, keeping the header row.
  • Keeps pre and code blocks whole, so code is not split mid-line.
  • Resolves link text ("see here") into something meaningful or stores the URL.
  • Records the heading path of each section and adds it to its chunks.

Carrying headings into chunks

LangChain's HTMLHeaderTextSplitter splits a page at headings and records the heading path in metadata. Prepend that path to the text so the embedding "knows" the section:

Python
from langchain_text_splitters import HTMLHeaderTextSplitterhtml = """<main><h1>Fixed Deposit FAQs</h1><h2>Premature withdrawal</h2><p>Allowed. The rate is reduced by 1%.</p><h2>Interest payout</h2><p>Interest is paid quarterly or at maturity.</p></main>"""splitter = HTMLHeaderTextSplitter(headers_to_split_on=[("h1", "page"), ("h2", "section")])for doc in splitter.split_text(html):    if doc.page_content in doc.metadata.values():        continue                                        # skip heading-only pieces    heading = " > ".join(doc.metadata.values())    doc.page_content = f"{heading}\n{doc.page_content}"  # chunk carries its path

The chunks become "Fixed Deposit FAQs > Premature withdrawal\nAllowed. The rate is reduced by 1%." and "Fixed Deposit FAQs > Interest payout\nInterest is paid quarterly or at maturity.". Without the prefix, "Allowed. The rate is reduced by 1%." could be about any product. The splitter also emits heading-only pieces, which the if skips.

This is a simple, free form of the contextual retrieval idea covered in the chunking section: give each chunk enough context to stand alone.

A real-life example

A bank's product-FAQ bot indexes the public site. The team counts tokens on a sample page: about 1,800 tokens of HTML text, of which about 1,100 are navigation, footer and legal links, identical on all 900 pages.

Before custom parsing, the query "credit card annual fee" returned the home-loan FAQ in second place, because both pages shared the same large block of menu text that includes "Credit cards" and "Fees". After parsing with a per-template content selector, table conversion and heading prefixes, the indexed text shrank to about 700 tokens per page. On their 60-question test set, the right page appears in the top 3 far more often, and the embedding bill for a full re-index fell by more than half.

The review habit they kept: after every crawl, print five random chunks and read them. It takes two minutes and catches template changes before users do.

Follow-up questions to expect

  • "How would you handle hundreds of sites with different layouts?" — Use a general main-content extractor as the default, and write custom rules only for the few high-traffic sites where quality matters most.
  • "Would you convert HTML to Markdown first?" — Often yes. Markdown keeps headings, lists and tables in a compact form that both splitters and models read well.
  • "How do you detect boilerplate automatically?" — Count how often each line or block appears across pages; text that appears on most pages is boilerplate.