RAG Systems

Course Content

RAG Systems

12 sections · 66 lessons

What is the specific role of BeautifulSoup in document ingestion?


What you need to know

A complete cleaning pass

A bank's FAQ page has a nav bar, a cookie banner, a footer and a tracking script around a short answer and a rate table. This code keeps only what a customer question could need:

Python
from bs4 import BeautifulSoupsoup = BeautifulSoup(open("page.html", encoding="utf-8").read(), "html.parser")for tag in soup(["script", "style", "nav", "footer"]):    tag.decompose()                                   # delete boilerplate elementsfor tag in soup.select(".cookie-banner"):    tag.decompose()main = soup.select_one("main") or soup.body           # keep the content areafor table in main.find_all("table"):                  # table -> one line per row    rows = [" | ".join(c.get_text(strip=True) for c in tr.find_all(["th", "td"]))            for tr in table.find_all("tr")]    table.replace_with("\n".join(rows))text = main.get_text("\n", strip=True)meta = {"title": soup.title.get_text(strip=True),        "last_modified": soup.find("meta", attrs={"name": "last-modified"})["content"]}

The resulting text is:

Text
Fixed Deposit FAQsCan I withdraw my FD before maturity?Yes. Premature withdrawal is allowed. The rate is reduced by 1% from the rate for the period held.Tenure | General | Senior citizen1 year | 6.60% | 7.10%3 years | 6.75% | 7.25%

Every table row keeps its column meaning ("1 year | 6.60% | 7.10%" under a header row). A plain get_text() would have produced "1 year6.60%7.10%" mixed in with "Home Loans Cards We use cookies…".

SoupStrainer in WebBaseLoader

Python
import bs4from langchain_community.document_loaders import WebBaseLoaderloader = WebBaseLoader(    web_paths=["https://example.com/help/fixed-deposits"],    bs_kwargs={"parse_only": bs4.SoupStrainer(["main", "title"])},)

parse_only tells BeautifulSoup to build the tree only for matching elements, so the rest of the page is never kept. That is faster, uses less memory, and removes boilerplate in one line when the site has a clean content element.

What BeautifulSoup cannot do

  • Run JavaScript. Single-page apps send almost empty HTML. Render them first with a headless browser (for example Playwright), then parse the result.
  • Guess the main content. You tell it which selector to keep. For thousands of different sites, a content-extraction library such as trafilatura, which guesses the main text, saves writing a rule per site.

A real-life example

An e-commerce company indexes 4,000 help-centre pages for its support bot. The first crawl used plain get_text(). Each page carried about 300 words of menu, footer and "Was this helpful?" text around a 150-word answer. The top result for "How do I return a damaged item?" was often a page about gift cards, because the menu text containing "Returns" appeared on every page.

With a BeautifulSoup pass that kept only the article element and turned tables into rows, the indexed text fell by about two thirds, embedding cost fell with it, and the returns page came first again.

Follow-up questions to expect

  • "Which parser backend would you use?" — html.parser needs no install; lxml is faster on large crawls. Both work with the same BeautifulSoup API.
  • "How do you keep links useful?" — Replace an anchor with its text plus URL, or store the URLs in metadata, so "click here" does not lose its target.
  • "How do you deal with a site that changes its layout?" — Monitor the share of pages where your selector finds nothing, and alert when it rises.