Course Content
RAG Systems
12 sections · 66 lessons
What is the specific role of BeautifulSoup in document ingestion?
What you need to know
A complete cleaning pass
A bank's FAQ page has a nav bar, a cookie banner, a footer and a tracking script around a short answer and a rate table. This code keeps only what a customer question could need:
1from bs4 import BeautifulSoup23soup = BeautifulSoup(open("page.html", encoding="utf-8").read(), "html.parser")45for tag in soup(["script", "style", "nav", "footer"]):6 tag.decompose() # delete boilerplate elements7for tag in soup.select(".cookie-banner"):8 tag.decompose()910main = soup.select_one("main") or soup.body # keep the content area11for table in main.find_all("table"): # table -> one line per row12 rows = [" | ".join(c.get_text(strip=True) for c in tr.find_all(["th", "td"]))13 for tr in table.find_all("tr")]14 table.replace_with("\n".join(rows))1516text = main.get_text("\n", strip=True)17meta = {"title": soup.title.get_text(strip=True),18 "last_modified": soup.find("meta", attrs={"name": "last-modified"})["content"]}The resulting text is:
Fixed Deposit FAQsCan I withdraw my FD before maturity?Yes. Premature withdrawal is allowed. The rate is reduced by 1% from the rate for the period held.Tenure | General | Senior citizen1 year | 6.60% | 7.10%3 years | 6.75% | 7.25%Every table row keeps its column meaning ("1 year | 6.60% | 7.10%" under a header row). A plain get_text() would have produced "1 year6.60%7.10%" mixed in with "Home Loans Cards We use cookies…".
SoupStrainer in WebBaseLoader
1import bs42from langchain_community.document_loaders import WebBaseLoader34loader = WebBaseLoader(5 web_paths=["https://example.com/help/fixed-deposits"],6 bs_kwargs={"parse_only": bs4.SoupStrainer(["main", "title"])},7)parse_only tells BeautifulSoup to build the tree only for matching elements, so the rest of the page is never kept. That is faster, uses less memory, and removes boilerplate in one line when the site has a clean content element.
What BeautifulSoup cannot do
- Run JavaScript. Single-page apps send almost empty HTML. Render them first with a headless browser (for example Playwright), then parse the result.
- Guess the main content. You tell it which selector to keep. For thousands of different sites, a content-extraction library such as trafilatura, which guesses the main text, saves writing a rule per site.
A real-life example
An e-commerce company indexes 4,000 help-centre pages for its support bot. The first crawl used plain get_text(). Each page carried about 300 words of menu, footer and "Was this helpful?" text around a 150-word answer. The top result for "How do I return a damaged item?" was often a page about gift cards, because the menu text containing "Returns" appeared on every page.
With a BeautifulSoup pass that kept only the article element and turned tables into rows, the indexed text fell by about two thirds, embedding cost fell with it, and the returns page came first again.
Follow-up questions to expect
- "Which parser backend would you use?" —
html.parserneeds no install;lxmlis faster on large crawls. Both work with the same BeautifulSoup API. - "How do you keep links useful?" — Replace an anchor with its text plus URL, or store the URLs in metadata, so "click here" does not lose its target.
- "How do you deal with a site that changes its layout?" — Monitor the share of pages where your selector finds nothing, and alert when it rises.