Course Content
Live Coding Interview Prep
7 sections · 50 lessons
Build a document parser for PDFs and extract text.
What you need to know
A PDF is a set of drawing instructions — "put these glyphs at these coordinates" — not a text file. Extracting text means reconstructing words and lines from positions, which is why results vary between libraries and why some PDFs produce nonsense.
Three kinds of PDF, three strategies:
| kind | how to recognise | strategy |
|---|---|---|
| Born-digital (exported from Word, a website) | extract_text() returns real text | text extraction |
| Scanned (a photo of paper) | extract_text() returns empty or a few characters | OCR: render the page to an image, run Tesseract or a vision model |
| Complex layout (columns, tables, forms) | text comes out interleaved or flattened | layout-aware tools: pdfplumber, PyMuPDF, Docling, unstructured |
Keep the page number. A RAG answer that says "see page 14 of the policy" is checkable; one that says "somewhere in the PDF" is not.
Encrypted PDFs. Many "protected" PDFs use an empty user password and only restrict printing or copying; decrypt("") opens them. In pypdf, decrypt returns a result instead of raising when the password is wrong, so check it.
1import re2from pypdf import PasswordType, PdfReader34def clean(text: str) -> str:5 text = re.sub(r"(\w)-\n(\w)", r"\1\2", text) # "refund-\nable" -> "refundable"6 text = re.sub(r"[ \t]+", " ", text)7 text = re.sub(r"\n{3,}", "\n\n", text)8 return text.strip()910def extract_pdf(path: str, min_chars: int = 20) -> dict:11 """Text per page, with the pages that look scanned listed in needs_ocr."""12 reader = PdfReader(path)13 if reader.is_encrypted and reader.decrypt("") == PasswordType.NOT_DECRYPTED:14 raise ValueError(f"{path} needs a password")15 pages, needs_ocr = [], []16 for number, page in enumerate(reader.pages, start=1):17 text = clean(page.extract_text() or "")18 if len(text) < min_chars:19 needs_ocr.append(number)20 pages.append({"page": number, "text": text, "source": path})21 return {"pages": pages, "needs_ocr": needs_ocr}2223def pdf_chunks(parsed: dict, split, size: int = 800) -> list[dict]:24 """Chunk each page separately so every chunk keeps its page number."""25 return [{"text": piece, "page": p["page"], "source": p["source"]}26 for p in parsed["pages"] for piece in split(p["text"], size)]The tricky parts:
page.extract_text() or ""— some pages returnNone; theorkeeps the rest of the code simple.- The hyphen regex only joins a hyphen that ends a line between two word characters, so a genuine "T-shirt" in the middle of a line is untouched.
min_charsis the scanned-page detector. A scanned page often yields an empty string or a few stray characters from a stamp or page number.- Chunking per page (
splitis any splitter, such asrecursive_splitfrom the chunking question) means a chunk never straddles two pages, so its page number is always right. The cost is that a sentence crossing a page break is split.
Complexity: O(total characters) for extraction and cleaning, O(pages) for the list. reader.pages is lazy, so memory can stay at one page at a time if you process pages as you go instead of collecting them.
A real-life example
A three-page PDF built for the test: page 1 has text with a hyphenated line break, page 2 is blank (standing in for a scanned image), page 3 has text:
1parsed = extract_pdf("policy.pdf")2for p in parsed["pages"]:3 print(p["page"], repr(p["text"]))4print("needs OCR:", parsed["needs_ocr"])5# 1 'Refunds are processed within 5 working days to the original payment method.'6# 2 ''7# 3 'Cash-on-delivery orders are refunded as store credit.'8# needs OCR: [2]| page | raw extraction | after clean | decision |
|---|---|---|---|
| 1 | …within 5 work-\ning days… | …within 5 working days… | index |
| 2 | '' | '' (0 chars, under 20) | send to OCR |
| 3 | Cash-on-delivery orders… | unchanged — the hyphens are mid-line | index |
Page 2 is never silently indexed as empty. In a real pipeline it goes to an OCR queue, and its chunks are added when OCR finishes.
Insurance companies processing claim forms and banks processing KYC documents hit exactly this mix: most pages are digital, a few are phone photos of paper, and those must go through OCR.
Follow-up questions to expect
- "The text comes out with two columns interleaved — what now?" — Use a layout-aware extractor that reads blocks with coordinates (PyMuPDF
get_text("blocks"),pdfplumberwith word positions), sort blocks by column, or use a document-AI parser. - "How do you handle tables?" — Extract them separately (for example with
pdfplumber's table finder) and store each as Markdown or CSV, as its own chunk with the table caption. - "Headers and footers repeat on every page — does that matter?" — Yes: "Confidential — Acme Corp — Page 3" pollutes every chunk's embedding. Detect lines that repeat on most pages and strip them.