Live Coding Interview Prep

Course Content

Live Coding Interview Prep

7 sections · 50 lessons

Build a document parser for PDFs and extract text.


What you need to know

A PDF is a set of drawing instructions — "put these glyphs at these coordinates" — not a text file. Extracting text means reconstructing words and lines from positions, which is why results vary between libraries and why some PDFs produce nonsense.

Three kinds of PDF, three strategies:

kindhow to recognisestrategy
Born-digital (exported from Word, a website)extract_text() returns real texttext extraction
Scanned (a photo of paper)extract_text() returns empty or a few charactersOCR: render the page to an image, run Tesseract or a vision model
Complex layout (columns, tables, forms)text comes out interleaved or flattenedlayout-aware tools: pdfplumber, PyMuPDF, Docling, unstructured

Keep the page number. A RAG answer that says "see page 14 of the policy" is checkable; one that says "somewhere in the PDF" is not.

Encrypted PDFs. Many "protected" PDFs use an empty user password and only restrict printing or copying; decrypt("") opens them. In pypdf, decrypt returns a result instead of raising when the password is wrong, so check it.

Python
import refrom pypdf import PasswordType, PdfReaderdef clean(text: str) -> str:    text = re.sub(r"(\w)-\n(\w)", r"\1\2", text)      # "refund-\nable" -> "refundable"    text = re.sub(r"[ \t]+", " ", text)    text = re.sub(r"\n{3,}", "\n\n", text)    return text.strip()def extract_pdf(path: str, min_chars: int = 20) -> dict:    """Text per page, with the pages that look scanned listed in needs_ocr."""    reader = PdfReader(path)    if reader.is_encrypted and reader.decrypt("") == PasswordType.NOT_DECRYPTED:        raise ValueError(f"{path} needs a password")    pages, needs_ocr = [], []    for number, page in enumerate(reader.pages, start=1):        text = clean(page.extract_text() or "")        if len(text) < min_chars:            needs_ocr.append(number)        pages.append({"page": number, "text": text, "source": path})    return {"pages": pages, "needs_ocr": needs_ocr}def pdf_chunks(parsed: dict, split, size: int = 800) -> list[dict]:    """Chunk each page separately so every chunk keeps its page number."""    return [{"text": piece, "page": p["page"], "source": p["source"]}            for p in parsed["pages"] for piece in split(p["text"], size)]

The tricky parts:

  • page.extract_text() or "" — some pages return None; the or keeps the rest of the code simple.
  • The hyphen regex only joins a hyphen that ends a line between two word characters, so a genuine "T-shirt" in the middle of a line is untouched.
  • min_chars is the scanned-page detector. A scanned page often yields an empty string or a few stray characters from a stamp or page number.
  • Chunking per page (split is any splitter, such as recursive_split from the chunking question) means a chunk never straddles two pages, so its page number is always right. The cost is that a sentence crossing a page break is split.

Complexity: O(total characters) for extraction and cleaning, O(pages) for the list. reader.pages is lazy, so memory can stay at one page at a time if you process pages as you go instead of collecting them.

A real-life example

A three-page PDF built for the test: page 1 has text with a hyphenated line break, page 2 is blank (standing in for a scanned image), page 3 has text:

Python
parsed = extract_pdf("policy.pdf")for p in parsed["pages"]:    print(p["page"], repr(p["text"]))print("needs OCR:", parsed["needs_ocr"])# 1 'Refunds are processed within 5 working days to the original payment method.'# 2 ''# 3 'Cash-on-delivery orders are refunded as store credit.'# needs OCR: [2]
pageraw extractionafter cleandecision
1…within 5 work-\ning days……within 5 working days…index
2'''' (0 chars, under 20)send to OCR
3Cash-on-delivery orders…unchanged — the hyphens are mid-lineindex

Page 2 is never silently indexed as empty. In a real pipeline it goes to an OCR queue, and its chunks are added when OCR finishes.

Insurance companies processing claim forms and banks processing KYC documents hit exactly this mix: most pages are digital, a few are phone photos of paper, and those must go through OCR.

Follow-up questions to expect

  • "The text comes out with two columns interleaved — what now?" — Use a layout-aware extractor that reads blocks with coordinates (PyMuPDF get_text("blocks"), pdfplumber with word positions), sort blocks by column, or use a document-AI parser.
  • "How do you handle tables?" — Extract them separately (for example with pdfplumber's table finder) and store each as Markdown or CSV, as its own chunk with the table caption.
  • "Headers and footers repeat on every page — does that matter?" — Yes: "Confidential — Acme Corp — Page 3" pollutes every chunk's embedding. Detect lines that repeat on most pages and strip them.