Building with LLMs

Mini Project: Build an Intelligent Document Q&A App with LangChain


Your company has a 200-page employee handbook. People ask HR the same forty questions about it every week. The obvious idea: put the handbook in front of a model and let people ask.

The obvious implementation pastes the whole thing into the prompt. Price it: 200 pages at roughly 500 words a page is 100,000 words, about 133,000 tokens. At 3 dollars per million input tokens that is 0.40 dollars per question — 400 dollars a day at a thousand questions, to answer things that occupy two paragraphs of the document.

It is also worse at the job than you expect. A single relevant clause competes with 130,000 tokens of irrelevant text, and answers get vaguer as the haystack grows. Add a second document — the expenses policy, the security handbook — and you are past what many models accept in one request at all.

The fix is to stop sending the document and start sending the part of the document that answers this question. Find the three or four passages that matter, put those in the prompt, and answer from them. That is retrieval-augmented generation, and the same request now costs about 0.011 dollars — a 36-fold reduction — while being more accurate, not less, because the model is reading two relevant pages instead of two hundred mostly-irrelevant ones.

You are going to build that: a document question-answering service where you upload PDFs, ask questions in plain English, and get answers with citations back to the source. It is the most common LLM application in industry, and building one end to end exercises everything that makes these systems hard — chunking, embeddings, vector search, prompt construction, an API, a UI, and a container that runs the lot.

The handbook, from PDF to a cited answerExtracttext per pageChunkwith overlapEmbed and storeRetrieve top kfor the questionAnswer, or sayI don't knowChunk size is the quality knob: too small loses the clause, too large drowns it in neighbours.
Everything left of the retrieval step runs once; everything right of it runs per question, and that split is the whole cost model.

What the finished thing does

Hold yourself to these. They separate a demo from something a team will use.

RequirementTargetWhy this number
Upload PDF, DOCX, TXTUp to 50 MBReal handbooks and contracts are large
Answer latencyUnder 3 s at p95Past about 5 s people stop trusting it
Every answer cites its sourcesDocument name and pageAn uncited answer cannot be verified, so it is not trusted
Says "I don't know"When retrieval finds nothing relevantA confident wrong answer is the failure mode that kills these projects
Handles concurrent users10+ simultaneouslyForces you to get the async path right
Cost per questionUnder 0.02 dollarsKeeps 1,000 questions a day under 20 dollars

The architecture, and why it splits where it does

Text
INGESTION (once per document)  PDF ──▶ extract text ──▶ split into chunks ──▶ embed each chunk                                   │                    │                                   ▼                    ▼                            chunk store            vector index                          (id → text, page)     (id → 1536 floats)QUERY (once per question)  question ──▶ embed ──▶ search index ──▶ top-k chunk ids                                              │                                              ▼                                      fetch text from chunk store                                              │                                              ▼                              prompt = question + chunks ──▶ model ──▶ answer + citations

The two paths run at completely different times and rates. Ingestion happens once per document and is slow and expensive. Query happens thousands of times and must be fast and cheap. Everything that can be precomputed belongs in ingestion.

Note the two stores — the piece most first attempts get wrong. A vector index stores vectors and returns identifiers and distances; it does not necessarily hand back the original text. Without a separate chunk store, search returns identifiers you cannot turn into prompt text, and you discover it at the last step, after everything else already works.

Text
document-qa/├── app/│   ├── config.py       # one validated settings object│   ├── documents.py    # extract text from files│   ├── chunking.py     # split text into overlapping chunks│   ├── embeddings.py   # text → vectors, batched│   ├── store.py        # vector index + chunk store together│   ├── qa.py           # retrieve, prompt, answer│   └── api.py          # FastAPI├── ui/app.py           # Streamlit, talks to the API over HTTP└── tests/

Setup

Bash
pip install fastapi uvicorn[standard] streamlit langchain langchain-openai \            chromadb pypdf python-docx tiktoken pydantic-settings \            python-multipart pytest
Python
# app/config.pyfrom pydantic_settings import BaseSettings, SettingsConfigDictfrom pydantic import SecretStr, Fieldclass Settings(BaseSettings):    model_config = SettingsConfigDict(env_file=".env", extra="ignore")    openai_api_key: SecretStr    chat_model: str = "gpt-6-luna"    embedding_model: str = "text-embedding-3-small"    chunk_size: int = Field(default=800, ge=200, le=2000)      # tokens    chunk_overlap: int = Field(default=150, ge=0, le=500)    top_k: int = Field(default=4, ge=1, le=20)    max_upload_mb: int = 50    persist_dir: str = "./storage"settings = Settings()

Commit a .env.example with the key names and no values, put .env in .gitignore, and never let a real key into the repository. A committed key is compromised permanently — deleting it in a later commit does nothing, because the history keeps it and automated scanners find public pushes within minutes.

Extracting text

Python
# app/documents.pyfrom pypdf import PdfReaderfrom docx import Document as Docxfrom pathlib import Pathdef extract(path: str) -> list[tuple[int, str]]:    """Return [(page_number, text)] — page numbers make citations possible."""    suffix = Path(path).suffix.lower()    if suffix == ".pdf":        reader = PdfReader(path)        pages = [(i + 1, (p.extract_text() or "").strip())                 for i, p in enumerate(reader.pages)]        pages = [(n, t) for n, t in pages if t]        if not pages:            raise ValueError(                "No text found. This PDF is probably a scan — it needs OCR first.")        return pages    if suffix == ".docx":        return [(1, "\n".join(p.text for p in Docx(path).paragraphs if p.text.strip()))]    if suffix == ".txt":        return [(1, Path(path).read_text(encoding="utf-8", errors="replace"))]    raise ValueError(f"Unsupported file type: {suffix}")

Returning page numbers rather than one blob of text is what makes "handbook.pdf, page 47" possible in a citation — and retrofitting citations once the page structure is gone means reprocessing everything.

The scanned-PDF check is the most common ingestion failure in practice. A scanned document is images of text; extract_text() returns empty strings, chunking produces nothing, the index stays empty, and the app answers every question with "I don't know" while reporting a successful upload. Detect it and say so.

Chunking: the decision that determines answer quality

Chunking is where most of the quality of a document Q&A system is won or lost, and it gets far less attention than prompt wording.

Python
# app/chunking.pyfrom langchain_text_splitters import RecursiveCharacterTextSplitterfrom app.config import settingssplitter = RecursiveCharacterTextSplitter.from_tiktoken_encoder(    encoding_name="cl100k_base",    chunk_size=settings.chunk_size,    chunk_overlap=settings.chunk_overlap,    separators=["\n\n", "\n", ". ", " ", ""],)def chunk_pages(pages: list[tuple[int, str]], doc_name: str) -> list[dict]:    out = []    for page_no, text in pages:        for i, piece in enumerate(splitter.split_text(text)):            out.append({                "id": f"{doc_name}:{page_no}:{i}",                "text": piece,                "metadata": {"document": doc_name, "page": page_no, "index": i},            })    return out

The separator list is doing real work. RecursiveCharacterTextSplitter tries each separator in turn and only falls through to a cruder one when a chunk is still too big — so it prefers to break at paragraph boundaries, then line breaks, then sentence ends, and only splits mid-word as a last resort. Splitting on a fixed character count instead cuts sentences in half, and half a sentence embeds to a vector that means something different from either the sentence or its neighbours.

Now the size trade-off, which is a genuine engineering choice rather than a default to accept:

Chunk sizeRetrieval precisionContext per chunkFailure mode
200 tokensHigh — very targetedToo littleRetrieves the right sentence without the paragraph that explains it
800 tokensGoodRoughly one sectionThe usual sweet spot for prose documents
2,000 tokensPoor — dilutedPlentyRelevant sentence buried in noise; four chunks fill the prompt

The overlap exists for a specific failure. Suppose a chunk boundary lands mid-way through the notice-period clause: the first chunk ends with "employees must give notice of" and the second begins "three months, except during probation". Neither chunk answers "how much notice must I give". With a 150-token overlap, the second chunk begins 150 tokens earlier and contains the whole clause. Roughly 15–20% of chunk size is a good starting overlap; the cost is storing and embedding about that much duplicate text.

If your Q&A system gives vague answers, look at the chunks it retrieved before you touch the prompt. Nine times out of ten the model answered the best it could from chunks that did not contain the answer.

Embeddings

An embedding turns text into a list of numbers positioned so that texts about similar things sit close together. "How much holiday do I get?" and "Annual leave entitlement is 25 days" share almost no words, yet their vectors are near neighbours — which is why this beats keyword search for questions asked in a user's own words.

Python
# app/embeddings.pyfrom langchain_openai import OpenAIEmbeddingsfrom app.config import settingsembedder = OpenAIEmbeddings(model=settings.embedding_model,                            api_key=settings.openai_api_key.get_secret_value())def embed_documents(texts: list[str], batch_size: int = 100) -> list[list[float]]:    vectors = []    for i in range(0, len(texts), batch_size):        vectors.extend(embedder.embed_documents(texts[i:i + batch_size]))    return vectorsembed_query = embedder.embed_query

Batching matters at ingestion scale. Embedding 4,000 chunks one at a time is 4,000 round trips; at 120 ms each that is eight minutes of pure network latency. In batches of 100 it is 40 requests and under a minute.

The cost is worth internalising, because people over-estimate it badly. The handbook is about 133,000 tokens; at 0.02 dollars per million for a small embedding model, indexing all of it costs 0.0027 dollars — under a third of a penny, and 500 such documents cost about 1.33 dollars. Embeddings are effectively free; generation is where your money goes.

Two rules prevent silent corruption. Use the same model for documents and queries — vectors from different models are not comparable, and mixing them makes retrieval essentially random. And if you change embedding model, re-embed everything; the old vectors are meaningless in the new space and nothing will warn you.

The store: vectors and text together

Python
# app/store.pyimport chromadbfrom chromadb.config import Settings as ChromaSettingsfrom app.config import settingsfrom app.embeddings import embed_documents, embed_queryclient = chromadb.PersistentClient(path=settings.persist_dir,    settings=ChromaSettings(anonymized_telemetry=False))collection = client.get_or_create_collection("documents",    metadata={"hnsw:space": "cosine"})def add_chunks(chunks: list[dict]) -> int:    vectors = embed_documents([c["text"] for c in chunks])    collection.add(        ids=[c["id"] for c in chunks],        embeddings=vectors,        documents=[c["text"] for c in chunks],       # the chunk store        metadatas=[c["metadata"] for c in chunks],    )    return len(chunks)def search(question: str, k: int | None = None, document: str | None = None):    res = collection.query(        query_embeddings=[embed_query(question)],        n_results=k or settings.top_k,        where={"document": document} if document else None,    )    return [        {"text": t, "metadata": m, "distance": d}        for t, m, d in zip(res["documents"][0], res["metadatas"][0], res["distances"][0])    ]

Passing documents= alongside embeddings= is what makes Chroma serve as both stores at once, which is why it is a good choice for a project this size. With a bare vector index you would keep a separate table mapping identifier to text — the "chunk store" from the architecture diagram — and the retrieval step would do two lookups instead of one.

hnsw:space: cosine sets the distance metric. Cosine distance compares direction rather than magnitude, which is what text embeddings want; a euclidean default gives subtly worse neighbours in a way that is very hard to notice from outside.

Answering, with citations and an honest "I don't know"

Python
# app/qa.pyfrom langchain_core.prompts import ChatPromptTemplatefrom langchain_core.output_parsers import StrOutputParserfrom langchain_openai import ChatOpenAIfrom app.store import searchfrom app.config import settingsPROMPT = ChatPromptTemplate.from_messages([    ("system",     "You answer questions using ONLY the provided context.\n"     "Rules:\n"     "- If the context does not contain the answer, reply exactly: "     "  'I could not find that in the uploaded documents.'\n"     "- Never use outside knowledge, even if you are confident.\n"     "- Cite the source of each claim as [document, page N].\n"     "- Be concise. Quote the document's wording where it matters."),    ("human", "Context:\n{context}\n\nQuestion: {question}"),])model = ChatOpenAI(model=settings.chat_model,                   api_key=settings.openai_api_key.get_secret_value())chain = PROMPT | model | StrOutputParser()MAX_DISTANCE = 0.55        # tune against your own documentsdef answer(question: str, document: str | None = None) -> dict:    hits = search(question, document=document)    relevant = [h for h in hits if h["distance"] <= MAX_DISTANCE]    if not relevant:        return {"answer": "I could not find that in the uploaded documents.",                "sources": [], "confidence": "none"}    context = "\n\n---\n\n".join(        f"[{h['metadata']['document']}, page {h['metadata']['page']}]\n{h['text']}"        for h in relevant)    return {        "answer": chain.invoke({"context": context, "question": question}),        "sources": [{"document": h["metadata"]["document"],                     "page": h["metadata"]["page"],                     "distance": round(h["distance"], 3)} for h in relevant],        "confidence": "high" if relevant[0]["distance"] < 0.35 else "medium",    }

Three things in there separate a trustworthy system from a plausible one.

The distance threshold. Vector search always returns k results — it returns the four nearest chunks whether or not any of them is relevant. Ask an employee handbook about the offside rule and you get four chunks about something else, and a model handed irrelevant context will still try to answer from it. Filtering by distance is what turns "always answers" into "answers when it knows". Tune the threshold against your own documents by logging distances for questions you know are answerable and questions you know are not; the gap between the two distributions is where the threshold goes.

Citations embedded in the context. Because each passage is prefixed with its source, the model cites accurately instead of inventing a page number, and anyone can check the answer in ten seconds.

The exact refusal string. Specifying the precise sentence makes refusals detectable in code and countable in metrics. A rising refusal rate means retrieval has degraded — often a scanned PDF that indexed to nothing.

The hard part of document Q&A is not answering — it is declining. Any system will produce an answer for every question; only one with a relevance threshold will tell you when the answer is not in your documents.

The API

Python
# app/api.pyfrom fastapi import FastAPI, UploadFile, File, HTTPExceptionfrom pydantic import BaseModel, Fieldimport tempfile, os, asyncio, logging, timefrom app import documents, chunking, store, qafrom app.config import settingsapp = FastAPI(title="Document Q&A", version="1.0.0")class AskRequest(BaseModel):    question: str = Field(min_length=3, max_length=1000)    document: str | None = None@app.post("/documents")async def upload(file: UploadFile = File(...)):    data = await file.read()    if len(data) > settings.max_upload_mb * 1024 * 1024:        raise HTTPException(413, f"File exceeds {settings.max_upload_mb} MB")    with tempfile.NamedTemporaryFile(            suffix=os.path.splitext(file.filename)[1], delete=False) as tmp:        tmp.write(data); tmp_path = tmp.name    try:        pages = await asyncio.to_thread(documents.extract, tmp_path)        chunks = chunking.chunk_pages(pages, file.filename)        added = await asyncio.to_thread(store.add_chunks, chunks)    except ValueError as exc:        raise HTTPException(422, str(exc))    finally:        os.unlink(tmp_path)    return {"document": file.filename, "pages": len(pages), "chunks": added}@app.post("/ask")async def ask(req: AskRequest):    started = time.perf_counter()    result = await asyncio.to_thread(qa.answer, req.question, req.document)    result["latency_ms"] = int((time.perf_counter() - started) * 1000)    logging.info("ask latency=%dms sources=%d confidence=%s",                 result["latency_ms"], len(result["sources"]), result["confidence"])    return result

asyncio.to_thread is the important detail. PDF parsing and the synchronous Chroma and model calls all block; running them directly inside an async def handler freezes the event loop and stalls every other in-flight request on that worker.

Note also the finally: os.unlink. Without it, every upload leaves a temporary file behind and the container's disk fills up over weeks — a slow failure that is baffling when it finally arrives.

The interface

Python
# ui/app.pyimport streamlit as st, requests, osAPI = os.getenv("API_URL", "http://localhost:8000")st.set_page_config(page_title="Document Q&A", page_icon="📄")st.title("Document Q&A")with st.sidebar:    up = st.file_uploader("Upload a document", type=["pdf", "docx", "txt"])    if up and st.button("Index it"):        with st.spinner("Indexing..."):            r = requests.post(f"{API}/documents",                              files={"file": (up.name, up.getvalue())}, timeout=300)        if r.ok:            d = r.json()            st.success(f"{d['document']}: {d['pages']} pages, {d['chunks']} chunks")        else:            st.error(r.json().get("detail", "Upload failed"))if q := st.chat_input("Ask about your documents"):    st.chat_message("user").markdown(q)    with st.chat_message("assistant"):        with st.spinner("Searching..."):            r = requests.post(f"{API}/ask", json={"question": q}, timeout=60)        if not r.ok:            st.error("The service is unavailable.")        else:            d = r.json()            st.markdown(d["answer"])            if d["sources"]:                with st.expander(f"{len(d['sources'])} sources, {d['latency_ms']} ms"):                    for s in d["sources"]:                        st.caption(f"{s['document']} page {s['page']} "                                   f"(distance {s['distance']})")

The UI holds no model logic and no key — it is an HTTP client, so the same backend can serve a Slack bot or a mobile app later. Exposing the distances in the expander is not decoration either: it is the fastest debugging tool you have when someone reports a bad answer.

Running it

A slim Python base image, a non-root user, PYTHONUNBUFFERED=1 so logs are not swallowed, and uvicorn app.api:app --host 0.0.0.0 as the command. Then compose the two services:

Text
services:  api:    build: .    ports: ["8000:8000"]    environment:      - OPENAI_API_KEY=${OPENAI_API_KEY}    volumes:      - ./storage:/app/storage    restart: unless-stopped  ui:    build: { context: ., dockerfile: Dockerfile.ui }    ports: ["8501:8501"]    environment:      - API_URL=http://api:8000    depends_on: [api]

The volume on ./storage stops every redeploy wiping your index and forcing a full re-ingestion. The key arrives from the environment at run time — never baked into the image, where docker history exposes it to anyone who can pull it.

Testing it honestly

Write down twenty questions with known correct answers before you tune anything. Without that set you will "improve" the prompt by feel and never know whether you helped.

Python
CASES = [    # (question, must appear in answer, expected page)    ("How many days of annual leave do I get?", "25", 12),    ("What is the notice period during probation?", "one week", 8),    ("Can I carry over unused leave?", "five days", 12),    ("What is the offside rule?", "could not find", None),   # must refuse]def test_answers():    for question, expected, page in CASES:        result = qa.answer(question)        assert expected.lower() in result["answer"].lower(), question        if page:            assert page in [s["page"] for s in result["sources"]], question

The refusal case is the most valuable test in the file. A system that answers everything is easy to build and worthless; one that declines when it should is the whole point, and it breaks silently whenever you loosen the distance threshold.

Three failure modes and how to tell them apart, since they look identical from the outside:

SymptomCheckLikely cause and fix
Every answer is a refusalcollection.count()Zero chunks — scanned PDF or failed ingestion. Add OCR or reject the file clearly.
Answers are vague but sources look rightPrint the retrieved chunk textChunks too large, answer diluted. Reduce chunk size, raise k.
Answers cite the wrong pageChunk metadataPage numbers lost during extraction — chunk per page, not per document.
Right chunk exists but is never retrievedSearch a phrase from it directlyQuestion and document use different vocabulary. Try a larger embedding model, or hybrid keyword-plus-vector search.
Latency above 5 sTime retrieval and generation separatelyAlmost always generation — reduce k, shorten chunks, use a faster model.

Making it faster and cheaper once it works

Do these in order, and measure between each one.

Cache question embeddings. The same forty questions recur constantly. Caching removes a 50–100 ms call from all repeat traffic, in about ten lines of code.

Cache whole answers. Generation is not fully deterministic, and many current models do not accept temperature=0 at all, but for a factual question over unchanged documents one good answer is all you need — and a cache makes every repeat identical by construction. Key on the question plus the document set, and give it a TTL so re-uploads invalidate. On FAQ traffic a 30% hit rate is realistic, which at 0.011 dollars a call and 1,000 calls a day saves about 3.30 dollars a day and, more visibly, returns those answers in milliseconds.

Tune k deliberately. Each extra chunk adds roughly 800 tokens to every prompt, so k = 4 to k = 8 doubles your context cost. Check it against your twenty test cases before paying for it; often it changes nothing.

Stream the answer. Generation dominates your latency. Streaming does not reduce total time, but it puts the first words on screen in a few hundred milliseconds instead of three seconds — the difference users actually perceive.

Log four numbers on every question: retrieval latency, generation latency, top distance, and whether it refused. Those turn every future complaint into a diagnosis instead of a guess.

What to build next, and what you now know how to judge

Finish the core version first, then treat these as the real extensions — each one addresses a limitation you will meet the moment other people start using it.

  • Hybrid search. Vector search struggles with exact identifiers — product codes, clause numbers, names. Running a keyword search alongside it and merging the results fixes the single most common "why did it not find that" complaint.
  • Multi-document filtering. Once ten documents are indexed, answers start mixing sources. The where filter is already in the search function; wire it to a selector in the UI.
  • Conversational follow-ups. "What about during probation?" only makes sense given the previous question, so rewrite it into a standalone question before embedding — otherwise you embed a fragment and retrieve nothing.
  • Re-ranking. Retrieve twenty chunks by similarity, then use a cheap model to score which four genuinely answer the question. One extra small call, and it usually improves precision more than any amount of prompt tuning.
  • Access control. Once two teams share the service, chunk metadata must carry who may see it and the search filter must enforce it — far easier before the index has documents in it.

The larger point is what this build teaches you to evaluate. When a vendor demonstrates a document Q&A product, you now know the questions that separate a real one from a demo: what is the chunk size and overlap, does it cite pages, what happens when the answer is not in the corpus, what is the distance threshold, and how does it handle a scanned PDF. Those five questions expose more than any feature list, and you know them because you have had to answer each one yourself.