Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

An LLM extraction pipeline processing millions of documents costs far more than budgeted. How do you cut the cost by about 10x?


Monthly bill in thousands of dollars, 4M documents602211601233,200 exampletokens every callcacheprefix, batch APIcascade, 18% escalatedistilled 8B on vLLM
Caching and batching are hours of work; the last big step comes from moving the examples into a small model's weights.

What you need to know

You cannot cut what you have not measured. The first job is to see where the tokens go.

Text
Per request (illustrative):  system prompt + schema      800 tokens   (same every call)  few-shot examples         3,000 tokens   (same every call)  the document              1,500 tokens   (unique)  output                      250 tokens

Here 3,800 of 5,300 input tokens are identical on every call. That static prefix is the first target.

The four levers, in order of effort

LeverWhat it doesTypical effortCatch
Prompt cachingThe provider reuses a repeated prefix; cached input is billed at a steep discountHoursPrefix must be byte-identical and placed first
Batch APISubmit jobs for asynchronous processing at about half priceA dayResults arrive within hours, not seconds
CascadeA small model does the work; low-confidence cases go to the big modelA week or twoNeeds a reliable confidence signal
DistillationFine-tune a small model on the big model's labelled outputsWeeksNeeds a training set, an eval and retraining when inputs drift

The major providers, including OpenAI, Anthropic and Google, offer batch interfaces at roughly a 50% discount for work that can wait. A nightly extraction job has no reason to pay real-time prices.

Why distillation is where 10x comes from

Caching and batching together might cut the bill by half or more. A cascade might send 75 to 85% of documents to a small model. But a small open model fine-tuned with LoRA on 10,000 or more good outputs can handle most documents on its own. The few-shot examples move into the weights, so the prompt shrinks too. Self-hosted on a GPU that stays busy, the per-document cost can drop by an order of magnitude.

  1. Measure — log tokens by prompt section, cost per document, and per tenant.
  2. Cache and batch — reorder the prompt, move offline jobs to the batch API.
  3. Cascade — small model first, escalate when its confidence or a verifier check fails; watch the escalation rate.
  4. Distil — train on accumulated high-quality outputs, evaluate field by field against the big model, then switch.

Guardrails so it stays fixed

Add a per-tenant token budget and a daily spend alert. A new customer uploading 5 million pages, or a code change that doubles the prompt, should be caught in hours, not on the monthly invoice.

A real-life example

Scenario (illustrative numbers). A logistics company extracts fields from 4 million invoices and delivery receipts a month. The bill is about $60,000 a month on a frontier model, far over the $8,000 budget.

The breakdown shows 3,200 tokens of few-shot examples on every call. Moving them into a cached prefix and running the nightly jobs through the batch API brings the bill to about $22,000. A cascade then sends documents to a small hosted model first and escalates the 18% where a verifier finds a missing or unsupported field; that reaches about $11,000.

After two months they have 400,000 verified outputs. A LoRA fine-tune of an 8B open model matches the frontier model's field accuracy within one point on a 2,000-document test set. Served on two GPUs with vLLM, the total cost is about $6,000 a month, roughly 10x below the start.

Follow-up questions to expect

  • "How do you know the small model's confidence is trustworthy?" — Don't rely on self-reported confidence alone. Use a verifier, such as checking each value appears in the source, and measure the error rate on documents it did not escalate.
  • "Why not distil on day one?" — You need thousands of trusted labels and an eval to prove the student works. The big model generates both while the cheaper levers save money.
  • "What about output tokens?" — Tighten the schema, drop explanations, and set max_tokens; output tokens usually cost several times more than input tokens.