Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your quantized 4-bit model cuts costs massively — but some reasoning tasks degrade unpredictably. How do you evaluate the trade-offs between quantization, latency, and reasoning quality?
What you need to know
Common formats
| Format | Memory vs 16-bit | Typical use |
|---|---|---|
| BF16 / FP16 | 1x | Reference quality |
| FP8 (weights and activations) | about 0.5x | Data-centre GPUs with FP8 support; usually a small quality loss |
| INT8 | about 0.5x | Broad hardware support |
| 4-bit weight-only (AWQ, GPTQ) | about 0.25–0.3x | Large cost savings; quality varies by task |
| GGUF Q4_K_M and similar | about 0.3x | CPUs and laptops via llama.cpp |
How you quantise matters. AWQ and GPTQ use a small calibration dataset to decide how to round; calibrating on your own traffic, not generic web text, often protects the behaviours you care about. Keeping sensitive layers (often embeddings and the output layer) or the KV cache at higher precision can recover part of the loss. Measure each choice; do not assume.
Why errors hit reasoning hardest
A small rounding error in one step barely matters for a one-line answer. In a 20-step calculation or a long chain of thought, errors can compound, and one wrong digit early changes the final answer. Strict formats break the same way: one wrong token makes JSON invalid.
The evaluation plan
- Sample real prompts — from production, stratified by task type and input length.
- Slice by capability — maths and multi-step reasoning, tool-call validity, retrieval faithfulness, long-context recall, safety refusals, each language you serve.
- Pair the runs — same prompts, same settings, both models; compare per item, which is far more sensitive than comparing two averages.
- Sample k times — for example 5 runs per prompt, and report the spread as well as the mean.
- Plot the curve — per-slice quality against cost per million tokens and p95 latency.
- Route — quantised by default; higher precision for requests classified as reasoning-heavy or when confidence is low.
1import statistics as st23def compare(prompts, ref, quant, grade, k=5):4 report = {}5 for slice_name, items in group_by_slice(prompts).items():6 deltas, spreads = [], []7 for p in items:8 r = [grade(p, ref(p)) for _ in range(k)]9 q = [grade(p, quant(p)) for _ in range(k)]10 deltas.append(st.mean(q) - st.mean(r))11 spreads.append(st.pstdev(q) - st.pstdev(r))12 report[slice_name] = {"mean_delta": st.mean(deltas), "extra_spread": st.mean(spreads)}13 return reportgrade scores one answer (exact match, JSON validity, a judge score). The report shows, per slice, how much worse the quantised model is on average and how much less consistent it is.
A real-life example
Scenario, numbers made up. A fintech's support model is moved from BF16 to 4-bit AWQ, cutting GPU cost by more than half. The overall eval score falls by only 1 point, and the change ships. Two weeks later, agents report wrong EMI calculations.
A paired eval on 2,000 production prompts shows the picture. General FAQ answers: no change. Multi-step EMI and interest calculations: 14 points worse, with the spread across 5 samples tripling. Tool-call JSON validity: 99.6% down to 97.1%. The team re-calibrates AWQ on their own traffic, which halves the maths gap, and routes calculation and tool-calling requests (about 15% of traffic) to a BF16 replica. They keep most of the savings and the complaints stop.
Follow-up questions to expect
- "Why does variance matter so much?" — Users see single answers. A model that is right on average but wrong one time in four on the same question feels unreliable, which is exactly "degrades unpredictably".
- "How do you route to the right model?" — A cheap classifier or rules on the request (calculation, tool use, long context) choose the path; low-confidence answers can also be retried on the full model.
- "Would a smaller unquantised model be better?" — Sometimes. Compare a quantised large model with a full-precision smaller one on the same curve; the answer depends on your slices.