Course Content
Live Coding Interview Prep
7 sections · 50 lessons
Track metrics like accuracy, latency, and cost in code.
What you need to know
Percentiles. The p95 latency is the value below which 95% of requests fall. LLM latency is long-tailed: most requests take 2 seconds, a few take 20 (long outputs, retries, a slow provider). The mean can look fine while one user in twenty waits 20 seconds. Report p50 (typical), p95 and p99 (the tail).
The nearest-rank method is simple and exact on stored data:
sort the values; p-th percentile = value at position ceil(p/100 × n) (1-based)Cost comes from token usage: input_tokens × input_price + output_tokens × output_price, with prices per million tokens. For example, Claude Opus 5 lists $5 per million input and $25 per million output tokens; prices change, so keep them in configuration.
Cost per successful request is the honest unit. If 10% of calls fail and are retried, cost per call looks lower than what each answered question really costs.
Accuracy needs labels. Online you rarely have them; use a golden set replayed nightly, a judge on sampled traffic, or proxies such as thumbs-down and escalation-to-human rates. Count only graded calls in the denominator.
1import math, time2from collections import defaultdict3from contextlib import contextmanager45PRICING = {"claude-opus-5": (5.00, 25.00), "claude-haiku-4-5": (1.00, 5.00)} # USD per 1M tokens67def cost_usd(model: str, input_tokens: int, output_tokens: int) -> float:8 rate_in, rate_out = PRICING[model] # unknown model -> KeyError, on purpose9 return (input_tokens * rate_in + output_tokens * rate_out) / 1_000_0001011def percentile(sorted_values: list[float], p: float) -> float | None:12 """Nearest-rank percentile of already-sorted values."""13 if not sorted_values:14 return None15 k = max(1, math.ceil(p / 100 * len(sorted_values)))16 return sorted_values[k - 1]1718class Metrics:19 def __init__(self) -> None:20 self.latency = defaultdict(list)21 self.calls, self.ok = defaultdict(int), defaultdict(int)22 self.graded, self.correct = defaultdict(int), defaultdict(int)23 self.usd = defaultdict(float)2425 def record(self, name: str, latency_ms: float, ok: bool = True, model: str | None = None,26 input_tokens: int = 0, output_tokens: int = 0, correct: bool | None = None) -> None:27 self.calls[name] += 128 self.ok[name] += ok29 self.latency[name].append(latency_ms)30 if model:31 self.usd[name] += cost_usd(model, input_tokens, output_tokens)32 if correct is not None:33 self.graded[name] += 134 self.correct[name] += correct3536 def report(self, name: str) -> dict:37 values, n = sorted(self.latency[name]), self.calls[name]38 if n == 0:39 return {}40 return {"n": n, "p50_ms": percentile(values, 50), "p95_ms": percentile(values, 95),41 "p99_ms": percentile(values, 99), "mean_ms": round(sum(values) / n, 1),42 "accuracy": round(self.correct[name] / self.graded[name], 3) if self.graded[name] else None,43 "usd_total": round(self.usd[name], 4),44 "usd_per_success": round(self.usd[name] / self.ok[name], 5) if self.ok[name] else None}4546@contextmanager47def timed(metrics: Metrics, name: str):48 box: dict = {"ok": True}49 started = time.perf_counter()50 try:51 yield box # caller fills model, tokens, correct52 except Exception:53 box["ok"] = False54 raise55 finally:56 metrics.record(name, (time.perf_counter() - started) * 1000, **box)The tricky parts:
max(1, ceil(...))— for p = 0 or tiny lists the rank would be 0, and index-1would silently return the largest value.PRICING[model]raises on an unknown model. A silent cost of zero for a new model is worse than an exception that forces someone to add its price.- Failures still record latency and cost (the
finally), but not success. Failed calls spend money too. self.ok[name] += ok—Trueadds 1,Falseadds 0.
Complexity: record is O(1) amortised. report sorts n latencies, O(n log n). Keeping every latency is O(n) memory; at high volume, use a histogram or a t-digest sketch, which gives approximate percentiles in fixed memory.
A real-life example
Twenty calls to a support bot: most fast, two slow, one failure:
1m = Metrics()2latencies = [900, 950, 1000, 1000, 1100, 1100, 1150, 1200, 1200, 1250,3 1300, 1300, 1350, 1400, 1500, 1600, 1800, 2000, 9000, 15000]4for i, ms in enumerate(latencies):5 m.record("answer", ms, ok=(i != 18), model="claude-opus-5",6 input_tokens=2000, output_tokens=300, correct=(i % 5 != 0))7print(m.report("answer"))8# {'n': 20, 'p50_ms': 1250, 'p95_ms': 9000, 'p99_ms': 15000, 'mean_ms': 2355.0,9# 'accuracy': 0.8, 'usd_total': 0.35, 'usd_per_success': 0.01842}| metric | computation | value |
|---|---|---|
| p50 | rank ceil(0.50 × 20) = 10 → 10th value | 1,250 ms |
| p95 | rank ceil(0.95 × 20) = 19 → 19th value | 9,000 ms |
| mean | 47,100 / 20 | 2,355 ms |
| cost per call | (2,000 × 5 + 300 × 25) / 1,000,000 | $0.0175 |
| cost per success | $0.35 / 19 successes | $0.01842 |
| accuracy | 16 correct of 20 graded | 0.8 |
The mean (2.4 s) looks acceptable; the p95 (9 s) shows that one user in twenty waits nine seconds or more. And the failed call raises the true cost per answered question from $0.0175 to $0.0184.
An online-grocery company tracking its order-help bot puts exactly these numbers on a dashboard per model, and alerts when p95 latency or cost per resolved conversation moves.
Follow-up questions to expect
- "For streaming, which latency matters?" — Both: time to first token is what the user feels; total time is what you pay for and what bounds throughput. Record them separately.
- "How do you count cached responses?" — Count them as requests with zero (or near-zero) marginal cost and very low latency, and report cache hit rate alongside, or the averages mislead.
- "Where do these metrics go?" — Emit them to a metrics system (Prometheus, Datadog, CloudWatch) as histograms and counters, labelled by model, prompt version and route.