Live Coding Interview Prep

Course Content

Live Coding Interview Prep

7 sections · 50 lessons

Track metrics like accuracy, latency, and cost in code.


What you need to know

Percentiles. The p95 latency is the value below which 95% of requests fall. LLM latency is long-tailed: most requests take 2 seconds, a few take 20 (long outputs, retries, a slow provider). The mean can look fine while one user in twenty waits 20 seconds. Report p50 (typical), p95 and p99 (the tail).

The nearest-rank method is simple and exact on stored data:

Text
sort the values; p-th percentile = value at position ceil(p/100 × n)   (1-based)

Cost comes from token usage: input_tokens × input_price + output_tokens × output_price, with prices per million tokens. For example, Claude Opus 5 lists $5 per million input and $25 per million output tokens; prices change, so keep them in configuration.

Cost per successful request is the honest unit. If 10% of calls fail and are retried, cost per call looks lower than what each answered question really costs.

Accuracy needs labels. Online you rarely have them; use a golden set replayed nightly, a judge on sampled traffic, or proxies such as thumbs-down and escalation-to-human rates. Count only graded calls in the denominator.

Python
import math, timefrom collections import defaultdictfrom contextlib import contextmanagerPRICING = {"claude-opus-5": (5.00, 25.00), "claude-haiku-4-5": (1.00, 5.00)}   # USD per 1M tokensdef cost_usd(model: str, input_tokens: int, output_tokens: int) -> float:    rate_in, rate_out = PRICING[model]                   # unknown model -> KeyError, on purpose    return (input_tokens * rate_in + output_tokens * rate_out) / 1_000_000def percentile(sorted_values: list[float], p: float) -> float | None:    """Nearest-rank percentile of already-sorted values."""    if not sorted_values:        return None    k = max(1, math.ceil(p / 100 * len(sorted_values)))    return sorted_values[k - 1]class Metrics:    def __init__(self) -> None:        self.latency = defaultdict(list)        self.calls, self.ok = defaultdict(int), defaultdict(int)        self.graded, self.correct = defaultdict(int), defaultdict(int)        self.usd = defaultdict(float)    def record(self, name: str, latency_ms: float, ok: bool = True, model: str | None = None,               input_tokens: int = 0, output_tokens: int = 0, correct: bool | None = None) -> None:        self.calls[name] += 1        self.ok[name] += ok        self.latency[name].append(latency_ms)        if model:            self.usd[name] += cost_usd(model, input_tokens, output_tokens)        if correct is not None:            self.graded[name] += 1            self.correct[name] += correct    def report(self, name: str) -> dict:        values, n = sorted(self.latency[name]), self.calls[name]        if n == 0:            return {}        return {"n": n, "p50_ms": percentile(values, 50), "p95_ms": percentile(values, 95),                "p99_ms": percentile(values, 99), "mean_ms": round(sum(values) / n, 1),                "accuracy": round(self.correct[name] / self.graded[name], 3) if self.graded[name] else None,                "usd_total": round(self.usd[name], 4),                "usd_per_success": round(self.usd[name] / self.ok[name], 5) if self.ok[name] else None}@contextmanagerdef timed(metrics: Metrics, name: str):    box: dict = {"ok": True}    started = time.perf_counter()    try:        yield box                                # caller fills model, tokens, correct    except Exception:        box["ok"] = False        raise    finally:        metrics.record(name, (time.perf_counter() - started) * 1000, **box)

The tricky parts:

  • max(1, ceil(...)) — for p = 0 or tiny lists the rank would be 0, and index -1 would silently return the largest value.
  • PRICING[model] raises on an unknown model. A silent cost of zero for a new model is worse than an exception that forces someone to add its price.
  • Failures still record latency and cost (the finally), but not success. Failed calls spend money too.
  • self.ok[name] += ok — True adds 1, False adds 0.

Complexity: record is O(1) amortised. report sorts n latencies, O(n log n). Keeping every latency is O(n) memory; at high volume, use a histogram or a t-digest sketch, which gives approximate percentiles in fixed memory.

A real-life example

Twenty calls to a support bot: most fast, two slow, one failure:

Python
m = Metrics()latencies = [900, 950, 1000, 1000, 1100, 1100, 1150, 1200, 1200, 1250,             1300, 1300, 1350, 1400, 1500, 1600, 1800, 2000, 9000, 15000]for i, ms in enumerate(latencies):    m.record("answer", ms, ok=(i != 18), model="claude-opus-5",             input_tokens=2000, output_tokens=300, correct=(i % 5 != 0))print(m.report("answer"))# {'n': 20, 'p50_ms': 1250, 'p95_ms': 9000, 'p99_ms': 15000, 'mean_ms': 2355.0,#  'accuracy': 0.8, 'usd_total': 0.35, 'usd_per_success': 0.01842}
metriccomputationvalue
p50rank ceil(0.50 × 20) = 10 → 10th value1,250 ms
p95rank ceil(0.95 × 20) = 19 → 19th value9,000 ms
mean47,100 / 202,355 ms
cost per call(2,000 × 5 + 300 × 25) / 1,000,000$0.0175
cost per success$0.35 / 19 successes$0.01842
accuracy16 correct of 20 graded0.8

The mean (2.4 s) looks acceptable; the p95 (9 s) shows that one user in twenty waits nine seconds or more. And the failed call raises the true cost per answered question from $0.0175 to $0.0184.

An online-grocery company tracking its order-help bot puts exactly these numbers on a dashboard per model, and alerts when p95 latency or cost per resolved conversation moves.

Follow-up questions to expect

  • "For streaming, which latency matters?" — Both: time to first token is what the user feels; total time is what you pay for and what bounds throughput. Record them separately.
  • "How do you count cached responses?" — Count them as requests with zero (or near-zero) marginal cost and very low latency, and report cache hit rate alongside, or the averages mislead.
  • "Where do these metrics go?" — Emit them to a metrics system (Prometheus, Datadog, CloudWatch) as histograms and counters, labelled by model, prompt version and route.