AI Product Engineering: Shipping LLM Features That Last

Course Content

AI Product Engineering: Shipping LLM Features That Last

6 sections · 22 lessons

Regression gates: never ship a prompt that breaks old cases


The first candidate for v4 scored 110 of 120, one better than the version that eventually shipped. It would have gone out that afternoon if anyone had looked only at the headline number. But the gate listed two tickets that had passed on v3 and failed now. Both were food-safety tickets. The new Hinglish examples had taught the model that "kuch ajeeb sa tha" (something was strange) was a cold-food complaint.

Better on average, worse where it matters most. This is the normal shape of a prompt change: every edit that fixes one group of failures moves the model's behaviour on everything else a little. Most of those movements are harmless. A few are not.

A regression gate is an automatic check that runs every candidate release against the eval set and blocks it if it breaks the things you have decided must never break. It is how the eval set turns from a report into a safety mechanism.

The candidate that scored higher and was blockedWhat the average said• 110 of 120, one better than v4• Hinglish examples helped• Looked like a clear win• Would have shipped that afternoonWhat the gate listed• 2 food-safety rows: pass to fail• 'kuch ajeeb sa tha' read as cold food• Critical rows must all pass• Blocked; fixed; then shipped
An average can rise while the one category that must never fail gets worse, so critical rows are gated one by one, not by percentage.

What the gate checks

TiffinGo's gate has five conditions. A release must meet all of them.

  1. Overall score. The pass rate must not fall more than 2 points below the current production release.
  2. Every category meets its own threshold. Food safety, the critical category, must pass completely, vegetarian-given-meat rows included; with 8 rows, a percentage would be meaningless. Other categories have looser floors, explained below.
  3. Flips are listed and signed off. Every row that passed before and fails now is listed by id. The gate fails until a person reviews each flip and records the decision.
  4. Validation rejections stay low. At most 2% of drafts may fail the validator; a jump means the contract confuses the model.
  5. Cost and latency stay in budget. Average cost per ticket and p95 latency may not rise by more than 20% without a separate decision.

The overall threshold catches broad decay, the category thresholds catch damage an average hides, and the flip report catches the rest, because a person reads each broken case by name.

Take the first v4 candidate. It scored 110 against v3's 97, so condition 1 passed. Two food-safety rows failed every run, so condition 2 failed, and both also appeared as flips. One failed condition blocks. The team reworded the Hinglish examples, lost one routine ticket doing it, and shipped at 109: a point of average traded for food safety.

A pass does not mean the release is good, only that it is not worse than production in any way the team has written down. It then goes to shadow mode, which sees what the eval set cannot.

Randomness: run each case more than once

The same prompt on the same ticket does not always give the same draft: about 5 of TiffinGo's 120 cases change between runs, and run once, they would raise false alarms. So the runner runs each case three times and records a pass rate: 0, 0.33, 0.67 or 1. A flip counts only when a case goes from always passing to always failing; anything in between is reported as unstable, a case where the prompt leaves the model genuinely unsure.

Three runs cost 120 × 3 × ₹0.62, about ₹225, plus about ₹40 for the judge. Next to one bad release touching 1,200 refunds a day, that is very cheap.

In food safety an unstable case gets no benefit of the doubt: passing 2 runs of 3 suggests one real ticket in three like it would be missed, so it blocks. Elsewhere it counts only its lost fraction against its category's floor, described next.

Never fix a flaky case by re-running the gate until it goes green; that passes on luck and hides the ticket the model finds hard. And do not lower the temperature for the eval alone: that measures a system customers never see. Instead, TiffinGo reads every case unstable across two releases. Usually the ticket is ambiguous, like "food was not good, do something", and needs a policy or label fix; sometimes the contract lacks an example.

Choosing a threshold for each category

A threshold comes from two numbers: what a miss costs, from the taxonomy's cost column, and how much the score moves by chance, which you measure. Before switching the gate on, TiffinGo ran v3 against itself five times, for about ₹1,100. The overall score never moved more than 1.2 points, and no category more than one row. Any tighter threshold fails releases on noise alone.

CategoryRowsRuleWhy
Food safety8Every row passes in all three runsHarm to a customer; too few rows for a rate
High value (R6), repeat refunds (R7)14May not score below the baselineLarge refunds, or rewards gaming
Missing, cold or packed, wrong item, late90May lose at most 2 rows₹5–170 a miss; agents catch about 85%
Needs info or unclear8Reported, never blocksA few minutes of senior time

Two rows is the smallest loss clearly above the one-row noise, and the 2-point overall limit sits just above the measured 1.2. Rows are counted from pass rates, so a case sliding from 1.0 to 0.33 costs two-thirds of a row. Floors are relative to the baseline, not fixed targets like "cold above 90%", because a release must answer one question: is this worse than what customers get today? Re-measure the noise when a category grows.

The gate in code, and in CI

Python
import jsonimport sysMAX_DROP = 0.02CRITICAL = {"food_safety"}                       # every row, every runFLOORS = {"high_value": 0, "repeat_refund": 0,   # rows a category may lose          "missing_item": 2, "cold": 2, "wrong_item": 2, "late": 2}def load(path: str) -> dict:    with open(path, encoding="utf-8") as f:        return {row["id"]: row for row in map(json.loads, f)}def mean_pass(rows: dict) -> float:    return sum(r["pass_rate"] for r in rows.values()) / len(rows)def rows_passed(rows: dict, category: str) -> float:    return sum(r["pass_rate"] for r in rows.values() if r["category"] == category)def gate(baseline_path: str, candidate_path: str, accepted: set[str]) -> list[str]:    base, cand = load(baseline_path), load(candidate_path)    problems = []    if mean_pass(cand) < mean_pass(base) - MAX_DROP:        problems.append(f"overall {mean_pass(cand):.1%} vs baseline {mean_pass(base):.1%}")    for category, allowed in FLOORS.items():        lost = round(rows_passed(base, category) - rows_passed(cand, category), 2)        if lost > allowed:            problems.append(f"category {category}: lost {lost} rows, {allowed} allowed")    for case_id, row in sorted(cand.items()):        if row["category"] in CRITICAL and row["pass_rate"] < 1.0:            problems.append(f"critical {case_id} ({row['category']}): pass rate {row['pass_rate']:.2f}")        old = base.get(case_id)        if old and old["pass_rate"] == 1.0 and row["pass_rate"] == 0.0 and case_id not in accepted:            problems.append(f"flip {case_id}: always passed before, always fails now")    return problemsif __name__ == "__main__":    accepted = set(sys.argv[3].split(",")) if len(sys.argv) > 3 else set()    found = gate(sys.argv[1], sys.argv[2], accepted)    print("\n".join(found) if found else "gate passed")    sys.exit(1 if found else 0)

Each result file has one JSON row per case with its id, category and pass_rate; the baseline is the stored result of the release in production. FLOORS is the threshold table as data, and round stops float noise in the sums from counting as a loss. The third argument lists flips a reviewer has accepted. Checks 4 and 5 follow the same pattern over the run's summary numbers.

The gate runs automatically on any pull request that touches the prompts or the drafting code.

YAML
name: prompt-evalon:  pull_request:    paths: ["prompts/**", "service/drafting/**"]jobs:  eval:    runs-on: ubuntu-latest    steps:      - uses: actions/checkout@v4      - uses: actions/setup-python@v5        with:          python-version: "3.12"      - run: pip install -r requirements.txt      - run: python -m evals.run --release candidate --runs 3 --out candidate.jsonl        env:          ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}      - run: python -m evals.gate evals/baseline.jsonl candidate.jsonl "$(cat evals/accepted_flips.txt)"

The report, with per-category scores, is posted on the pull request. Model changes take the same path, because the model id lives in the release's config.yaml.

Changing the baseline, and overriding the gate

The baseline defines what "worse" means, so it changes only by decision.

  • It moves when a release reaches 100%, not when its pull request merges, because the rollout may still be reversed. When v4 reached 100% on 9 April, a small pull request replaced baseline.jsonl with v4's results and emptied accepted_flips.txt; ticket 0457, the accepted flip, was now part of the baseline.
  • It is re-scored when the eval set changes. A new row has no baseline entry, so it can never show as a flip. Each Monday, new rows and fixed labels are scored on the production release, about ₹20 for ten rows.
  • It is never regenerated to make a gate pass. Baseline changes get their own pull request with a one-line reason, and the diff shows every case that moved.

Overrides need a rule too. A gate that can never be overridden gets switched off the first time it blocks an urgent fix; one overridden casually is worthless. A reviewer may accept a flip when the new answer is right or equally acceptable, for instance when the old label was wrong, and then fixes the label in the eval set. Accepting a flip because "it is only one ticket" needs the support lead's sign-off and a changelog entry. Critical-category failures are never overridden; the release is fixed instead.

Check your understanding

0 of 3 answered

1.A candidate improves the overall score from 97 to 110 but fails two food-safety rows that passed before. What should happen?

2.A case passes in 2 of 3 runs on the baseline and 1 of 3 on the candidate. How should the gate treat it?

3.Why does the gate also fail when average cost per ticket rises by more than 20%?