Building AI Features in Python Backends

Prompts under version control, with tests and evals in CI


Two months after launch, a ShipFast engineer improved the classifier. Complaints were being labelled reschedule too often, so they added a sentence to the prompt: "If the customer expresses frustration, prefer complaint." They checked ten messages by hand, all better, and merged it on a Friday. By Monday, damaged-parcel recall had fallen from 97% to 81%, because customers with broken parcels are, understandably, frustrated. Claims were arriving in the escalations queue, where nobody was watching for them.

Nothing in the process could have caught it. The prompt was a string in a Python file, the change looked like a one-line text edit, the unit tests used a fake model and passed, and nobody measured the real model before or after. A prompt is code that runs on someone else's computer and whose behaviour you can only measure. This lesson gives it the same review and testing as code, plus the one test code does not need: an evaluation.

What runs, and whenPrompt lock: released files never changeContract tests: prompt, enum, routesUnit tests with FakeLLM — every commitReal-model eval — when prompts change
Only the per-label eval catches the Friday change: overall accuracy rose while damaged-parcel recall fell to 81%.

Prompts as versioned files

The prompts move from Python strings to files, one file per released version.

Text
shipfast/prompts/  classify/v3.md  classify/v4.md  extract/v2.md  draft/v2.mdprompts.lock.json
Python
# shipfast/prompts.pyimport hashlibfrom dataclasses import dataclassfrom functools import cachefrom pathlib import PathPROMPT_DIR = Path(__file__).parent / "prompts"@dataclass(frozen=True)class Prompt:    name: str    version: str    text: str    @property    def id(self) -> str:        return f"{self.name}-{self.version}"    @property    def sha256(self) -> str:        return hashlib.sha256(self.text.encode()).hexdigest()@cachedef load_prompt(name: str, version: str) -> Prompt:    text = (PROMPT_DIR / name / f"{version}.md").read_text(encoding="utf-8").strip()    return Prompt(name, version, text)

In classify.py, the constants become PROMPT = load_prompt("classify", "v3"), SYSTEM = PROMPT.text and PROMPT_VERSION = PROMPT.id. Nothing else changes, so every result still records classify-v3.

Why files, if the strings were already in git? Three reasons. A prompt in its own file shows up in code review as a readable text diff, not a change buried in Python. Non-engineers, such as the support lead who knows what "damaged" means to customers, can read and propose changes. And each version gets a stable identity that is recorded with every result.

The rule that makes this work: never edit a released version. To change classify-v3, create v4, point the code at it and let tests and evaluation decide. Old results keep pointing at the exact text that produced them, and rolling back is a one-line change.

Contract tests, on every commit

These tests are fast and free. They check that prompts, schemas and code agree with each other.

Python
# tests/test_prompts.pyimport jsonfrom pathlib import Pathfrom shipfast import classifyfrom shipfast.prompts import load_promptfrom shipfast.routing import ROUTESfrom shipfast.schemas import IntentLOCK = json.loads(Path("prompts.lock.json").read_text())   # {"classify-v3": "9f2c...", ...}def test_released_prompts_are_unchanged():    for prompt_id, digest in LOCK.items():        name, version = prompt_id.rsplit("-", 1)        assert load_prompt(name, version).sha256 == digest, f"{prompt_id} was edited: add a new version"def test_classifier_prompt_defines_every_intent():    for intent in Intent:        assert f"- {intent.value}:" in classify.SYSTEM, f"prompt does not define {intent.value}"def test_every_intent_has_a_queue():    assert set(ROUTES) == set(Intent)

The first test enforces the "never edit a released version" rule mechanically: the lock file stores a hash per released prompt, and any edit to a released file fails CI with a message saying what to do. The second catches the most common drift, adding an intent to the enum but not to the prompt, which would make the model unable to choose it. The third makes sure routing covers every label.

Evaluations, when prompts or models change

Unit and contract tests say nothing about accuracy. For that you call the real model on a labelled set and compare. ShipFast's classification set has 300 real messages, redacted as described in the logging lesson, each labelled by two senior agents, with disagreements settled by a third.

Python
# tests/evals/test_classify_eval.pyimport asyncioimport jsonimport osfrom collections import Counterfrom pathlib import Pathimport pytestfrom shipfast.classify import classifyfrom shipfast.config import settingsfrom shipfast.llm import LLMClientpytestmark = pytest.mark.skipif(not os.getenv("RUN_EVALS"), reason="set RUN_EVALS=1 to call the model")MIN_ACCURACY = 0.91MIN_RECALL = {"damaged_parcel": 0.95, "reschedule": 0.93, "address_change": 0.90}async def label_all(cases: list[dict]) -> list[tuple[str, str]]:    llm, slots = LLMClient(settings.llm_model, timeout_s=30), asyncio.Semaphore(8)    async def one(case: dict) -> tuple[str, str]:        async with slots:            return case["label"], (await classify(llm, case["text"])).intent.value    return await asyncio.gather(*(one(c) for c in cases))def test_classifier_meets_thresholds():    cases = [json.loads(line) for line in Path("evals/classify.jsonl").read_text().splitlines()]    pairs = asyncio.run(label_all(cases))    accuracy = sum(t == p for t, p in pairs) / len(pairs)    totals, hits = Counter(t for t, _ in pairs), Counter(t for t, p in pairs if t == p)    recall = {label: hits[label] / totals[label] for label in totals}    print(json.dumps({"accuracy": round(accuracy, 3), "recall": recall}, indent=2))    assert accuracy >= MIN_ACCURACY    for label, floor in MIN_RECALL.items():        assert recall[label] >= floor, f"{label} recall {recall[label]:.2f} is below {floor}"

Each line of evals/classify.jsonl is one case, such as {"text": "box came wet, charger missing", "label": "damaged_parcel"}. The thresholds are per label, not only overall, because the Friday change would have passed an overall accuracy check: the total barely moved. The damaged-parcel floor of 0.95 is what fails it.

A run costs about 300 × $0.003, roughly $1, and takes about a minute with 8 calls in parallel. That is cheap enough to run on every pull request that touches a prompt or the model setting, and too expensive and slow to run on every commit. The workflow runs it only when those files change:

YAML
# .github/workflows/evals.ymlname: prompt-evalson:  pull_request:    paths: ["shipfast/prompts/**", "shipfast/config.py", "evals/**"]jobs:  evals:    runs-on: ubuntu-latest    steps:      - uses: actions/checkout@v4      - uses: actions/setup-python@v5        with: {python-version: "3.11"}      - run: pip install -r requirements.txt      - run: pytest tests/evals -s        env:          RUN_EVALS: "1"          ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}

Where the eval set comes from

The best source is production: messages agents re-routed, extractions they corrected, drafts they rewrote. Each one is a case the current prompt got wrong, already labelled by a person. ShipFast adds about 20 such cases a week, redacted, and checks every six months that the set still reflects real traffic. The set is a product asset; it is what lets the team change prompts and models with confidence.

Check your understanding

0 of 3 answered

1.A prompt change raises overall accuracy from 92.7% to 93.0% but drops damaged-parcel recall from 97% to 81%. Which check catches it?

2.Why does ShipFast forbid editing a released prompt file?

3.Why run the real-model evaluation only when prompt, model or eval files change, instead of on every commit?