AI Product Engineering: Shipping LLM Features That Last

Course Content

AI Product Engineering: Shipping LLM Features That Last

6 sections · 22 lessons

Versioning prompts like code


On a Tuesday afternoon, a support lead fixed a typo in the prompt through an admin page that edited a database row. While there, she changed "reason: at most 40 words, showing the arithmetic" to "reason: short and friendly". By Thursday, unchanged approvals had fallen from 74% to 61%. Reasons had become polite and empty: "Sorry for the trouble, refund issued for affected items."

Nobody could say what had changed, because the row had no history. It took a day to find the edit and another to agree on what the old text had been. Two days of worse drafts, 2,400 tickets, from a change nobody reviewed.

A prompt is code that controls behaviour. It needs what all such code needs: history, review, tests, a way to release gradually, and a fast way back.

The Tuesday edit, two waysPrompt in a database row• Edited live, no review• No history; a day to find the change• Model id changed separately• Approvals fell from 74% to 61%Prompt as a release folder• Pull request with the gate report• Version and hash on every draft• Text, schema andmodel versioned together• Shadow, 10%, 100%, one-flag rollback
The same 'friendlier reasons' request later shipped safely because the gate showed the arithmetic score falling before any agent saw it.

What a release is

A prompt's behaviour depends on more than its text. Change any of these and the output changes:

  • the system prompt text
  • the output schema
  • the model id
  • parameters such as max_tokens
  • few-shot examples, if any
  • the version of the policy it quotes

So the unit you version is a prompt release: all of these together, under one version name, with one content hash. "v3" means exactly one combination. If someone swaps the model and keeps the text, that is v4, not "v3 on a new model".

Keep releases in the repository

Text
prompts/order_issue/  v3/    system.txt        the contract from the first lesson of this section    schema.json       DRAFT_SCHEMA, as JSON    config.yaml       model id, max_tokens, policy version    examples.jsonl    few-shot examples (empty in v3)  v4/    ...  CHANGELOG.md        one entry per version: what changed, why, eval result
YAML
# prompts/order_issue/v3/config.yamlmodel: claude-sonnet-5   # chosen in section 5; any chat model id works heremax_tokens: 4000policy_version: "2026-03"

Releases are never edited once they serve traffic. To change anything, copy the folder to a new version and change it there. That makes a comparison between v3 and v4 a plain diff.

The code that loads a release computes a hash of its contents, so a quiet edit to a live folder is detectable.

Python
import hashlibimport jsonfrom dataclasses import dataclassfrom pathlib import Pathimport yamlROOT = Path("prompts/order_issue")@dataclass(frozen=True)class PromptRelease:    version: str    system: str    schema: dict    model: str    max_tokens: int    content_hash: strdef load_release(version: str) -> PromptRelease:    folder = ROOT / version    system = (folder / "system.txt").read_text(encoding="utf-8")    schema_text = (folder / "schema.json").read_text(encoding="utf-8")    config = yaml.safe_load((folder / "config.yaml").read_text(encoding="utf-8"))    blob = system + schema_text + json.dumps(config, sort_keys=True)    digest = hashlib.sha256(blob.encode("utf-8")).hexdigest()[:12]    return PromptRelease(version, system, json.loads(schema_text),                         config["model"], config["max_tokens"], digest)

load_release reads the three files, hashes them together with the config in a fixed key order, and returns one frozen object. The service loads each release once at start-up and passes it to draft_for_ticket.

Every draft, log line and eval result records release.version and release.content_hash. When approvals drop, the first question, "what changed?", has an answer in one query.

The changelog is short, and it is the first thing a new team member reads:

Text
v4  (2026-04-02)Changed: money arithmetic moved to code; model returns units and rule only.         Added 3 Hinglish examples and 2 veg-order-got-meat examples.Why:     v3 error analysis: 9 price/arithmetic, 4 Hinglish, 2 missed veg cases.Eval:    109/120 (v3: 97/120). Pass-to-fail flips: 1 (ticket 0457, reviewed, accepted).Rollout: shadow 3 days, 10% 2 days, 100% on 2026-04-09.

Each entry says what changed, why, what the eval said, and how it was rolled out. Six months later, when someone asks why the model no longer computes amounts, the answer is one search away.

Review a prompt change like a code change

A prompt change goes through a pull request, like any other change. The review asks three questions:

  1. What problem does it fix? Link the failures from the error analysis (section 4) that motivated it.
  2. What did the eval say? The regression gate's report is attached: overall score, per-category scores, and every case that went from pass to fail.
  3. Is the change the smallest one that fixes it? A rewrite of the whole prompt to fix one rule makes it hard to know what caused any later change.

Support leads should still be able to propose changes. They know the policy better than engineers do. Give them the pull request, not the database row.

Roll out gradually, roll back instantly

  1. Offline eval — the new release runs on the eval set and passes the regression gate.
  2. Shadow mode — for two or three days, v4 runs on live tickets next to v3. Agents see only v3; you compare the two drafts and agent outcomes offline.
  3. Percentage rollout — a flag sends 10% of tickets to v4, then 50%, then 100%, watching unchanged-approval rate and edits at each step.
  4. Rollback — the flag goes back to v3 with one config change, and no deploy is needed.
Python
import randomLIVE = {"stable": "v3", "candidate": "v4", "candidate_share": 0.10}def pick_release(ticket_id: int) -> str:    rng = random.Random(ticket_id)  # same ticket, same release, every time    return LIVE["candidate"] if rng.random() < LIVE["candidate_share"] else LIVE["stable"]

Seeding by ticket id keeps a ticket on one release if it is re-opened, so an agent never sees its draft change under them. In production, LIVE lives in your feature-flag system, not in code.

Shadow mode doubles model cost while it runs: at ₹0.62 a ticket, three days of shadowing 1,200 tickets a day costs about ₹2,200. That is cheap insurance for a change that touches every refund.

Prompt in a database row

  • Edited live, no review
  • No history; "what changed?" takes a day
  • Model id changed separately, untracked
  • Rollback means retyping the old text

Prompt as a release

  • Pull request with eval report attached
  • Every draft stamped with version and hash
  • Model, schema and text versioned together
  • Rollback is one flag change

The right-hand column is more work to set up: perhaps two days for the loader, the flag and the CI job. It pays for itself the first time someone asks "what changed?" and the answer takes one query instead of one day.

Check your understanding

0 of 3 answered

1.The team switches from one model to another but keeps the prompt text identical. How should this be released?

2.Why does pick_release seed the random generator with the ticket id?

3.What is the main purpose of shadow mode?