Course Content
LLMOps & Deployment
6 sections · 40 lessons
How do you manage prompt versions in production systems?
What you need to know
Why prompts need versioning
A prompt is code that the model executes. A one-word change can alter tone, length, cost, language or safety. If prompts are edited in place, you cannot answer the first question of every incident: what changed?
The working setup
- Store templates, not strings. A template has declared variables (
{language},{documents}) and is rendered at request time. - Immutable IDs. Once a version is published, it never changes. A new edit is a new version. A content hash makes accidental edits visible.
- Attach evals to the version. The eval run for
pension-faq@7.0.0is stored with it. - Log it on every trace.
prompt_version,model,temperature,index_version. - Release by flag. The flag decides which version each user gets; the old version stays loaded for instant rollback.
- Pin the model. Use dated model snapshots, not floating aliases like "latest", so a provider update does not silently change your prompt's behaviour.
1import hashlib2from dataclasses import dataclass34@dataclass(frozen=True)5class PromptVersion:6 name: str7 version: str8 template: str910 @property11 def id(self) -> str:12 digest = hashlib.sha256(self.template.encode()).hexdigest()[:8]13 return f"{self.name}@{self.version}+{digest}"1415pension_v7 = PromptVersion(16 name="pension-faq",17 version="7.0.0",18 template="You answer questions about pension schemes in {language}. "19 "Use only the documents below.\n{documents}\nQuestion: {question}",20)21print(pension_v7.id) # pension-faq@7.0.0+ddfb9348frozen=True stops code from changing a version after it is created. The hash in the ID changes if anyone edits the text, even if they forget to bump the version number.
Git or a registry?
- Git + flags is enough when engineers own the prompts. You get review and history for free.
- A registry (Langfuse, LangSmith and similar tools have prompt management) helps when product or content people edit prompts, because they get a UI, version labels such as "production", and links from each version to its traces.
Either way, the same rules hold: immutable versions, review, eval, logged IDs.
A real-life example
A government pension chatbot kept its prompt in a database row that a content officer could edit. One Friday she added "Always mention the new online portal." On Monday, complaints arrived: for elderly users asking about paper forms, the bot now pushed the portal and skipped the paper process.
Nobody could tell which answers used which text, because the prompt had no version. The team moved to a registry: each edit creates a new version, runs the 250-case eval set (including 40 paper-form cases) and goes live only through a 10% flag. The same edit tried again scored 71% on paper-form cases against 92% before, and was fixed before any citizen saw it.
Follow-up questions to expect
- "How do you version prompt, model and index together?" — Define a release bundle (prompt version, model snapshot, parameters, index version) with one ID, and flag that bundle.
- "How fast can you roll back?" — Seconds: switch the flag back to the previous bundle, which is still deployed.
- "Do you version few-shot examples separately?" — They are part of the prompt template or a versioned file it references; either way, any change produces a new version.