Course Content
AI Safety & Guardrails
5 sections · 50 lessons
What is knowledge editing, and how can it fix model behavior?
What you need to know
The method families
| Family | Examples | Idea |
|---|---|---|
| Locate-and-edit | ROME, MEMIT | Treat MLP layers as key–value memories; find where the fact is stored and apply a targeted weight update. MEMIT edits many facts at once. |
| Meta-learning / hypernetwork | MEND | Train a small network that turns a fine-tuning gradient into a precise edit. |
| Memory-based | SERAC and similar | Keep edits in an external memory; a classifier decides whether a query touches an edit and routes it to a patch model. |
| Adapter | LoRA with corrections | Fine-tune a small adapter on corrected examples; removable. |
How editing is evaluated
- Reliability: the edited question gives the new answer.
- Generalisation: paraphrases give the new answer too.
- Locality: unrelated facts are unchanged.
- Portability / ripple effects: facts that depend on the edit update correctly ("Who does the new CEO report to?").
Why it is rarely the right production fix
- Edits often pass the exact phrasing and fail the paraphrase.
- Related facts break or stay inconsistent.
- Hundreds of edits degrade the model over time.
- You cannot review an edit the way you review a config or document change.
Better options for most teams
- Retrieval: update the document; the next answer changes. Versioned and easy to revert.
- Rules layer: for a few exact facts that must never be wrong, return a fixed answer in code.
- Prompt: short-term fixes in the system prompt, with evals.
- Fine-tuning for behaviour or style changes across many cases.
A real-life example
A bank's customer chatbot uses a fine-tuned open model. After an interest-rate change, it keeps saying the savings account rate is 3.5% instead of 3.0%. An engineer applies a ROME edit in an afternoon. The direct question now answers 3.0%, but "What rate will I earn on my savings?" still says 3.5%, and one test shows the FD rate answer also changed wrongly.
The team reverts the edit. Rates now come from a get_rates() tool that reads the rates table, and the model is instructed to always call it for rate questions. The next rate change needs no model work at all, and an eval checks 30 phrasings of rate questions after every deploy.
Follow-up questions to expect
- "When would you use knowledge editing?" — Research, safety studies, or an urgent fix on a model you cannot retrain, only with thorough regression tests.
- "Can editing remove harmful knowledge?" — Related methods ("unlearning") try, but information often remains recoverable with different prompts or light fine-tuning. Treat it as reduction, not removal.
- "How is it different from fine-tuning?" — Fine-tuning updates many parameters on many examples; editing targets one fact with a small, local change.