Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Your boss wants fine-tuning on 100 labeled examples. Early results look great but production accuracy is terrible. When is fine-tuning the right answer — and when is it actively harmful?


Claim-email accuracy at each stage956178890123leaked testproductionprompt, 24examplesLoRA on3,000 real14 of the 20 test emails had a near-twin in the training set.
The 95 percent measured one analyst's writing style; the frozen production eval set is what made every later number believable.

What you need to know

Why the early results lied

"Early results look great" means the test set looked like the training set. With 100 examples, the usual cause is that the test cases were split from the same 100, or written by the same person on the same day in the same style. The model learned that person's phrasing. Production users write differently, so accuracy collapses.

A quick check for near-duplicates between train and test:

Python
import numpy as npdef leaked_pairs(train_vecs, test_vecs, threshold=0.92):    """Both are L2-normalised embedding arrays; returns test rows with a near-twin in train."""    sims = test_vecs @ train_vecs.T    best = sims.max(axis=1)    return np.where(best >= threshold)[0], float((best >= threshold).mean())# idx, share = leaked_pairs(embed(train_texts), embed(test_texts))# print(f"{share:.0%} of test cases have a near-duplicate in training")

What fine-tuning is good at

Fine-tuning changes the model's weights with your examples. It is good at behaviour: always returning a given JSON shape, a house tone, a fixed set of labels, or copying a larger model's behaviour into a smaller one (distillation). It is weak at knowledge: a few examples of a fact do not reliably teach the fact, and you cannot update or cite it later.

NeedBest first toolFine-tune when
Answers from company documentsRAG (retrieval)Rarely; maybe to teach citation style
Fixed output format or tonePrompt plus few-shot examplesPrompting is not consistent enough at your volume
Narrow classification or extractionPrompt, then fine-tuneYou have low thousands of varied labelled rows
Lower cost or latencySmaller model with a good promptDistil a large model into a small one with many examples

Rough floors: a few hundred varied examples for style, low thousands for a task where accuracy matters.

When fine-tuning actively harms

  • Facts go stale. A policy change means a new training run instead of editing one document.
  • Forgetting. Narrow training can weaken general abilities the product also needs (called catastrophic forgetting).
  • Confident errors. The model learns the house voice and says wrong things in it.
  • Upkeep. Every base-model upgrade means retraining and re-testing.

The order I'd work in

  1. Freeze an eval set — 200 real production queries with human-labelled answers, from different users and weeks.
  2. Improve the prompt — clear instructions, the output schema, and a few examples from the 100.
  3. Add retrieval — if answers depend on documents, RAG over the sources.
  4. Fine-tune last — only if the remaining gap is format or behaviour, with LoRA on a larger, deliberately varied dataset, measured on the frozen set.

A real-life example

Scenario, numbers made up. An insurance company wants to classify claim emails into 12 categories. An analyst writes 100 examples, fine-tunes, and tests on 20 of them: 95%. In production, accuracy is 61%.

The leak check shows 14 of the 20 test emails have a near-duplicate in training. The team builds a 200-email eval set from real traffic across three months. A careful prompt with the 12 category definitions and 24 examples reaches 78%. They then label 3,000 real emails, fine-tune with LoRA, and reach 89% on the frozen set, with a clear gain on the four categories the prompt kept confusing.

Follow-up questions to expect

  • "Is 100 examples ever enough?" — For a light style change on a strong model, sometimes. For accuracy on a real task, it rarely covers the variety users produce, and you cannot even measure it well with so few.
  • "LoRA or full fine-tuning?" — LoRA trains small adapter matrices, so it is cheaper, easier to roll back, and forgets less. Full fine-tuning is for large datasets and big behaviour changes.
  • "How do you convince the boss?" — Show the leak check and the score on real production queries. Numbers on the right eval set end the debate.