Course Content
LLM Evaluation
6 sections · 50 lessons
How do you evaluate multi-turn conversations effectively?
What you need to know
Many assistant failures only appear across turns:
- Context loss — the user said "I'm on probation" in turn 2; the answer in turn 6 ignores it.
- Contradiction — turn 3 says 12 days of leave, turn 7 says 15.
- Re-asking — asking for the employee ID the user already gave.
- Topic-switch confusion — after a side question, the assistant loses the main task.
Method 1: replay
Store real multi-turn transcripts and send the recorded user messages to your system one by one. Cheap and deterministic, but once your new system answers differently in turn 2, the recorded turn 3 may not make sense any more.
Method 2: simulated users
Simulation scales to hundreds of conversations and reacts to what your system says. Its weakness: simulated users are more patient, clearer and more cooperative than real people. Give some personas hard traits — vague, impatient, switching topics — and keep checking a sample against real transcripts. LangSmith (through its openevals package) and DeepEval both provide simulation helpers.
Scoring
- Conversation-level rubric over the full transcript: goal reached? consistent? constraints kept? clarifying questions when needed and only then?
- Turn-level checks for known rules, such as "never ask for data already given".
- Efficiency: turns needed to reach the goal.
LLM judges lose accuracy on long transcripts, so score specific yes/no questions, not "rate the conversation", and have humans review a sample.
A real-life example
The HR policy assistant is tested with 120 simulated employees. One persona: "You want to know if you can encash unused leave. You joined 4 months ago and are on probation — reveal this only if asked. You get annoyed if asked the same thing twice."
Single-turn evals had shown 90% correct answers. The simulation found that in 22% of conversations the assistant gave the full encashment rule without asking about employment status, although the policy differs during probation. Where the user mentioned probation early, 9% of conversations still dropped it by turn 5. The team added a rule to ask one clarifying question when the policy depends on status, and a check that the final answer uses facts stated earlier. Correct final answers rose to 84% of simulated conversations.
Follow-up questions to expect
- "How do you stop the simulated user from being too easy?" — Write varied personas with hidden facts, vague wording and impatience, and compare simulated transcripts with real ones on turn count and question types.
- "How do you get repeatability?" — Fix the persona set and seeds, pin the simulator model, and run several conversations per persona; report averages with intervals.
- "What single metric would you track?" — Goal completion rate, plus a consistency check, sliced by persona type.