Behavioral Interview

Course Content

Behavioral Interview

15 sections · 30 lessons

Deep dive: what interviewers score and how to choose the story


The deep dive is the most technical competency in the round and the one most likely to be run by a senior engineer who will keep asking until they hit the edge of what you know. That is the point of it.

Because the interviewer controls the depth, your preparation is mostly about choosing the right story and knowing it all the way down. This lesson covers what the deep dive is testing, the signals the interviewer writes down, the phrasings that let you choose your own story, and the tests and sketch that make a story ready.

Drilling until it reaches the edgeYour story, as toldWhy this design, not that?What broke, exactlyHow did you know?The edge of what you know
Reaching the edge is the intended outcome; the score comes from whether you say so plainly when you get there.

The definition

The mistake almost everyone makes is optimising for impressiveness. Candidates pick the biggest, most complex thing they have touched, because it sounds senior. Then the drilling starts and it turns out they were adjacent to the complexity rather than inside it.

What it is not

It is not "the hardest problem you've solved". Even when that is the literal question. The interviewer will happily spend forty minutes on a subtle bug in a small service, and will learn more from it than from a shallow tour of a distributed platform. Depth of understanding is the axis, not scale.

It is not a system design interview. You are not being asked to design something. You are being asked to explain something that exists, including the parts of it that are ugly and the decisions you would take back.

It is not initiative or delivery. The same project supplies all three stories. Here the emphasis is entirely on how you reasoned: what you observed, what you hypothesised, how you tested, and how you knew when you were right.

Story choiceLikely outcome
"The distributed tracing platform my team built" — I owned one collectorCollapses at layer two
"A 0.3% booking loss nobody could reproduce" — I found itSurvives five layers
"Our microservices migration"Too broad; no single line of reasoning to follow
"Why our p99 doubled after a dependency upgrade"Excellent — bounded, evidential, has a resolution

Why interviewers weight this so heavily

Two reasons worth understanding.

It is the hardest thing to fake. A borrowed story survives the summary and dies at the third "why". Someone who was in the room remembers the dead ends; someone who read the postmortem remembers the conclusion. Interviewers know this and drill for it explicitly.

It predicts on-call behaviour. How you narrate a diagnosis is a fair proxy for how you will behave at 3am when a system is failing and nobody knows why. An engineer who guesses, changes three things at once, and declares victory when the symptom disappears is a specific and recognisable risk.

Signals scored

So what does the interviewer write down while you reason? Four positives, four negatives, and one signal that experienced interviewers weight above all the others.

What survives the third follow-upPositive indicators• Explains mechanism, not the label• Names the alternative rejected• Says clearly where knowledge ends• Reasoning holds three layers deepNegative indicators• Vocabulary without mechanism• Only one option ever considered• Guesses rather than saying so• Detail thins as drilling deepens
Experienced interviewers weight one signal above the rest: whether you can name the boundary of what you know.

The positive indicators

1. Systematic narrowing rather than guessing. The observable form is a sequence where each step eliminates possibilities. The API below is an application programming interface — the endpoint one service calls on another:

"The five hops were the booking API, the queue, the consumer, the write to primary, and the read from the replica. I put a correlation identifier through all five and logged at each boundary, so the next missing booking would tell me which hop it died at rather than me guessing. It took two days to get the instrumentation in and about six hours to get an answer once a real case arrived."

The interviewer writes: instrumented to localise before hypothesising.

2. Evidence-based conclusions with the evidence named. Not "we figured out it was the consumer" but "the correlation identifier appeared at the queue boundary and never at the write boundary, for all eleven cases we captured".

3. Knowing the boundary of your knowledge. The strongest single move available:

"I understand the commit semantics of the consumer well because I read that code carefully. I do not deeply understand how the rebalance protocol chooses partition assignments — I know it well enough to know the timing window existed, not well enough to explain the algorithm."

4. Honest uncertainty about the conclusion itself. "We are fairly confident this was the cause because the fix removed the symptom for four months, but we never reproduced it deterministically, so there is a version of this where we fixed a different bug that happened to overlap."

The negative indicators

What you sayWhat gets written
"We tried a few things and eventually it went away"Changed multiple variables; no causal claim
"I knew straight away it would be the database"Intuition presented as method; no falsification
"The team found the root cause"Cannot separate own contribution
Confident detail that turns out to be wrong under a follow-upSerious. Recalibrates every other claim

The one that decides it

Saying "I don't know" cleanly, and then giving the shape of what you do know.

This is counter-intuitive enough that it is worth stating plainly: at senior and above, admitting the limit of your knowledge scores higher than an extra correct answer. The reason is practical. An engineer who fabricates under mild pressure in an interview will fabricate under real pressure in an incident, and that is expensive.

The strong form has two parts — the admission and the useful residue:

"I don't know what the partition count was. I know it was more than one, because the rebalance mattered, and I know it wasn't large because we were running two consumer instances. If it matters for what you're asking, I'd guess low single digits, but that's a guess."

Questions asked

Seven phrasings. Note that several of them are invitations to choose your own story, which makes selection the dominant factor.

Seven phrasings, most inviting your own storyYou choosethe storyHardest problemA bug you chased downA design you ownedA decision you regretSomething you debuggedExplain it simply
When the question lets you pick, selection dominates delivery — pick the system you can defend three layers down.
  1. "Tell me about the hardest technical problem you've solved."
  2. "Walk me through a production incident you debugged."
  3. "Tell me about a decision you made with incomplete data."
  4. "Describe a system you built end to end. Start with why it exists."
  5. "Tell me about a bug that took you a long time to find."
  6. "Tell me about a technical decision you got wrong."
  7. "Pick something from your resume and take me as deep as you can go."

What each is really probing

QuestionEmphasisWhere candidates lose it
1 Hardest problemDepth of understandingChoosing by size rather than by depth
2 IncidentDiagnostic method under time pressureNarrating the timeline with no hypotheses
3 Incomplete dataReasoning under uncertaintyDescribing waiting, or describing a guess
4 System end to endWhether you know why, not only whatReciting an architecture with no rationale
5 Long-running bugPersistence and systematic narrowing"Eventually we found it" with no path
6 Got it wrongHonesty and updatingChoosing a decision that was actually fine
7 Pick from your resumeWhether your resume is defensibleChoosing something you inflated

Question 7 is the one to prepare for

"Pick something from your resume and go deep" removes every excuse. It also means your resume determines the question, which is a good reason to run the evidence audit from How to Write a Good Resume before this loop.

The correct response is not to pick the most impressive line. Pick the line you can defend furthest, say so, and offer the choice:

"The two I could go deepest on are the booking-loss investigation and the scheduling rewrite. The booking one has more interesting reasoning in it; the rewrite has more scale. Which is more useful?"

The follow-on that always comes

Whatever you choose, expect: "can you draw that for me?" — on a whiteboard or a shared document. The story-building steps below cover preparing for it. A candidate who cannot sketch the system they just described in five boxes has told the interviewer something.

Building the story

Two tests, and one preparation task that candidates almost never do.

The preparation almost nobody doesPick a system you ownedDraw it from memoryDefend three layers deepDescribe the resolution
Drawing it before rehearsing it finds the components you cannot actually explain, in private rather than in the room.

Test 1 — you can defend it three layers deep

Take your candidate story and answer these, out loud, without notes:

  1. What was the symptom, precisely? Including how often, since when, and how you knew.
  2. What did you observe that narrowed it? Named evidence, not intuition.
  3. What did you rule out, and how? This is the layer most people cannot reach.
  4. Why did the fix work? Mechanically. Not "it stopped happening".
  5. What is the failure mode of your own fix? Every fix has one.
  6. What don't you understand about it, even now?

If you stall on 3 or 5, the story is not ready. If you cannot answer them at all, this is not your story — you were nearby when someone else solved it, which is common, honest, and not usable here.

Test 2 — it has a resolution you can describe mechanically

"It went away after we upgraded" is not a resolution. A resolution has a mechanism:

"Committing the offset before the write meant a rebalance during a deploy dropped whatever was in flight. Moving the commit after the write turns that into at-least-once delivery instead of at-most-once, so the failure mode flips from losing a booking to writing it twice — which is fine, because the write is keyed on the booking identifier and is idempotent."

That is two sentences and it is the difference between a hire and a lean hire.

The preparation nobody does — draw it first

Expect "can you draw that?". Prepare the sketch in advance, on paper, and be able to redraw it in ninety seconds.

Rules for the interview sketch:

  • Five to seven boxes, no more. You are drawing the path the failure took, not the architecture.
  • Label every arrow with what flows along it. An unlabelled arrow carries no information.
  • Mark where the problem was, with an X or a circle, as you talk.
  • Draw the boundary of your knowledge. A dashed box labelled "someone else's service, I know its contract but not its internals" is a strong, honest move.

Where to look for material

The best deep-dive stories are usually not the biggest projects. Look for:

  • A bug that took more than a week and had a non-obvious cause
  • Anything that was intermittent, because intermittency forces method
  • A performance regression you traced to its source
  • An incident where the first three hypotheses were wrong
  • A design decision you can still explain the trade-offs of, years later
  • Something you fixed that two other people had already tried to fix

The last one is a strong shape. "Three people had looked at this before me" builds difficulty into the setup without you claiming anything about yourself.