Behavioral Interview

Course Content

Behavioral Interview

15 sections · 30 lessons

Deep dive: level calibration, anti-patterns and a worked answer


The deep dive calibrates level on the boundary you reasoned across, not on the difficulty of the code.

This lesson shows that with one investigation told three ways, then covers the two anti-patterns that end deep dives — both usually invisible to the candidate until the third follow-up — and finishes with a full senior answer drilled five layers deep.

MIDInside our booking service — logs, a profiler, thequery planThe bug and its fixOne bug closedSENIORAcross the queue boundary, following a correlationID through 5 hops — hop 3 was not mine, so Iworked from its contractThe diagnosis, the fix, and the conversation withthe team that owned hop 3Bug closed, plus the consumer pattern fixed in 2other servicesSTAFFAcross a class of incidents — 11 postmortems readside by side, 3 sharing one causeThe systemic cause, and getting someone to ownitThe incident class removed from the register; a lintrule in CIWHERE I LOOKEDWHAT I OWNEDWHAT CHANGEDAFTERBOUNDARY CROSSEDinside one serviceacross servicesacross a class of incidentssaying what you do NOT own is a positive signal,not a gap
The debugging skill is identical in all three; the boundary the candidate was willing to cross is not.

The same investigation at three levels

The symptom: about 0.3% of appointment bookings confirmed to the user never appeared in the clinic's list.

Mid — debugging within a service.

"I got the ticket for the missing bookings. I could not reproduce it, so I added logging around the write path in our booking service and waited. When a case came in, the log showed the write had never been attempted. I read the consumer code carefully and found the offset was committed before the database write. I moved the commit after the write, made the write idempotent on the booking identifier, and the reports stopped."

Real diagnosis, real fix, one service. That is a good mid answer.

Senior — diagnosing across service boundaries.

"Three people had looked at this over four months and each had checked their own service and found it clean, which was the problem — nobody had looked at the joins. So the first thing I did was not debug. I put a correlation identifier through all five hops and logged at each boundary, so the next occurrence would tell us where it died instead of us arguing about whose service it was.

Eleven cases later the answer was unambiguous: present at the queue boundary, absent at the write boundary, every time. That eliminated the booking application programming interface, the queue itself, and the replica-lag theory that everyone had assumed, which mattered because two engineers had spent a sprint on read-after-write consistency for nothing.

Then I correlated the timestamps against our deploy log: ninety-four percent of losses were within ninety seconds of a consumer restart. The consumer committed its offset before doing the write, so a partition rebalance during a deploy dropped whatever was in flight. We deployed about twice a day, which is where 0.3% comes from.

The fix was to commit after the write and make the write idempotent on booking identifier — which flips the failure mode from losing a booking to occasionally writing it twice, and a duplicate is harmless here. Then I grepped for the same pattern and found it in two other consumers, one of which was handling payment events."

Same root cause. The difference is the boundary crossed, the hypotheses explicitly eliminated, and the generalisation to two other services.

Staff — the systemic cause behind a class.

"The booking loss was one instance. What made me look wider was that I'd seen three postmortems in the previous year with 'messages lost during deploy' in them, all closed as one-off fixes.

I read eleven postmortems side by side. Nine of them involved a consumer written against our shared client library, and the library's example in the internal documentation committed the offset first. Every team had copied the example. So the systemic cause was not any team's code — it was a default in a library plus an example that taught the wrong thing.

What I did about it: changed the library default to commit-after-processing, which is a breaking semantic change, so that took a written proposal and a migration path. Added a lint rule that fails the build on manual pre-commit. And rewrote the documentation example, which took twenty minutes and was arguably the highest-value part.

The incident class hasn't appeared on the register since. What I'd do differently: I read those eleven postmortems eight months after the third one. We had the data to see the pattern a year earlier and nobody was looking across them, which is now a quarterly review I set up."

What changed between them

MidSeniorStaff
Boundary reasoned acrossInside one serviceFive hops, three teamsEleven incidents, one library
What was eliminatedNothing explicitlyReplica lag, the API, the queuePer-team explanations
GeneralisationNoneTwo other consumersA library default and a lint rule
Reflection"Add more logging earlier""Instrument before hypothesising""Nobody was reading postmortems across teams"

Anti-patterns

A well-calibrated story can still fail if it does not hold under questioning. Two anti-patterns end deep dives, and both are usually invisible to the candidate until the third follow-up.

The five follow-ups that end a deep diveWhy that approach?What did you rule out?What broke first?How did you know that?What would you change?
A story you did not build, or credit borrowed from a teammate, both fail at the same depth — usually the third question.

Anti-pattern 1 — the story you cannot explain under questioning

Candidate: "...so we identified that the issue was in the consumer's offset handling and corrected it."

Interviewer: How did you identify that?

Candidate: We looked at the logs and traced it through.

Interviewer: What did the logs show, specifically?

Candidate: That the message wasn't being processed properly.

Interviewer: What did you rule out before that?

Candidate: I mean, we checked a few things... I'd have to look back at the notes.

Three questions, and the write-up is already written. Note that the candidate said nothing untrue. They may well have been on the team. But the reasoning is not theirs, so there is nothing under the summary.

The tell is abstraction that increases under pressure. When a real memory is questioned it gets more specific. When a reconstructed one is questioned it gets vaguer.

Anti-pattern 2 — credit borrowed from a teammate

Distinct from the first, and worse, because it involves a claim. A candidate tells their colleague's investigation in the first person, usually with a rationalisation ("I was on the team, it was a joint effort").

Interviewers catch it two ways. Follow-ups about the dead ends — which the borrower does not know, because postmortems record conclusions, not failed hypotheses. And follow-ups about feelings and timing: "what did you think it was on day one?" is unanswerable secondhand.

The honest version costs you nothing and is often better:

"This was Priya's investigation, not mine — I want to be straight about that. My part was the fix and the generalisation to the other two consumers, and I can go deep on both. Do you want the diagnosis as I understood it, or shall I take you through the part I owned?"

A worked example

Here is the senior version of the booking-loss investigation as a full answer, then five escalating follow-ups.

The five follow-ups

1. Interviewer: How did you know it wasn't replica lag? That's the obvious explanation for something appearing to be missing.

Ola: Two things. The direct evidence: the correlation identifier never reached the write boundary, and replica lag would mean the write happened and the read was stale. Lag can't explain an absent write log line.

The confirming evidence: we re-queried those eleven booking identifiers against the primary a week later and they weren't there either. Lag resolves. This didn't. I want to be fair to the people who chased it, though — lag was a reasonable first hypothesis, because the symptom presents identically and our replica lag did spike during deploys, which is a real correlation with the actual cause.

2. Interviewer: You said you decided not to try to reproduce it. Walk me through that decision — I don't think it's clearly right.

Ola: It isn't clearly right, no. The argument for reproducing is that a deterministic reproduction is worth more than any amount of logging, and if I'd found one in a day I'd have saved a week.

What I weighed: three people had failed to reproduce it, so my prior on succeeding in a day was low. The event rate was high — forty thousand a day at 0.3% is over a hundred cases a day — so instrumentation would give me real cases within hours of shipping, not weeks. And the instrumentation had residual value: we still use those correlation identifiers.

If the event rate had been one a month, I'd have made the opposite call.

3. Interviewer: Why commit-after-write rather than a transactional outbox? That's the textbook answer for exactly this.

Ola: It is, and we considered it. The outbox pattern would have given us exactly-once-ish semantics between the write and the publish, and it's the right answer if you need the write and the publish to be atomic.

Our problem was one step downstream of that — we weren't losing the publish, we were losing the consume. An outbox on the producer side wouldn't have fixed it. What would have fixed it is transactional consumption, and our infrastructure at the time didn't support it in a way I trusted.

The honest reason, though, is cost. Commit-after-write plus an idempotent write was about a day of work and it moved the failure mode to one we could tolerate. An outbox across three consumers was a two-month project. I'd rather take the cheap change that makes the bad outcome impossible than the correct architecture that ships next quarter.

4. Interviewer: So what's the failure mode of your fix? You've made duplicates possible.

Ola: Correct, and that was the trade I made deliberately. Committing after the write means at-least-once delivery, so a crash between the write and the commit replays the message.

The write is idempotent on the booking identifier, which is generated by the client, so a replay is an upsert with identical content. I checked the two things that could break that: the booking record has an updated_at field, so a replay does bump the timestamp — cosmetic — and there was a notification side effect. The consumer sent a confirmation email after the write. A replay would have sent a second email. I moved the email to a separate consumer keyed on a de-duplication table with a twenty-four hour window.

There's a residual case I did not solve: if the de-duplication table write fails after the email is sent, you get a duplicate email. We accepted that. A duplicate confirmation email is a support ticket; a lost appointment is a patient standing in a clinic.

5. Interviewer: Where does your understanding of this run out? What don't you know?

Ola: Two places, and one of them bothers me.

The one I'm relaxed about: I don't understand the rebalance protocol properly. I know enough to know the window exists and roughly why — the consumer group reassigns partitions and in-flight work on a revoked partition is abandoned — but I could not tell you how the assignment strategy actually works or what the timing guarantees are. I never needed to, because the fix doesn't depend on it.

The one that bothers me: I never reproduced it deterministically. Everything I have is correlational — the boundary evidence, the ninety-four percent deploy correlation, and the symptom disappearing after the change. That's a strong case but it isn't a proof. There is a version of the world where I fixed a different bug that happened to have the same shape and the real one is dormant. I put a metric on it — a counter for messages consumed but not written — so if I'm wrong we'll find out from a graph rather than from a clinic. It has been flat for eight months, which raises my confidence but doesn't close it.

What the write-up said

"Problem Solving — strong hire. Chose a bounded problem and held depth across five escalating follow-ups without degrading. Made an explicit, defensible decision not to reproduce and could argue the counterfactual (would have chosen differently at low event rate). Named what was eliminated and how. Knew the failure mode of their own fix, including the second-order email side effect and the residual case they accepted, with a stated reason for accepting it. Volunteered two limits of their own knowledge, distinguishing the one that doesn't matter from the one that does, and had instrumented against being wrong. No defensiveness on the outbox question. Also strong evidence for Delivery and Learning."

The last follow-up is the one that produced the highest-value paragraph in the interview. It is worth noticing that the paragraph is entirely about what Ola does not know.