Course Content
Behavioral Interview
15 sections · 30 lessons
Delivery: level calibration, anti-patterns and a worked answer
Delivery is the competency where the level gap is widest, because the mid-level version is genuinely good work that does not reach past the individual.
This lesson takes one project through all three levels, then looks at the two anti-patterns candidates most often believe are their best story, and finishes with a full senior answer and the drilling that followed it.
The same project at three levels
The project: a payment provider decommissioning its v1 application programming interface (API, the contract other systems call) on a fixed date fourteen weeks out, against about twenty-two weeks of scoped work across three teams.
Mid — delivering your own task.
"I owned two of the six workstreams: the refunds path and the webhook receiver. The refunds work was blocking two other people, so I re-ordered my own backlog to land it first even though the webhook piece was more interesting. In week four I realised refunds was going to take nine days rather than five, because the v2 API returns partial refunds differently, and I told my tech lead that week rather than at the end of the sprint. Both my pieces were done two weeks before cutover."
Good work. Scope: own tasks. Autonomy: sequencing within an assignment. Influence: unblocked two people. This is mid, and it is a solid mid answer.
Senior — unblocking the critical path.
"I ended up owning the cutover across all three teams, mostly because I was the one who put the dependency graph on a wall. Twenty-two weeks of work, fourteen weeks of calendar, and no extension — the API returns errors after the thirtieth, that's it.
I used one criterion for the cut list, and I stated it up front so the arguments were about the criterion rather than about people's features: does checkout return an error without this? Three of the six workstreams were in that set. The payment-method selector redesign, the saved-card UX work, and analytics parity were not, so they were deferred. I wrote a one-page note with what was deferred and when we'd come back to it, and I had both affected product managers acknowledge it before we started, in writing.
The dependency risk was the fraud team's integration, scheduled for August. In week two I asked for their sprint plan rather than their roadmap date, and it was clear August was optimistic. So I built a compatibility shim that let us cut over without them.
We cut over eleven days early. The three deferred items shipped in Q4."
Scope: three teams. Autonomy: set the cut criterion. Influence: two product managers and a dependency team re-sequenced.
Staff — making it deliverable at all.
"My read after two weeks was that it was not deliverable as scoped by any arrangement of the three teams — twenty-two weeks of work, fourteen of calendar, and the estimate had no slack for the unknowns in a payment cutover. So the useful thing I could do was not manage it harder, it was change the shape.
I restructured it into three independently shippable phases. Phase one was the minimum that keeps checkout working. Phase two was everything with a customer-visible benefit. Phase three was the analytics and reporting parity work, which I proposed we cancel outright rather than defer, because deferred work of that kind sits in a backlog for two years and costs a re-planning conversation every quarter.
Cancelling was the contentious part. I wrote the rationale down — what we lose, who is affected, what we would do if it turns out to matter — and took it to the three engineering managers and the director before the group meeting. Two of them disagreed at first. The data that moved them was that the reporting in question had four internal users and they had a workable manual path.
Phases one and two shipped. Phase three was cancelled with a written rationale, and nobody has asked for it in the eighteen months since. And nobody worked a weekend."
Scope: the program. Autonomy: changed the shape of the work. Influence: three managers, a director, a cancellation decision.
What changed between them
| Mid | Senior | Staff | |
|---|---|---|---|
| What was at risk | Two workstreams | The cutover | The program's feasibility |
| What they could change | Their own sequence | The scope | The structure of the work |
| Cut authority | None | Deferred three items with sign-off | Cancelled work outright |
| Evidence of communication | Told their lead in week 4 | Written deferral note, acknowledged | Written rationale, sceptics first |
The escalation is in what you were allowed to change, not in how hard the engineering was.
Anti-patterns
Two anti-patterns, both extremely common, both of which the candidate believes are their best story.
Anti-pattern 1 — everything went fine
"We had a big launch last year. I led the backend work. We planned it out, we hit our milestones, and it shipped on the date we said. It went really smoothly."
There is nothing wrong with this as work and nothing in it as evidence. The write-up:
"Delivery — no evidence. Candidate described a project with no constraint, no trade-off, and no adaptation. Prompted twice for something that went wrong; got 'nothing major'. Cannot assess judgement under pressure."
The instinct behind it is understandable: candidates think they are being asked to prove they are reliable. They are being asked to prove they can make decisions when the plan stops working. A frictionless project does not contain that evidence, so choosing one is choosing to be unscoreable.
If the smooth project is genuinely your biggest, tell it — but find the tension inside it. There is always a week where something was at risk.
Anti-pattern 2 — the death march as a badge
"It was brutal. We were doing sixty-hour weeks for two months, I was working most Sundays, but we got it out the door. That's the kind of thing I'll do when it matters."
Told as a strength, scored as three separate concerns: the estimate was wrong and not corrected, the escalation did not happen or did not work, and the candidate is signalling that sustained overwork is their answer to pressure. For a manager candidate, this is close to disqualifying — it describes what they will do to a team.
The salvage, when this really is the story: keep the facts and move the emphasis to what you changed. "We did two months of sixty-hour weeks, which was a failure of the plan, not a triumph. What I did about it was —" and then the escalation, the cut, or the structural change you made afterwards so it did not recur.
A worked example
Here is the senior version of the payments cutover told as a full answer, then drilled.
The follow-ups
Interviewer: You said you cut three of six workstreams. Who actually made that decision?
Marco: The criterion was mine and I proposed the cut list. The decision was made by the two product managers and my engineering manager in a forty-minute meeting. I don't want to overclaim — I couldn't have cancelled the selector redesign on my own authority. What I did was make the decision easy: I brought one criterion, the six items scored against it, and a written proposal with dates for the deferred work. The meeting was short because there wasn't much left to decide.
Interviewer: Did anyone fight it?
Marco: One of the product managers, on the saved-card work, and their argument was legitimate — it had been promised to a specific enterprise customer for that quarter, in writing, by someone in sales. My criterion didn't have a way to see that. What we did was keep the promise but change its shape: I found that about eighty percent of what had been promised was achievable with a small change to the existing flow rather than the redesign, so we shipped that in October and the full redesign in Q4. That was their idea, not mine.
Interviewer: Let's go back. At what point did you know fourteen weeks wasn't enough for twenty-two weeks of work?
Marco: Immediately — that arithmetic was clear on day one. The thing I didn't know on day one was whether the twenty-two-week estimate was real. So the first week was spent testing it rather than acting on it. I asked each of the three teams for their estimate broken into pieces of a week or less, because estimates in units of "about a month" are usually a month of not having thought about it. That exercise moved the total from twenty-two weeks to about nineteen, which was still not fourteen, and that was the point I stopped trying to fit it and started cutting.
Interviewer: And when did you tell your director?
Marco: End of week one, with the cut proposal attached. I've learned not to bring a problem of that size without at least one option, because the meeting where you bring only the problem generates three days of anxiety and no decisions.
Interviewer: What went wrong?
Marco: Two things. The smaller one: the shim I built for the fraud path leaked memory under sustained load and we found it in shadow traffic about a week before cutover. That was a fair catch by our load test and it cost two days.
The bigger one is the one I'd want to tell you about. Refunds. The v2 API handles partial refunds with a different idempotency model — v1 keyed on our reference, v2 keys on theirs. We migrated the charge path and the refund path in the same release, and for about six hours after cutover, a retry on a partial refund could issue a second refund. We caught it because I'd asked for a reconciliation job comparing our ledger to theirs hourly during the cutover week, which was the one piece of paranoia I'm glad I insisted on. Eleven duplicate refunds, about £3,400, all recovered.
Interviewer: Why didn't the tests catch it?
Marco: Because our integration tests used their sandbox, and the sandbox accepted our idempotency key without enforcing it. It behaved like v1. That's on me — I knew the sandbox was lenient in other areas and I didn't think through what that meant for a semantic change rather than a syntactic one. What I do now is: for any external migration, I write down the behaviours the sandbox does not reproduce, before writing tests against it. That list is usually short and it is always the interesting part.
What the write-up said
"Delivery — strong hire. Named an external immovable constraint with a quantified gap (14 wks vs 22). Tested the estimate before acting on it rather than accepting it — good instinct. Explicit, stated-in-advance cut criterion, and did not overclaim the decision authority when asked directly. Detected a dependency risk by asking for the sprint plan rather than the roadmap date, and routed around it instead of escalating. Escalated to director in week 1 with options attached. Volunteered a real production defect with a money figure, explained the root cause down to sandbox fidelity, and produced a generalised rule from it. Also strong evidence for Problem Solving and Earning Trust."
Notice what the write-up rewards: the estimate test, the stated criterion, the honest boundary on decision authority, and the refunds defect volunteered rather than extracted.