Behavioral Interview

Course Content

Behavioral Interview

15 sections · 30 lessons

Initiative: level calibration, anti-patterns and a worked answer


The same category of work reads as mid, senior, or staff depending on scope, autonomy, and influence. This is where candidates most often tell a technically strong story that scores a level below the one they applied for.

This lesson takes one story — flaky continuous integration (CI, the automated build-and-test pipeline that runs on every change) — through all three levels. Then it looks at the two anti-patterns that feel like strong initiative from the inside, and ends with the full senior answer, the follow-up drilling, and the write-up it earned.

MIDOur team's build pipeline failed randomly about 3times a dayI fixed the four flakiest testsTold my tech lead in standupOur pipeline became reliable; the team stoppedre-running buildsSENIOR38% of runs on main failed unrelated to the change— about 9 engineer-hours a weekRoot-caused all four, then added a quarantinepolicy so it could not come backTech lead plus the two maintainers, before starting38% → 4%; policy in the shared template; two otherteams adopted itSTAFFThe flake rate was a symptom — no team ownedtest infrastructure across 11 repositoriesProposed a test-platform owner and acompany-wide flake budgetWrote a proposal, ran the review with the threeloudest sceptics first, got a director to fund half aheadcountFlake budget in the engineering standards doc; 11repos reporting; the class of problem gonePROBLEM FOUNDDECISIONBUY-INRESULTBLAST RADIUSone pipelineone teameleven repositoriesthe shift is from fixing the instance to owning theclass
The problem barely changes between panels; what changes is how much of it the candidate treated as theirs.

The same event at three levels

The work is the same category in all three. What differs is what the candidate chose and who moved.

Mid.

"Our CI kept failing on unrelated changes, maybe two or three times a day, and it was costing everyone time. Nobody had picked it up, so I spent a couple of afternoons on it. I found four tests that were timing-dependent, fixed them, and mentioned it in standup. After that our pipeline was reliable and people stopped blind-re-running."

Scope: one pipeline. Autonomy: chose the task. Influence: nobody. This is a clean, honest, mid-level story and it is a perfectly good answer at mid.

Senior.

"I noticed I was re-running CI a couple of times a day, so I pulled six weeks of pipeline history: 38% of runs on main failed for reasons unrelated to the change. That's about nine engineer-hours a week across seven people. I took that number to my tech lead and to the two people who maintain the pipeline, because if they weren't going to own it afterwards there was no point starting. Fixing the four root causes took a week. The part that made it stick took another three: a quarantine policy that auto-disables a test after two flakes and files a ticket, and getting that into the shared pipeline template. Flake rate went to 4% and stayed there. Two other teams picked up the template."

Scope: one team plus adoption. Autonomy: reframed from "fix tests" to "stop this recurring". Influence: two teams.

Staff.

"The flakiness on our pipeline was real, but I'd seen the same complaint in three other teams' retros, so I treated it as a class. I pulled flake rates across all eleven repos — the median was 30%-plus and nobody owned test infrastructure anywhere. Writing a fix for our repo would have been a week and would have changed nothing structurally. So I wrote an RFC proposing a named owner for test platform and a flake budget teams report against, like an error budget. I ran the review with the three people most likely to object first, one at a time, which is why it passed the group review in twenty minutes. It cost me about six weeks, mostly conversations, and half a headcount from a director. A year later the flake budget is in the engineering standards doc and all eleven repos report against it."

Scope: eleven repos. Autonomy: chose the problem class. Influence: an org practice and a funding decision.

What actually changed between them

MidSeniorStaff
Evidence"it kept failing"six weeks of pipeline historyflake rates across 11 repos
Framingfix the testsstop the recurrenceown the class of problem
Buy-inmentioned in standuplead plus maintainers, before startingRFC, sceptics first, director funding
Durabilityuntil the next flaky testpolicy in the shared templatea reported standard

The technical difficulty is roughly constant. The score is not.

Anti-patterns

Level is one way a good story under-scores. The other is choosing a story that looks like initiative and is not. Two anti-patterns account for most low scores on this competency, and both feel like strong stories from the inside.

Two stories that feel strong from the insideThe weekend heroic• A broken release saved by hand• The underlying gap never fixed• Reads as a process failure you hidThe unsanctioned rewrite• Rebuilt a system nobody sanctioned• No owner agreed to maintain it• Reads as ignoring the org, not leading
Both prove you can act alone; neither proves the organisation was better off for it.

Anti-pattern 1 — heroics that should not have been necessary

"We had a release going out on the Friday and the migration script failed at 6pm. I stayed until 3am rewriting it by hand, row by row, and we shipped on time. My manager said it saved the launch."

The interviewer's write-up:

"Candidate volunteered a story where they personally absorbed a systemic failure. When asked what changed so it couldn't recur, said 'we were more careful next time'. No process change, no test, no rollback plan. Initiative: no evidence. Some concern about judgement — a 3am manual data edit on production is a decision, and it wasn't discussed with anyone."

Heroics stories are seductive because they contain effort, drama, and a good outcome. They score badly because the rubric is looking for choosing what to work on, and this candidate chose nothing — the situation chose for them. The same event becomes a strong story with one addition: what you built afterwards so it could not happen again.

Anti-pattern 2 — the unsanctioned rewrite

"The legacy sync layer was awful, so over about six weeks I rewrote it properly in my own time and then showed the team. They didn't merge it in the end, but I learned a lot."

This scores negatively rather than neutrally. The interviewer records: acted at large scale without agreement, produced no adoption, and the failure is attributed to the team's reception rather than to the approach. At senior and above it raises a real question about whether the candidate can work inside a team.

The tell is the phrase "and then showed the team". Buy-in came after the work, so there was never any.

A worked example

Here is the senior version told in full, with the drilling that followed.

The follow-ups

The story is the entry point. This is where the score is decided.

Interviewer: You said thirty-eight percent. Where did that number come from, exactly?

Devi: The CI system keeps run history with the exit status and the failing job. I pulled six weeks of runs on main into a spreadsheet — about 210 runs a week — and classified the failures. The hard part was deciding what counted as "unrelated to the change". I used a proxy: a run counted as flaky if the identical commit passed on a re-run with no code change. That undercounts, because some real failures also pass on retry, so 38% is a floor rather than a point estimate. I said so at the time.

Interviewer: And the nine hours a week?

Devi: That one's softer. Fourteen minutes is the median pipeline duration, which is measured. The context-switch cost I estimated rather than measured — I used ten minutes per interruption from memory of what the team said, not from data. If someone had pushed on it I'd have said the defensible number is the pipeline time, about five hours a week, and the rest is judgement.

That answer scores well specifically because Devi separated the measured number from the estimated one without being pushed to.

Interviewer: Who pushed back on the quarantine rule, and what was their argument?

Devi: One of the pipeline maintainers, and their argument was the right one: a quarantined test is a test that isn't protecting you, and the failure mode is that it stays quarantined forever and something real ships. They'd seen exactly that at a previous company — a suite where about a third of tests were permanently disabled and nobody knew which.

Interviewer: So how did you resolve it?

Devi: I didn't argue them out of it, because they were right. I changed the proposal. The cap on concurrent quarantines came directly from that conversation, and so did putting the count on the dashboard where the team sees it daily. I also agreed a rule that a quarantined test older than two sprints becomes a blocking item in the team's planning. The thing I got wrong initially was proposing a mechanism without a way for it to fail loudly.

Interviewer: What did you stop doing to make room for six weeks of this?

Devi: Honestly, not enough at first — the first two weeks came out of my evenings, which was a bad call and not sustainable. When I took the numbers to my lead I asked for it to be real work, and we dropped a piece of endpoint cleanup that had been sitting in the backlog for two quarters. That was the right trade: the cleanup was worth maybe a day of future time, and this was worth five hours a week forever. I should have had that conversation in week one instead of week three.

What the write-up said

"Initiative — strong hire. Identified the problem from data rather than annoyance; quantified before acting; explicitly separated measured from estimated figures unprompted. Sought buy-in from the eventual owners before investing, which is the behaviour we want at senior. Changed their own proposal in response to a good objection rather than defending it. Carried through past the interesting fix into policy and adoption (2 teams). Named a real cost trade-off and a mistake in how they funded the first two weeks. Also usable evidence for Earning Trust."

Three things did the work: a number with its measurement described, an opponent's argument stated at full strength, and a mistake named without being asked.