Behavioral Interview

Course Content

Behavioral Interview

15 sections · 30 lessons

Innovation: anti-patterns and a worked answer


Two anti-patterns account for most negative innovation scores, and both are told by candidates who believe they are describing their most inventive work.

This lesson takes them apart, then follows a worked example built entirely on deletion — the kind of story the previous lesson argued is usually the strongest — through four follow-ups.

The worked example is a deletionA subsystemnobody usedRemovalrisk reasonedDeletedbehind a flagNeverrebuilt sinceContrast: new technology chosen first, problem found later.
Innovation scores on the judgement, not the novelty, and deleting something is often the harder judgement to defend.

Anti-pattern 1 — novelty for its own sake

"We were building a new service and I pushed for us to use an event-sourced architecture with separate read and write models — CQRS, command query responsibility segregation. It was more interesting than another straightforward create-read-update-delete service and it meant we could replay state. It took longer than expected and honestly the team found it hard to work with, but I learned a lot."

The write-up:

"Innovation — negative signal. Adopted a substantially more complex architecture without a stated problem it solved. When asked what the simpler alternative would have cost, could not answer. Acknowledged team difficulty but framed it as a learning outcome for themselves. This is a judgement concern rather than an innovation gap."

The tell is the ordering. The candidate reached for the technique and then looked for a justification. Innovation stories run the other way: a problem, then a constraint, then a technique.

Anti-pattern 2 — the unjustified rewrite

"The old service was five years old and nobody understood it, so we rewrote it. It took about nine months. It's much better now."

Three problems in one story. There is no stated problem beyond age. There is no alternative considered — incremental refactoring, strangler-pattern replacement, or leaving it alone. And nine months is expensive, so the absence of a cost-benefit argument is conspicuous.

Rewrites can be excellent innovation stories. What makes the difference is a named failure of the old system, a costed alternative, and a validation step:

"The problem wasn't age, it was that the median change took eleven days to ship because every change required a full regression suite that took six hours and was 40% flaky. We costed three options: fix the tests, strangle it endpoint by endpoint, or rewrite. We spent two weeks trying to fix the tests first, precisely so we could say we had, and got the suite to four hours and 25% flaky, which wasn't enough. Then we strangled it over seven months rather than rewriting, so we were never in a position where neither system worked."

A worked example

The follow-ups

Interviewer: What would have made you abandon the on-demand approach?

Bo: Two thresholds, and I set them before I ran the replay.

If the ninety-six percent had been under about sixty, the saving wouldn't have justified the change and I'd have gone with shrinking the variant list, which was a two-day fix. And if the request distribution had been flat rather than concentrated — if no small set of variants dominated — then the warm-set idea doesn't work and every request is a cold generation, which would have made the latency argument much harder.

There was a third one I didn't set in advance and should have: what happens under a traffic spike on a new article. That's the case where on-demand is worst, because every variant is cold at once. I discovered that in load testing rather than in planning.

Interviewer: And what happened there?

Bo: It was worse than I expected. A new article going out to the newsletter produced a burst where about four hundred distinct variant requests hit within a minute, all cold. Our generation workers saturated and p95 went to about two seconds for ninety seconds. We fixed it by pre-warming the three common variants at publish time rather than at upload time, which is a small piece of the old pipeline coming back — and I think that's the honest framing. I didn't delete the concept of pre-generation, I moved it from every image at upload to a few variants at publish.

Interviewer: Who owned the tiering plan you displaced, and how did that conversation go?

Bo: A colleague on my team had scoped it, and they'd done the work properly — it was a reasonable plan for the problem as understood. I want to be careful here because this is the part I handled worst.

What I did was bring the log replay to our team meeting, which meant they found out their plan was being displaced in front of five people. That was thoughtless. They were fine about it publicly and it was clear afterwards that it had landed badly, which is fair.

What I should have done, and what I do now, is show the numbers to the person whose plan they affect first, privately, before anyone else sees them. It costs one conversation and it changes the whole thing from "your plan is wrong" to "look what I found, does this change your read?".

Interviewer: You said compute for something requested once a month is close to free. Did you check that?

Bo: Roughly, and I'd flag it as the weakest number in what I told you. I estimated generation cost per image from the batch's existing compute, which was about eight hundred milliseconds of CPU per variant, and multiplied by the request volume for the long tail. That came out around four hundred a month against the eleven thousand saved, so the margin was wide enough that I didn't refine it.

What I didn't model was the cost of the workers sitting idle between bursts, because we run them on shared capacity. If we'd needed dedicated capacity for this, the arithmetic would have been much closer and I'd have needed a real number rather than an estimate.

What the write-up said

"Invent and Simplify — strong hire. Identified and named the unstated assumption (that all variants are requested) and traced why it had gone unexamined. Validated in an afternoon using data that already existed, with pre-registered thresholds for abandoning the approach — this is the behaviour we most want and rarely see. Costed three options, not one. Outcome is a deletion: 800 lines and an on-call surface removed. Volunteered a load-test failure the planning missed and described the partial reintroduction of pre-generation honestly rather than claiming a clean win. Volunteered a real interpersonal mistake in how the displacement was handled, with a specific changed practice. Named their own weakest number unprompted. Also usable for Earns Trust and Customer Focus."

Two things carried it: thresholds set before the test rather than after, and the honest framing that the deletion was partial rather than total.