What AI automates, what it does not, and what to build a career on

JR

Jai Rao

August 22, 202616 min read

A concrete look at which parts of engineering work AI actually absorbs, which parts resist it, and which skills keep compounding once producing code gets cheap.


The question people ask about AI and engineering careers — will this replace me — is the least useful version of the question, because it treats a job as one indivisible thing. Jobs are bundles of tasks, and the tasks are not equally exposed. Writing a well-described utility function from a clear signature is very exposed. Working out why a reconciliation job silently drops a fraction of a percent of records on the last day of the month is not, and that does not change quickly, because the hard part of that work is not producing text.

So the practical move is to stop reasoning about job titles and start reasoning about task properties. If you can say what makes a piece of work automatable, you can look at your own last two weeks, sort them, and see which half is shrinking. That is a decision you can act on this month rather than a prediction you have to wait for. What follows is a description of those properties, what happens to the shape of engineering work as a result, and which skills actually accumulate value instead of being priced away.

Three properties that make a task automatable

A task is exposed to automation to the degree that it has three properties at once. Any one of them alone is not enough; the combination is what makes handing the work off safe.

Well-specified. There is a description of the desired result that does not require a conversation to interpret. "Convert this snake_case API response into camelCase TypeScript interfaces" is well-specified. "Make onboarding feel less heavy" is not — that is an opening position in a negotiation with a designer, a support lead, and whoever owns the activation metric. The work in the second case is producing the spec, and the coding is the easy tail end of it.

Verifiable. You can tell whether the output is right without redoing the work yourself. A pure function with a decent test suite is verifiable in seconds. A change to how a distributed lock is acquired is not: it looks fine, it passes, and you find out three weeks later when two workers process the same job during a network partition. What matters is the cost of a wrong answer multiplied by how long it takes to notice. When that product is small, delegation is cheap. When it is large, the review is the job.

Low-context. The information needed to do the task fits in what you can actually hand over. Most genuinely hard engineering context was never written down: why that table has a redundant column (a migration that half-finished years ago), why the retry count is three and not five (an undocumented vendor rate limit), why nobody touches that service on a Friday. Anything working without access to that will produce something locally reasonable and globally wrong, and it will do so confidently.

Scoring real tasks on those three axes is more informative than any general claim about the industry. A rough pass over four tasks from an ordinary sprint:

TaskSpecified?Verifiable?Context neededExposure
Add cursor pagination to a REST endpoint that has twelve siblings doing it alreadyYesYes, contract testsLowVery high
Backfill a new column across forty million rows without locking writesMostlyPartly — correct on a sample, unknown under production loadMediumModerate
Find why checkout completion dropped after Tuesday's releaseNoOnly in hindsightHighLow
Decide whether to shard now or buy eighteen more months with a bigger instanceNoNot for eighteen monthsVery highNear zero

The first row is work you should already be delegating and reviewing rather than typing. The last row is work that no tool takes from you, because the difficulty is not in generating the options — there are only two — but in weighing them against a roadmap, a hiring plan, and a team's tolerance for operational pain.

The work that stays hard, and the reason it stays hard

Three categories resist automation for structural reasons rather than temporary capability gaps.

Ambiguous requirements. When a request arrives as "reporting is confusing", the deliverable is a decision about what reporting should mean, and that decision is contested. Two stakeholders want different things and neither has the authority to settle it. Producing a defensible answer requires knowing who is actually affected, which of the stated wants are proxies for something else, and what you are willing to not build. Handing that to a tool produces a plausible interpretation of an underdetermined problem, which is exactly the failure mode you were trying to avoid.

Cross-system debugging. A page shows stale data after a mobile release. The chain runs through a client-side cache, a CDN that varies on a header somebody changed, a read replica with variable lag, and a feature flag evaluated per request. No single system's logs contain the answer; the answer lives in the mismatch between them. Progress comes from forming a hypothesis and choosing which instrument to reach for next, paying for that choice in minutes and in production access. Expensive steps, partial evidence, and a mental model deciding where to look next: close to the opposite of a well-specified task.

Judgement under conflicting constraints. Latency against cost, deadline against the team's ability to operate the thing at three in the morning, correctness against a customer who is threatening to leave. These have no correct answer, only defensible ones, and the defence has to hold up to people who will remember it in six months. What you are paid for is the willingness to be accountable for the call, and accountability does not delegate.

The centre of the job moves from producing code to specifying and reviewing it

The bottleneck used to be typing. It is now reading. That sounds like a smaller job and it is not, because reading unfamiliar code you are accountable for is slower than writing familiar code you already understand. A senior engineer who reviews thirty pull requests a week already knows this; the change is that the ratio of code-you-wrote to code-you-approved is shifting for everyone, including people who never had a review load before.

Two failure modes show up immediately. The first is rubber-stamping: the diff looks like something you would have written, the tests are green, you approve. The second is silent rewriting: you distrust the output, redo it by hand, and get none of the speed. Both come from the same root cause, which is that the accept-or-reject decision is expensive. Making it cheap is an engineering problem with familiar answers — keep diffs small enough to hold in your head, make invariants fail loudly instead of degrading quietly, put the risky behaviour behind a test that would actually catch it, and prefer changes whose blast radius you can state in one sentence.

The other half of the shift is upstream. A precise specification is a promise about behaviour at the boundaries: what happens on empty input, on a duplicate, on two concurrent callers, on a payload ten times larger than expected, on the dependency being down. Engineers who were already good at writing tickets that do not come back with questions are good at this. Engineers who described work loosely and filled in the gaps while coding now discover that the gaps get filled in for them, by something with no stake in the outcome.

The blunt version: a diff you did not read is not a time saving. It is a liability with your name in the blame output.

Verification is the skill that quietly decides everything

Here is a task small enough to fit in a paragraph and ambiguous enough to be dangerous. A room-booking view needs overlapping time spans merged so the calendar can draw them as blocks. The request, as it usually arrives, is one sentence: merge overlapping spans in a list of start and end hour pairs. This is the implementation you get, and it is the implementation most people would write by hand:

Text
def merge(spans):    out = []    for start, end in sorted(spans):        if out and start <= out[-1][1]:            out[-1][1] = max(out[-1][1], end)        else:            out.append([start, end])    return out

Read it closely and it is clean, sorted correctly, handles containment, and runs in O(n log n). It is also correct for one reading of the request and wrong for another, because nothing in that sentence said whether a span ending at ten and a span starting at ten overlap. Now the tests, which are where the decision has to live:

Text
def test_touching_spans_stay_separate():    # Bookings are half-open [start, end): 09:00-10:00 followed by    # 10:00-11:00 is two distinct meetings, not one two-hour block.    assert merge([(9, 10), (10, 11)]) == [[9, 10], [10, 11]]def test_real_overlap_merges():    assert merge([(9, 11), (10, 12)]) == [[9, 12]]def test_contained_span_disappears():    assert merge([(9, 17), (12, 13)]) == [[9, 17]]

The first test fails. With start <= out[-1][1], the two bookings collapse into a single nine-to-eleven block and the room shows as busy for a slot that is free. The fix is one character, < instead of <=, and the character is not the point. The point is that a product decision — touching spans are distinct — existed nowhere in the request, and the only participant capable of knowing it was a person who understood how rooms get booked. Reviewing generated code is mostly this: finding the decisions the request never made, and pinning them down somewhere durable before they turn into a support ticket.

Reading the implementation carefully still only compares code against a sentence; a test compares it against a behaviour someone committed to. That gap is why verification skill outlasts nearly everything else here — it is what converts "looks right" into "is right".

The apprenticeship gap this creates

Junior engineers historically learned by doing the exposed tier of work. The third variation on a form, the CRUD endpoint, the small parser, the fiddly date handling: unglamorous, well-specified, cheap to get wrong. Getting it wrong in small ways and then reading the review comments is how calibration develops — a feel for which mistakes are likely, which parts of a codebase are load-bearing, how much a seemingly trivial change can cost. That tier is precisely the tier that scores high on all three properties, which means it is the tier that is thinning out first.

This is a real problem, and the ladder is losing its bottom rungs while interviews still test for the judgement those rungs used to build. Two responses actually help.

If you are early: write the thing yourself first, then generate a version and diff them. The diff is the lesson, and it arrives in two minutes instead of two days of waiting for review. Volunteer for debugging work specifically — it is the least automatable category and the fastest way to build a real model of a system, because you cannot fake your way through a stack trace that spans three services. And read code you neither wrote nor generated; a codebase you can navigate under pressure is worth more than a framework you have skimmed.

If you lead a team: stop counting merged output, because it is now nearly free and therefore nearly meaningless as a signal. Give juniors end-to-end ownership of something small, including its alerts and its on-call pages, since ownership is what produces judgement. And treat review as the teaching surface it now is — a comment explaining why an approach is risky is worth more per minute of senior time than it was when the junior had already spent a day writing the code.

Skills that compound

Four capabilities get more valuable as the volume of code nobody on the team wrote by hand goes up. They compound because each one makes the others cheaper to apply.

Reading unfamiliar systems fast. Dropped into a repository you have never seen, can you find where a request enters, where it touches storage, what it does on failure, and which parts are load-bearing — in an hour? This underwrites everything else, including the review load described above. It is also trainable in a way people neglect: pick a dependency you use every day and trace one call from its public API down to the system call, without skipping the layers you find boring.

Writing specifications that pin down boundaries. Not documents. The habit of asking, before anything gets built, what the inputs are that break this, what the behaviour is on empty and duplicate and concurrent and oversized, and what should happen when the dependency below is down. A large fraction of what gets called good prompting is just this discipline performed in a text box, which is why people who already had it adapted in about a week.

Evaluation instinct. Knowing what evidence would convince you that something works, and refusing to be convinced by less. For deterministic code that is tests and invariants. For a probabilistic component it is a labelled set, a baseline to beat, and a metric that actually moves when quality moves — plus the discipline to say out loud that "it looked good on the five examples I tried" is not evidence of anything. This transfers directly into model work, evaluation harnesses, and any system whose output is a distribution rather than a value, which is why it is the most underpriced skill on this list.

Data literacy. Not a statistics syllabus: the reflex to ask how a number was computed, what is in the denominator, what got excluded, and who survived to be counted. Much of AI-era engineering is deciding whether a measurement means what the person presenting it believes it means. The practical floor is being able to write the query yourself instead of asking someone for a number you then cannot interrogate.

Why "prompt engineer" is a weak thing to become

The technique is real and worth being good at. The career is a bad bet, for four reasons that have nothing to do with whether prompting works.

It sits on the least stable layer in the stack: phrasings and tricks that mattered for one model generation stop mattering for the next, so the knowledge depreciates faster than you can accumulate it. It has almost no depth to grow into — a capable engineer gets most of the available skill in a few weeks, and anything learnable in weeks cannot sustain a wage premium once everyone has learned it. The tooling absorbs it: what was a carefully tuned prompt becomes a library default, a fine-tune, or simply behaviour the next model has by default. And it comes with no ownership surface: no service to run, no metric to be accountable for, no number that is yours. Skills that pay well are ones where being much better than average produces much more value. Prompt phrasing plateaus early.

What survives from that work is the part that was never really phrasing. Deciding which context to retrieve and how to keep it fresh. Decomposing a task into steps a system can check between. Designing the evaluation that tells you whether a change helped or just felt better. Building the fallback for when the output is nonsense, and the guardrail for when it is confidently wrong. Those are engineering skills that happen to be pointed at a model, and they belong to categories that were already valuable. Describe yourself as someone who builds and evaluates systems with a probabilistic component in them. That is both more accurate and more durable than the job title.

Own outcomes, not artefacts

Artefacts are pull requests, design documents, services, dashboards. Outcomes are the checkout error rate, the days it takes to onboard a merchant, the cost per thousand requests, whether the migration finished. When producing artefacts gets cheap, being the person who produces artefacts gets cheap with it. Being the person who can be handed a bad number and trusted to turn it into a good number does not, because that job contains all three of the resistant categories at once.

In practice this means attaching yourself to a metric, knowing its current value without looking it up, and being able to say what you changed and what happened — including the times it did not work. It also means getting comfortable with the parts of the job that are not coding: deciding what not to build, telling someone their request is underdetermined, and writing the two paragraphs that get four people to agree on a definition.

The complement to owning outcomes is having depth somewhere. General-purpose tools are strongest exactly where you are broad and shallow, which is most of your surface area. What makes you hard to substitute is being the person in the building who genuinely understands one thing: the payment rails, the video pipeline, the tax rules, the training infrastructure, the query planner. Depth is also what makes your review trustworthy, and review is the bottleneck. Pick something with a long half-life when you choose — the specifics of one vendor's API will be stale in two years, while how consensus works, how a planner chooses a join order, or how cost scales with context length will still be true.

How to tell whether you are actually getting stronger

Career advice is easy to nod along with and hard to check. These signals are checkable, and they lag by months rather than days, which is the correct timescale for this kind of change.

  • You catch things in review that the tests do not. If all your review comments are about naming and style, you are proofreading, not reviewing.
  • You are the person the ambiguous ticket gets brought to, because someone expects you to come back with a definition rather than a question.
  • You can reject a generated solution and say why in one sentence that a colleague immediately agrees with.
  • The systems you understand are getting bigger, not just the tools you have tried. "I have used six agent frameworks" is not progress; "I know what our retry policy does during a partial outage" is.
  • Your estimates are getting less wrong, which is the cleanest available proxy for whether your model of the system matches the system.

The warning signs are the mirror image. You are merging changes you could not explain if challenged. Your account of your own value depends on a tool name. You have not debugged something end to end in months. You cannot remember the last time you were the person who decided what "done" meant for a piece of work.

None of this rests on a prediction about how capable the models get, which is the useful thing about framing it as task properties instead of forecasts. The bet is the same either way: the closer your week sits to well-specified, verifiable, low-context work, the more of it moves away from you, and the more your value lives in judgement, verification, and depth, the more these tools make you faster rather than cheaper. Leverage multiplies whatever you are already able to evaluate. Spend the next six months getting better at evaluating.