- MantraMindAI
- Blog
- AI Applications & Strategy
Coding with an AI assistant: where it helps, where it costs you
Jai Rao
August 22, 202621 min read
A working guide to AI coding assistants: the jobs they genuinely accelerate, the failure modes and their mechanisms, and the review discipline that keeps quality.
Two hours of work collapsed into forty seconds of typing, and the tests passed. That is the experience that sells coding assistants, and it is a real experience — the tool genuinely wrote the adapter, the fixtures and the twelve test cases you were dreading. What the first week does not show you is the other shape this takes: a pull request that read cleanly, looked idiomatic, passed CI, and quietly dropped a handful of records on every batch for three weeks before a customer noticed.
Those are the same tool doing the same thing. It emits the most plausible continuation of the code it can see. That one sentence predicts nearly everything about where these assistants pay for themselves and where they take time back out of your week, which makes it worth holding onto instead of treating each surprise as a one-off. What follows is about working practice: which jobs to hand over, which to keep, and the review discipline that decides which of those two stories you get.
The jobs where it earns its keep
Sort the work in front of you by a single question: does writing this correctly require knowledge that is not visible to the model? Where the answer is no, the tool is genuinely excellent, and the gains are not small.
Code you could write but would rather not. The repository class with six near-identical methods. The argument parser. The mapper between a vendor's JSON and your internal type. A migration that adds four columns and backfills them. In all of these, the correct output is fully determined by things already on screen — the schema, the interface, the surrounding file — and the only real cost is keystrokes. Keystrokes are exactly what the tool removes.
Format conversions. A curl command into a typed client call. A JSON fixture into a set of dataclasses. A CSV header row into a table definition. A Postman collection into integration tests. These are mechanical transformations with a high tedium-to-thought ratio, and a mistake usually announces itself immediately.
Scaffolding against an API you have not used before. You know what you want; you do not know how this library spells it. A generated first draft gets you to something that runs, which is a much better place to read the documentation from than an empty file. Treat the draft as a lead, not an answer.
Tests from a specification you wrote. This one is real but carries a large asterisk, big enough that it gets its own section. The order matters enormously: tests derived from a spec you own are useful, tests derived from an implementation the model just wrote are close to worthless.
Explaining code that predates you
This is the most underrated use and the one with the best ratio of value to risk. Paste a three-hundred-line function with no comments and ask what it does, what each branch is for, and which parameters are load-bearing. The model is good at this because the answer is present in the text it was given — summarisation rather than invention. You also get vocabulary: the names of the concepts the original author had in mind, which makes the rest of the codebase searchable.
It is still a hypothesis. Confirm the two or three claims your change depends on by reading the code yourself. But going from "no idea" to "a specific claim I can check in five minutes" is a large step, and the tool takes it reliably.
Refactors that are typing, not thinking
Renaming a concept across sixty files. Converting callbacks to promises. Threading a new argument through a call chain. Migrating an assertion style. Splitting one over-loaded parameter into two. These changes are repetitive, and repetition is where humans get careless — you are sharp at file three and sloppy at file forty-one. The tool has no file forty-one.
Two preconditions make this safe: you can state the transformation precisely enough to recognise a violation, and there is a mechanical check at the end — a compiler, a test suite, a grep whose count you can predict. If neither holds, this is not a mechanical refactor. It is design, and you should be at the wheel.
Invariants that exist nowhere in the context window
Your system is held together by rules that are true but unwritten. The user id in this service is always the tenant-scoped id, never the global one. This function must not be called inside the outer transaction, because the handler retries and the retry would double-charge. Timestamps in events are UTC, but the ones in legacy_events are local, because a migration in 2019 was never finished. The queue is at-least-once, so every consumer has to be idempotent. Nobody wrote any of that down where you are working.
The mechanism is worth stating precisely, because it is not "the model is careless". The model produces code consistent with everything it can see, and when a constraint is absent from the context there is no signal distinguishing "this is fine" from "this is fine given rules I was never shown". Absence reads exactly like permission. A human joining your team has the same problem in week one, with one difference: the human is visibly uncertain and asks. The generated code arrives with no uncertainty attached.
Two families of bug follow from this more often than any other. The first is concurrency. Nothing in a single file says that two copies of this handler can run at the same instant, so you get read-then-write code like this, which is correct in a test and wrong in production.
def reserve_seat(conn, event_id): row = conn.execute( "SELECT remaining FROM events WHERE id = ?", (event_id,) ).fetchone() if row is None: raise EventNotFound(event_id) if row[0] <= 0: raise SoldOut(event_id) conn.execute( "UPDATE events SET remaining = remaining - 1 WHERE id = ?", (event_id,) ) conn.commit()Read it as an interleaving: two requests both read remaining = 1, both pass the check, both decrement, and you have sold a seat that does not exist. The check and the write are not atomic, and no amount of staring at the function reveals that, because the missing information — that this runs concurrently — was never in the room. The fix moves the condition into the write and reads the result.
def reserve_seat(conn, event_id): cur = conn.execute( "UPDATE events SET remaining = remaining - 1 " "WHERE id = ? AND remaining > 0", (event_id,) ) conn.commit() if cur.rowcount == 0: raise SoldOut(event_id) # or EventNotFound; distinguish with a second readOne statement, one round trip, and the database enforces the invariant instead of the application hoping for it. Note that the fix is not cleverer code; it is the same code plus a fact about the deployment.
The second family is lifetimes and ownership: a resource released while something still holds it. A generator returned from inside a with open(...) block, so the file is closed before anything iterates it. A connection handed to a background task that outlives the request. A closure capturing a loop variable by reference. These are hard to see for the same reason — what makes them wrong sits outside the fragment being written.
Confidently wrong about APIs, quietly out of date about patterns
A second class of weakness comes from the model's own knowledge rather than from your system. Three versions of it, in increasing order of how long they survive.
Methods that do not exist. You get client.batch_update(items, retry=True) — plausible name, plausible signature, consistent with everything else in that library's naming style, and completely absent from the library. Plausibility and existence are different properties, and only one of them is being optimised. The good news is that this fails loudly on the first run, so it costs seconds. The expensive variant is subtler: a parameter that exists but means something else, a method that exists on a sibling class, an option whose default flipped between major versions.
Patterns that were correct three years ago. Training data is a snapshot, and it is weighted by volume rather than recency — there is far more old code in the world than new code. So you get the deprecated auth helper, the pre-async idiom, the ORM call removed two majors ago, a pinned version with a known advisory against it, datetime.utcnow() in a codebase that moved to timezone-aware objects. This code compiles, often works, and usually passes review, which is precisely why it accumulates.
Code that is idiomatic for the language and wrong for your repository. It reaches for the standard logger while everything around it uses a structured one carrying request context. It raises a bare ValueError while your service has an error taxonomy mapped to status codes. It reads an environment variable directly while every other module goes through a settings object. Each passes review by anyone who does not already know the convention, and each is a small permanent tax on the next person. The mechanism is the same as before: a convention is only visible if a file carrying it is in the context.
Tests that agree with the code they were written against
If you take one thing from this, take this one. A test suite is worth exactly what its independence from the implementation is worth. When the same model writes the implementation and then the tests, both come from the same understanding of the problem — including the same misunderstanding. The suite then confirms that the code does what the code does. It goes green, you feel covered, and no information about correctness was produced.
Here is a function of the kind you would accept without a second look. The task was to split a list of ids into batches for an endpoint that accepts at most fifty at a time.
def chunk(ids, size): """Split ids into lists of at most size items for the bulk endpoint.""" if size <= 0: raise ValueError("size must be positive") chunks = [] for i in range(len(ids) // size): chunks.append(ids[i * size:(i + 1) * size]) return chunksIt validates its argument, it has a docstring, it handles the empty list correctly, and the loop looks like every batching loop you have ever read. It also silently discards the remainder: len(ids) // size counts only the whole batches, so with 101 ids and a size of 50 you get two batches and one id vanishes. No exception, no log line, no failing request — just a record that never got written.
Now the test that shipped alongside it, which passes.
def test_chunk_splits_into_batches(): assert chunk([1, 2, 3, 4], 2) == [[1, 2], [3, 4]] assert chunk([1, 2, 3, 4, 5, 6], 3) == [[1, 2, 3], [4, 5, 6]] assert chunk([], 2) == []Look at the inputs. Four items in batches of two, six in batches of three, and the empty list. Every case divides evenly, so the remainder path is never exercised. The tests are not dishonest; they were written from the shape of the implementation, and the implementation's shape has a blind spot, so the tests have the same blind spot. That is the circularity, and it is not fixed by asking for more tests — more tests from the same source have the same hole.
What breaks the circle is deriving the assertion from the requirement instead of the code. The requirement here is not "produce batches"; it is "every id ends up in exactly one batch, order preserved, no batch larger than the limit". Write that.
import pytest@pytest.mark.parametrize("n", [0, 1, 2, 3, 5, 7, 100, 101])@pytest.mark.parametrize("size", [1, 3, 50])def test_chunk_loses_nothing(n, size): ids = list(range(n)) out = chunk(ids, size) assert [i for c in out for i in c] == ids # nothing dropped, order kept assert all(0 < len(c) <= size for c in out) # no empty or oversized batchThis fails immediately on n=7, size=3 and on n=101, size=50, and it would have failed on the day the function was written. The fix is one line — iterate the input rather than a computed batch count — and the parametrised boundaries are what turn it from a bug someone finds in production into a red test.
for i in range(0, len(ids), size): chunks.append(ids[i:i + size])The practical rule that follows: you own the assertions, the tool can own the plumbing. Fixtures, parametrisation tables, mock setup, the tedious arrange-and-teardown scaffolding — hand all of that over. But the sentence that says what must be true is the whole value of the test, and it has to come from your understanding of the requirement. If you do want generated tests, give the model the specification and not the implementation, then read what it asserts rather than whether it is green.
Three specific tells that a generated test is not testing anything. It asserts on a mock, so it passes because the mock returned what the mock was configured to return. It wraps the call in a broad except and asserts no exception escaped. Or it asserts only that the result is not None — a condition almost no bug in the world violates. When you see these, the test is decoration; delete it or replace the assertion.
Read the diff, not the vibe
Reviewing generated code is a different skill from reviewing a colleague's code, and the difference trips up experienced reviewers specifically. Human code broadcasts its own uncertainty: the naming gets awkward where the author was unsure, the style wobbles at the hard part, a comment admits "I think this handles the retry case". You have spent years using that signal to decide where to slow down, and generated code does not emit it. Worse, the signal inverts. The passages that read most fluently are where the model was most firmly in the groove of the common pattern, which is exactly where an uncommon requirement gets flattened.
So invert the habit deliberately. Be most suspicious of the code that looks most finished. Concretely, stop at every index expression, every range, every len(), every comparison that could be off by one, every error branch, every place a lock or a transaction is implied, and every default value. For each hunk, answer one question out loud: what input makes this wrong? If you cannot produce a candidate, you have not read it — you have skimmed it and recognised the shape.
Read the diff itself, hunk by hunk, not the summary the tool wrote about it. That summary comes from the same context and inherits the same blind spots. Then keep the unit small: the cost of producing a nine-hundred-line diff has collapsed while the cost of reviewing one has not moved at all, and that asymmetry is new. A change too big to review is too big to accept, and it is now trivially easy to be handed one.
The floor rule: never merge code you cannot explain. If you cannot say in one sentence why each non-obvious line is there, you have two options — understand it or write it yourself. There is no third option where you merge it and hope, because in six months you are the person debugging it at three in the morning with no memory of ever having read it.
Different failures are caught by different checks, and it is worth knowing which is which rather than trusting review in general.
| Failure | Why it gets through | What actually catches it |
|---|---|---|
| Violated system invariant | Nothing in context contradicted it | A test written from the invariant; a reviewer who knows the system |
| Off-by-one, dropped remainder, empty case | Happy-path inputs pass | Boundaries: 0, 1, exact multiple, one over, duplicates |
| Race or lifetime bug | Single-threaded tests pass every time | Reasoning about interleaving; conditional writes; concurrent stress test |
| Local convention violation | Reads as idiomatic | A sibling file in context; a linter; a familiar reviewer |
| Method that does not exist | Signature looks plausible | Running it, once, immediately |
| Outdated pattern | Works today | Deprecation warnings; current docs; version pins in review |
| Injection or unsafe deserialisation | Functionally correct | Security lint; a hard rule that all queries are parameterised |
| Non-existent or abandoned dependency | Name sounds real | Checking the registry before it enters the lockfile |
Point it at the right files and say the constraints
The single practical skill worth building is not phrasing. Rewording a request buys you very little; changing what the model can see changes the answer completely. The same question asked with the interface definition and one sibling implementation in view produces work of a different quality from the same question asked cold, and no amount of clever wording closes that gap.
What belongs in the context is short and specific: the file being changed, the interface or type the result must satisfy, the test file, the real error output with its stack trace, the relevant schema. And above all, one existing implementation of the pattern you want followed. A sibling file encodes your conventions — logging, error taxonomy, transaction handling, naming — far more effectively than any description of them, because it shows the convention instead of asserting it.
Everything else does not belong. More context is not better: irrelevant files dilute attention and actively mislead, because a deprecated helper left in view looks like current practice from inside the context window. Stale copies are worse — long sessions accumulate versions of code that no longer exists, and the model keeps confidently editing the file as it was twenty messages ago. Re-read the file, or start clean.
Then say the invariants, because typing them is the only way they enter the context at all. This handler is retried, so it must be idempotent. This runs in four processes. This table has forty million rows, so no full scan. This project targets an older runtime, so that syntax is unavailable. Callers hold the lock, so do not take it again. Every one of those is a bug you did not get.
A useful test before you delegate: could you write down what a competent contractor would need to know to get this right on the first attempt? That list is your context. If you cannot produce the list, you do not yet understand the change well enough to hand it over — which is worth knowing before you have a diff to review rather than after.
Four hazards that outlive the pull request
These deserve separate treatment because the code involved is functionally correct. Tests pass, review nods, and the problem surfaces somewhere other than your test suite.
Injection and unsafe deserialisation. Data-access code assembled by string concatenation is enormously abundant in the material these models learned from — tutorials, scripts, forum answers — so it is a highly plausible continuation whenever the surrounding file does not demonstrate the parameterised alternative.
def find_orders(conn, customer_name, status): sql = ( "SELECT * FROM orders WHERE customer_name = '" + customer_name + "' AND status = '" + status + "'" ) return conn.execute(sql).fetchall()A customer name containing a quote breaks it; a customer name chosen by an attacker owns your database. The fix is boring and absolute: pass values as parameters, never as string fragments, no exceptions for internal-only fields. The same shape recurs as yaml.load instead of safe_load, pickle.loads on anything from the network, eval on a config value, and a subprocess call with a shell and interpolated arguments. Nothing in the file being edited says the input is attacker-controlled, so nothing pushes back.
Dependencies that do not exist, or should not be used. A suggested package name can be perfectly plausible and simply unpublished — and because assistants suggest the same plausible names repeatedly, someone may publish it later precisely to catch whoever installs it. Others exist but were abandoned years ago, or carry a licence your product cannot ship. No dependency enters the lockfile on a suggestion: look it up, and check who maintains it, when it last released, its licence, and how many transitive packages it drags in.
Verbatim reproduction of licence-encumbered code. The risk concentrates on distinctive, widely copied code — a specific algorithm implementation, a well-known utility — rather than on boilerplate, and it arrives with no attribution and no indication that anything was copied. If a generated block looks like a known implementation you could name, that is the signal. Write it yourself, or take it from a source you can cite and comply with.
Secrets in prompts. The config file pasted whole. The stack trace with a token in a URL. The .env shared to explain a startup failure. A customer record pasted to help debug a serialisation error. Anything that leaves the machine should be treated as disclosed: rotate the credential rather than reasoning about whether it was probably fine. Keep a scratch file of placeholder values, redact before pasting, and find out what your organisation's retention and training settings actually are instead of assuming the comfortable answer.
Why it feels faster than it measures
The feeling of speed is not imaginary, but it is measuring the wrong thing. What collapses is the visible, effortful part of the job: the empty file, the remembered syntax, the eighty lines you did not want to type. That is the part your body registers as work. What grows is reading, checking, and holding a hypothesis about code you did not write — which feels lighter while costing more. The time did not disappear. Writing went down and reviewing went up, and if reviewing went up further than writing went down, you got slower while feeling faster.
Which direction it lands in is predictable. On boilerplate in a familiar codebase, writing drops sharply and review is cheap, so the win is real. On unfamiliar code with dense invariants and real consequences, review dominates and you can end up behind. Sorting the work before you ask matters more than any technique.
Attribution makes it worse. The forty-second win is vivid and clearly credited to the tool. The ninety minutes lost to a subtly wrong abstraction gets filed under "debugging", which is where that time always went, so it never gets charged to the decision that caused it. Nobody is being dishonest; the accounting is just asymmetric by default.
Second-order costs compound quietly. Diffs get bigger, so more code exists, and every line will be read and paged in by someone later. Familiarity with your own codebase drops, because you cannot navigate from memory a module you skimmed once. And review load drifts onto whoever is most conscientious, which is a durable way to lose that person's throughput.
If review is the bottleneck, treat it as the constraint. Keep the unit of generated work small enough to review in one sitting, and finish one piece before asking for the next. Spend the saved time on the checks that are hard to fake — invariant tests, boundary cases, an argument about concurrent access. Front-load constraints so review has less to catch. And apply the honest rule: if reviewing it carefully costs more than writing it, write it. Then measure something other than volume — escaped defects, reverts, how often a generated change gets redone, or just a weekly look at which of your last ten uses actually saved time.
Habits that hold up at the keyboard
Everything above compresses into a handful of things you can do today.
- Classify before asking: is this code you could write but would rather not, or code whose correctness depends on what only you know? Hand over the first kind without hesitation.
- Put one sibling implementation in the context, every time. Conventions travel by example, not by description.
- State the invariants in the request. If you cannot list them, you are not ready to delegate the change yet.
- Keep the unit small enough to review in a single sitting, and review it before asking for more.
- Write the assertions yourself; let the tool write fixtures, parametrisation and mock setup.
- Test boundaries on reflex: empty, one, exact multiple, one over, duplicates, unsorted, missing.
- Reread the passage that looks most polished. That is where the inverted signal lives.
- Run it early. An invented method dies in two seconds on the first execution and can survive an hour of reading.
- No dependency enters the lockfile on a suggestion.
- Redact before pasting; rotate anything that got away.
- Never merge a line you cannot explain.
The pattern underneath all of them is the same. These tools are unusually good at the part of programming that is typing, and unusually bad at the part that is knowing things about your particular system. Most of the craft is keeping that boundary in view while you work — and noticing, when a session starts to feel effortless, that effortlessness is the condition under which the boundary is easiest to forget.