The rules are written down. The agent reads them on every prompt. They still do not make it into the code, and nothing is broken enough for the bot to report. A rules file is not configuration. It is a suggestion with an unknown success rate.

Our commits carry a co-authored-by line with the model's name. Nobody hides it. Everyone on the team knows which parts came out of a model and which parts someone wrote by hand.
Which should settle things. I know what I'm reviewing. There's a review bot. There's a project rules file the agent reads on every prompt. Configuration exists, rules exist, automation exists.
And the best part is that it all passes. I open the merge request and the pipeline is green.
Then I start reading the code, and not one of those rules made it into the diff.
Unit tests. A wall of describe blocks, one after another, testing the same
behaviour with a different input each time. The rules file says it plainly: use
it.each. Not for runtime. The suite runs at the same speed either way. It's so
that adding a case is one row in a table instead of another copied block that
somebody has to read end to end in six months just to work out it's a variant of
the one above it.
Iterating over data. I work on views where the datasets run into hundreds of thousands of records, with several tables and charts alive on the same screen. Every iteration has to earn its place. Complexity and performance aren't textbook topics here. They're the difference between a UI that glides and one that stutters.
What comes back is an iteration chain with a lookup buried inside it. Roughly this:
// Example written for this post.
const enriched = orders.map(order => ({
...order,
customerName:
customers.find(customer => customer.id === order.customerId)?.name ?? '',
}));It's correct. It passes the tests, it passes the types, it passes review. And every one of us has written it by hand at some point, because it's the first thing that comes to mind when you need to join two lists.
The trouble is that a find inside a map is a linear scan inside a loop.
O(n·m). Two collections of a thousand items each turn a thousand comparisons
into a million. At the sizes I work with, that isn't a difference you need to
measure to notice.
Build the index once:
const customerNames = new Map(
customers.map(customer => [customer.id, customer.name]),
);
const enriched = orders.map(order => ({
...order,
customerName: customerNames.get(order.customerId) ?? '',
}));Same logic, one line more, a different complexity class. No trick, no secret knowledge. Just a decision somebody has to make deliberately, knowing the scale of the data. And the scale of the data isn't visible in the diff.
Nested ternaries. The rules say don't. They come back anyway. They work, they're compact, they even look tidy, but past three levels nobody reads the condition any more. They guess it.
Every time, the code passes. Tests green, types fine, linter quiet, review bot with nothing to report. Nothing is broken and nobody violated a rule in the sense a tool understands the word.
Look at what the three have in common. A describe block is the most commonly
written shape of a test. A lookup inside a loop is the most natural thing to
reach for when joining two lists. A nested ternary is the most common way a
condition gets written in JSX.
The model didn't choose badly. It chose what occurs most often, because that is
precisely what it is. it.each, a Map instead of a scan, a condition broken
out into branches, all of them need something beyond correctness: you have to
know the alternative exists and decide it fits here. You have to know the scale
of the data, which the diff doesn't show, and know that someone will be reading
this file a year from now.
Which is where I had it wrong in my own head. I put those rules in a configuration file and started treating them as configuration, the kind of thing that either applies or fails loudly. It isn't a contract. It's a weight on a distribution. It raises the odds I get what I asked for, and that's all it does. A prohibition in that file isn't a prohibition. It's a suggestion with a success rate nobody will ever quote you.
That changes what review is for, more than I'd appreciated. If a rule is a guarantee, review is there to catch the exceptions. If a rule is a suggestion, review is the only place it's ever actually enforced.
I run the full review bot again, by hand. If it finds anything, the report goes into the merge request. And when the merge request is genuinely large, I read it line by line after the bot has been through it.
A green bot stopped being permission to merge. It became the precondition for reading. The bot tells me there are no violations, not that the code is good, and on a large merge request that's where I start reading, not where I stop.
I won't pretend this is an elegant process. It's babysitting the model, and I do catch myself wondering whether that time should go somewhere else. There is always more to do than there is time for, and here I am reading code a machine produced in thirty seconds.
For a long time I didn't have a good answer. I had a polite one about ownership of quality, and I didn't entirely believe it myself.
The answer arrived from somewhere else. AI tooling has downtime. And instead of the obvious response, fine, it's down, I'll sit and write it myself because I know how, something else happens. Panic, or a full stop. It's been down half an hour, I'll wait for it to come back.
Half an hour. Not two days, not a week. Half an hour was enough.
I caught myself doing exactly that, and it wasn't a pleasant thing to notice.
Because that's the bill. Taking the most common option is free for as long as the tool is up. The code passes, the feature ships, nobody pays. The ability to choose something else isn't needed, so it stops being used. Downtime just reveals that it's already gone.
So no, I don't think it's over-engineering any more. Reading that code line by line is currently the only time that muscle does any work at all.
Was this helpful?