In April 2024 I wrote about why code reviews break down. On a long-running headless multisite project, reviews had become the bottleneck for the entire team. We fixed it by agreeing what a review was for, starting with a hierarchy of what reviewers should check. We also asked authors to review their own work first, keep PRs small and explain the what and the why in the description. After that, open PRs settled at between 15 and 30.
Everything in that post assumed a person had written the code. Most of the code I ship now is written by an agent. I plan the work, set the constraints and check the result, but I don’t type most of it. On a recent build, those agent-written PRs often waited three or four days for review, sometimes a week. So I’ve started asking a question the old post never had to answer: what is peer review for, when nobody on the team wrote the code?
My answer is still provisional. This post is me thinking it through, and I’d like to hear where you think I’ve got it wrong.
The hierarchy, one layer at a time
The hierarchy had five layers. For each one, the question now is who or what does the checking.
Readable. This layer asked whether code is easy to understand, well named and consistent. Linters and coding standards now enforce most of the consistency. Given enough context and constraints, agents follow a codebase’s existing patterns and can explain any part of it on request. I’m less sure about readability across the years of a long agency project, as engineers rotate on and off. I haven’t yet seen whether a high-level understanding, plus an agent to explain the detail, holds up over five years.
Robust. This layer asked whether the code does what it should, handles edge cases and has tests. Unit and integration tests now catch regressions, and QA tests the feature against the specification. Performance is the gap in this layer, and I’ll come back to it.
Secure. Static and taint analysis catch the known patterns, such as missing escaping, unsanitised input, unsafe queries and missing capability checks. They can’t tell you whether the authorisation logic is correct. An adversarial review by an agent is useful for that kind of question. I’d use models from more than one vendor, because different models tend to miss different things.
Elegant and inspiring. The top two layers asked whether you’d be proud of the code and whether it left the codebase better. Those layers were partly about teaching, because junior engineers learned what good code looked like from review comments. When an agent types the code, I’m not sure who that feedback teaches. It can teach the agent if someone turns the lesson into a skill or a style guide, but it doesn’t obviously teach the engineer. I come back to this when I look at where the approach breaks.
The purposes nobody writes down
Code review also has purposes that no hierarchy lists. It spreads knowledge of the codebase across the team. Some clients rely on it as evidence that a second person approved every change. I also suspect it shares responsibility, because when two people approve a change, neither of them owns a mistake alone.
With agent-written code, shared responsibility is the purpose that worries me. My old post said your reviewer was probably not involved in the fix, and that the PR description should close that gap. With agents, the gap is wider. The author may not have read every line of a diff several thousand lines long. But they planned the change, set its constraints, steered the agent and read the tests. The reviewer arrives with the same lines and none of that decision context, so the author is still the person best placed to judge the change.
Ownership stays with the author when an agent types the code. If I ship it, I own it, and I’m expected to understand what it does.
Two builds
I’ve worked on two recent builds that handled review very differently. They had different teams and different clients, so they aren’t a controlled comparison.
The first build kept the workflow we’d used when we hand-coded everything. An engineer finished a ticket and it went to code review. After review, it went to staging for testing, then on to UAT if it passed or back to the engineer if it failed. I used agents for the whole build. The agent work moved quickly, and then PRs often waited three or four days for approval, sometimes a week. Because testing on staging only started after review, each of those PRs spent days waiting before anyone learned whether the feature worked.
The second build was me and one front-end engineer, so there wasn’t really a peer to review the back end. We sometimes asked each other for a quick look, but around 90% of the back end merged without peer review. We built each feature in a branch and merged it once linting and tests passed and QA had signed it off. Delivery was noticeably faster than on builds with a formal review step, and no work sat waiting for a reviewer.
I didn’t choose that setup to test a theory. The team was small and a reviewer wasn’t available. The interesting part is the constraint: with nobody else to review the code, I had to decide what I actually check.
What I check instead
When I review my own agent-written PR, I start with scope. I look at how much of the codebase the agent touched and compare that with the size of the feature, because a small feature that changes dozens of files needs an explanation. Then I look through the diff for anything absurd. I run a Copilot review, or something similar, as a second opinion. Last, I read the tests and check that they test what the feature needs.
My check is deliberately shallow on correctness. I ask whether the change looks like it does what it should, and I leave line-by-line correctness to the tests and QA. Each of my checks depends on knowing what the feature was meant to do, and the author knows that better than anyone else on the team.
Leaving correctness to the tests has a weakness. If the agent misreads the requirement, it writes the code and the tests from the same misreading, and the tests pass while confirming the wrong behaviour. Code and tests from the same interpretation only count as one piece of evidence. So I read the tests against what the feature needs, and I don’t treat a green test run as proof. QA is stronger evidence, because the tester works from the specification and never saw the agent’s interpretation. An adversarial agent given only the specification might add another independent check.
The approach has let things through. Performance problems have passed my self-review more than once. QA found them, and we either fixed them before go-live or scheduled a follow-up task soon after launch.
Where this breaks
Those follow-up fixes make the trade look easier than it is. One authorisation bug that exposes personal data costs more than a hundred slow queries. The fair comparison is days of waiting on every PR against the chance that a peer would have caught a serious defect, weighted by how bad that defect would be. I think that chance is small for most changes, but I haven’t measured it, and I suggest a way to check at the end. For some changes, the chance is clearly large.
Some failures look like success. QA catches a broken payment form, but it won’t catch a working payment form with an exploitable script path, because the form still takes payments. Authorisation logic that shows data to the wrong user can pass QA in the same way, and so can a migration that completes but writes some records incorrectly. For payments, authentication and authorisation, personal data and migrations, I still want a human reviewer. Some clients also require evidence that a second person approved each change, and a team would need to check its contracts before changing its process.
Keeping review for critical flows only works if someone notices when a change reaches one. In a routine PR, an agent can change authorisation code as a side effect, and an author skimming the diff can miss it. The scope check is how I catch this, because an unexpected file in the authentication code is a reason to stop and read properly. A team could also enforce it, for example with CODEOWNERS rules that require a reviewer for changes in those directories. My old post suggested a PR section called “What I want reviewers to focus on”. On critical flows, that section lets the author point the reviewer straight at the risky part.
Performance has a similar limit. QA caught my performance problems because they showed up during testing, but some problems only appear under production load. Tests can check part of this, for example by asserting how many queries a request makes, as long as someone decides to write them.
An agent can also confidently solve a problem next to the one you asked about. Everything passes, including the tests the agent wrote from the same misunderstanding, and the code is good code for the wrong feature. A peer reading the diff rarely catches this, because the diff looks fine. It gets caught earlier, when people agree the specification. On client projects, discovery already produces that specification, and I think human scrutiny there now does more good than a second reader on the diff.
The last failure is slower to appear. My checks depend on experience. I can tell when a change touches too much because I’ve spent years reviewing other people’s code. Engineers build that judgement by writing code, reviewing it, being reviewed and watching their own changes fail. Agents take away many of those repetitions, and supervising an agent needs exactly that judgement. I don’t know where a junior engineer gets it if they rarely write, review or get reviewed, and of all the questions in this post, it’s the one I find hardest.
Where I’d put the checks
On most projects I work on, every change still goes to peer review, including changes no person typed and the author only skimmed. My old hierarchy started from what a reviewer checks. I now find it more useful to start from the failure, then decide where a check has the best chance of catching it. For most agent-written changes, that place is somewhere other than a second person reading the diff. This is how I’d place the checks today:
| What we’re trying to catch | Where the check lives now | Human peer review? |
|---|---|---|
| Style and consistency | Linters and coding standards | Rarely |
| Regressions and functional bugs | Unit and integration tests, QA against the specification | Risk-based |
| Known security patterns | Static and taint analysis | Rarely |
| Wrong authorisation logic or exposed personal data | Adversarial agent review, deliberate tests | Yes |
| Broken payments or migrations | Tests, QA, and a reviewer told where to look | Yes |
| Performance under production load | Deliberate tests, such as query-count assertions | Sometimes |
| Solving the wrong problem | Agreeing the specification | At the specification stage |
| Spreading knowledge of the codebase | Unclear | Unresolved |
| Building engineers’ judgement | Unclear | Unresolved |
The last two rows have no answer yet. Most of the other rows rest on my guess that peer reviewers rarely catch serious defects in routine agent-written changes. You can test that guess on your own team. Take the comments on your last few agent-written PRs and sort them into style, real bugs, design decisions and missing context. Tools can handle most style comments. If the real bugs cluster in particular kinds of change, those are the changes that still need a second reader.
So, for the changes your team reviews today, what is the review for?