Skip to main content
How AI Is Changing Code Review — Anselm Fowel
AI & Technology

How AI Is Changing Code Review

10 min read
856 views
Share:

Code review has been one of the quiet constants of software engineering for most of my career. The mechanics changed over the years, from emailed patches to pull requests, but the core ritual stayed the same: a human reads another human's diff, asks questions, and decides whether it is safe to merge. In the regulated fintech systems my teams build, that ritual is not optional. It is part of how we demonstrate to auditors and to ourselves that money moves correctly.

How AI Is Changing Code Review
How AI Is Changing Code Review

Over the last two years, AI has started to reshape that ritual in ways that are more substantial than the autocomplete demos suggest. I want to be precise about what is actually changing, what is hype, and how a CTO who is accountable for production payment systems should think about adopting these tools without lowering the bar.

What Code Review Was Always Protecting

Before talking about AI, it helps to remember what review is for. In my experience the value of a pull request review is rarely the catching of obvious bugs. Compilers, type checkers, linters, and tests catch most of those before a human ever looks. The real value is harder to automate: confirming that the change matches intent, that it fits the existing architecture, that the author understood the edge cases, and that the next engineer to touch this code will be able to follow it.

In payments work there is an additional layer. A reviewer is implicitly asking whether the change preserves invariants that no test fully expresses. Does a refund still reconcile against the original authorization? Does this retry path risk a double settlement? Are we logging something we should not be logging under our data handling obligations? These questions require context that lives in people's heads and in compliance documents, not only in the diff.

I keep this in mind because it tells me where AI can genuinely help and where it cannot. AI is good at the broad, mechanical, repetitive layer of scrutiny. It is far weaker at the contextual judgment that made senior review valuable in the first place. The tools that disappoint are the ones sold as a replacement for judgment. The tools that earn their place augment the mechanical layer so humans can spend their attention on judgment.

From Linters to Language Models

The first wave of automated review was static analysis: linters, style checkers, security scanners. They were rule based and deterministic. If you wrote a SQL string with concatenated user input, a scanner flagged it. These tools were valuable but brittle. They produced enormous false positive volumes and they only knew the rules someone had written down in advance.

Language models change the shape of this because they reason over the diff in natural language rather than matching patterns. An AI reviewer can read a function, infer its intent from the name and the surrounding code, and notice that the null check three lines up does not actually protect the dereference below it. That is a qualitatively different capability than a regex. It can also explain its reasoning, which matters enormously for whether an engineer trusts and acts on the feedback.

The shift is from tools that check whether code matches a rule to tools that attempt to understand what the code is trying to do. That is more powerful and more dangerous at the same time, because the failure mode is now confident plausibility rather than a noisy false positive.

Where AI Review Genuinely Helps

I have watched AI review tools earn their keep in specific, bounded ways across my teams. The pattern is consistent: they are strongest where the problem is local to the diff and does not require deep system context. When I look at the categories of feedback that actually changed code before merge, a few stand out.

  • Catching obvious defensive gaps, such as an unhandled null, an unclosed resource, or a swallowed exception that hides a failure.
  • Flagging inconsistent error handling, where one branch returns a typed result and another throws, in the same method.
  • Spotting missing test cases for boundary conditions the author clearly intended to cover but did not.
  • Surfacing copy paste mistakes, like a logging statement that references the wrong variable after a block was duplicated.
  • Improving the readability of the pull request description and summarizing large diffs so human reviewers orient faster.

That last point is underrated. A surprising amount of the friction in review is simply understanding what a large change is doing. A good AI summary of a forty file diff, grouped by concern, lets a senior engineer decide where to focus within seconds rather than minutes. The AI is not making the judgment. It is reducing the cost of the human making it.

A Concrete Example From a Settlement Service

Let me make this tangible. Here is a simplified piece of C# from a settlement service, the kind of method an AI reviewer recently commented on usefully. The author had added a retry around a downstream call and an AI reviewer pointed out that the retry would re-execute a side effecting operation, risking a duplicate ledger entry.

public async Task<SettlementResult> SettleAsync(SettlementRequest request)
{
    for (var attempt = 0; attempt < MaxAttempts; attempt++)
    {
        try
        {
            // Side effect: writes a ledger row before the network call
            await _ledger.RecordPendingAsync(request.Id, request.Amount);
            var response = await _gateway.PostAsync(request);
            return SettlementResult.FromGateway(response);
        }
        catch (TransientGatewayException) when (attempt < MaxAttempts - 1)
        {
            // Retries the whole block, including RecordPendingAsync
            continue;
        }
    }

    throw new SettlementFailedException(request.Id);
}

The AI flagged that the ledger write sits inside the retry loop, so a transient gateway failure would record a second pending row on the next attempt. The fix is to move the idempotent recording outside the loop, or to key the ledger write on a deduplication identifier so repeats collapse. None of this is exotic, but it is exactly the kind of subtle ordering issue a tired reviewer skims past at the end of a long day. The AI did not understand our reconciliation rules. It understood that a side effect was being repeated, and that was enough to start a valuable conversation between two engineers.

Enjoying this article?

Get more like it in your inbox — practical engineering leadership, fintech, and AI. No spam, unsubscribe anytime.

The False Confidence Problem

The risk that worries me most is not that AI review misses things. Humans miss things too. The risk is that AI review produces confident, fluent feedback that is subtly wrong, and that this fluency lends it unearned authority. An engineer who is uncertain will often defer to a tool that sounds certain. In a regulated environment that is a dangerous dynamic.

I have seen AI reviewers confidently suggest a change that would have introduced a race condition, recommend removing a check that looked redundant but was load bearing for a compliance requirement, and assert that a piece of code was untested when the tests lived in a file the model had not been shown. Each of these would have been caught by a careful human, but each also created friction and, occasionally, a near miss when a junior engineer accepted the suggestion without challenge.

The lesson my teams have internalized is that AI feedback is an input to review, never the output of it. We treat an AI comment exactly as we would treat a comment from a new hire who is sharp but lacks context: worth reading, frequently useful, and never authoritative on its own.

What This Does to Engineers and Their Skills

There is a longer term concern that I think about as a leader more than as a technologist. Review is how junior engineers learn. When a senior engineer explains why a change is risky, that explanation transfers judgment. If AI absorbs the mechanical layer of review, juniors get faster feedback, which is good. But if AI absorbs the explanatory layer too, we risk producing engineers who can satisfy a tool without ever building the mental models that the tool is standing in for.

I do not think this is inevitable, but it requires intent. We deliberately keep senior humans in the loop on architecturally significant changes precisely so that the teaching continues. We also encourage engineers to argue with the AI in the pull request thread, because the act of articulating why a suggestion is wrong is itself a form of learning. A reviewer who can explain why the model is mistaken understands the system better than one who simply accepts or rejects.

The teams that adopt these tools well treat them as a forcing function for clearer thinking, not as a way to think less. That distinction shows up in code quality within a quarter.

Governance, Provenance, and the Audit Trail

In fintech, how a decision was made matters as much as the decision itself. When an AI tool participates in code review, several governance questions follow immediately, and I expect any team in a regulated context to have answers before they roll the tool out broadly.

First, what is sent to the model. If your review tool transmits proprietary source code or, worse, code that embeds secrets or customer data, you need a contractual and technical guarantee about retention and training use. We restrict review tooling to vendors with clear data handling terms and we keep certain repositories off limits entirely. Second, provenance. If an AI suggested a change that later caused an incident, our post incident process needs to reconstruct who accepted it and why. We log AI participation in the pull request record so it is part of the audit trail, not an invisible influence.

Third, accountability does not move. A human approver still owns the merge. The AI is a tool the human used, in the same sense that a static analyzer is a tool. No auditor will accept the model made me do it, and neither will I. Making that explicit in policy prevents the quiet erosion of ownership that otherwise creeps in when tooling gets good enough to feel authoritative.

How We Actually Rolled It Out

We did not flip a switch and trust AI review across the organization. We piloted it on a single team working on internal, non sensitive services. We measured whether it caught real defects, how often its comments were acted on versus dismissed, and whether review cycle time changed. The early signal was a high dismissal rate, which we treated as a tuning problem rather than a verdict.

Once we tuned the tool to comment only above a confidence threshold and scoped it to the categories where it was reliable, the dismissal rate fell and engineers stopped treating it as noise. Only then did we expand it to teams working on more sensitive systems, and even there we kept human review mandatory and unchanged. The AI shortened the path to a good human review. It never replaced the human signature on the merge.

The metric that mattered most was not lines reviewed per hour or comments generated. It was whether senior engineers felt they had more attention to spend on the hard questions. When the answer became yes, we knew the tool was doing its actual job, which was to clear the mechanical underbrush so human judgment could reach further.

Anselm Fowel, CTO and fintech architect
Anselm Fowel — CTO & fintech architect

Conclusion

AI is changing code review, but not in the way the loudest demonstrations imply. It is not replacing reviewers, and in a regulated payments environment I would be deeply skeptical of any vendor who claims it should. What it is doing is shifting the human's attention up the value chain, away from the mechanical scanning that machines now do well and toward the contextual judgment that still requires people who understand the system, the regulations, and the consequences of getting it wrong. Used with discipline, clear governance, and an unbroken line of human accountability, AI makes review faster and frees senior engineers to focus where they are irreplaceable. Used carelessly, it offers confident wrongness at scale. The technology is real and useful. The judgment about how to wield it remains, as it always has, ours.

Enjoyed this article? Share it with others!

Share:

Get new posts in your inbox

Occasional, practical notes on engineering leadership, fintech, and building with AI. No spam, unsubscribe anytime.

Comments (3)

Leave a Comment

Comments are moderated and will appear after review.

Ashley Wilson

August 18, 2026

Enjoyed this one. A small nit on "What Code Review Was Always Protecting": worth mentioning write amplification on batched flushes — otherwise the pattern degrades under real load.

Kwabena Adjei

August 16, 2026

Quick q on "From Linters to Language Models" — how do you handle backpressure when the downstream service times out? We're on Kotlin + Ktor and ops keep asking for manual replay tooling.

Ngozi Iheanacho

July 27, 2026

Good topic. Fractional CTO — done 3 fintech engagements in the last 5 years here. What we do differently: keep an append-only audit log and rebuild state from it on demand on AWS org unit split. It is not universally better; the on-call story is worse, but the audit story is dramatically better and that pays for itself the first time you have to answer a SOC 2 auditor question at 4am.

About the author

Anselm Fowel

Anselm Fowel

Chief Technology Officer & fintech architect. 16+ years leading engineering across AlliancePay, Mondu, Transalliance, Global Accelerex, and Fidelity Bank — writing here about engineering leadership, fintech architecture, and AI in production.

Read next