AI code review in a real team: CodeCritic lessons
I've been building CodeCritic - an AI code review tool - and running it on real pull requests, including my own teams', for about a year. Here's what survived contact with reality, and what didn't.
The short version: the model was never the hard part. The hard part was everything around it - what context it sees, what it's allowed to say, and whether the team believes it after week two.
The noise problem is everything
The first version commented on everything: naming, missing types, style preferences. Developers muted it within a week. A review bot lives or dies by one metric: of the comments it leaves, how many would a senior engineer agree with?
We cut the comment volume by ~80% and usage went up. The rule that made the difference: the bot only speaks when it's confident and the issue is actionable. "Consider renaming this variable" is noise. "This query will do a full table scan at 2M rows - here's the index" is gold.
The counterintuitive part: for adoption, precision beats recall. A bot that misses half the bugs but never cries wolf gets read every day. A bot that catches everything but leaves twenty comments per PR gets muted in a week - and once muted, it's effectively dead, no matter how good its catches were. Every threshold we later tuned was in service of that one trade: fewer comments, each one worth reading.
Context beats model size
A diff alone is nearly useless for review. The same three changed lines can be a bug or a no-op depending on the caller, the schema, the tests. The biggest quality jump didn't come from switching models - it came from feeding the reviewer more context.
What we assemble per review:
- the diff itself;
- the full content of every touched file - not just the changed hunks;
- tests that reference the changed functions;
- the last few commits that touched the same files (so the bot can tell "new bug" from "pre-existing mess").
One example that sold the team on it: a PR changed a default timeout from 30 to 5 seconds. In the diff - a one-character change, harmless. With the full file in context, the bot found the retry wrapper around that call: 5s timeout × 3 retries × 40 concurrent jobs = a queue that stalls every deployment. A diff-only reviewer waves that PR through every time.
If you're building something similar: spend your effort on context assembly, not on prompt wording. The prompt is 10% of the quality. The other 90% is what the model gets to look at.
Trust is built in small increments
Nobody reads a wall of AI comments. But a bot that catches one real bug per week, quietly, becomes part of the process. Our adoption curve followed a simple pattern: silence for a while, then one caught production-bound bug, then the team starts reading every comment.
Practical consequence: don't measure the bot by comments per PR. Measure it by bugs caught before merge and by how often humans act on its comments. Those two numbers are the entire product. Comment volume is a vanity metric - and usually an inverse one.
Where AI review actually saves time
- Pre-human filtering. The bot catches the mechanical stuff (null handling, N+1 queries, missing error paths) so humans review architecture and intent. The senior engineer's attention is the scarcest resource in the loop - everything mechanical should be drained from it before a human looks.
- First-pass on trivial PRs. Dependency bumps and typo fixes get reviewed in seconds, freeing the reviewer for real work. Roughly a third of PRs in a typical week are mechanical; that's a third of review requests that never interrupt anyone.
- Consistency checks. "We agreed to do it this way" rules are perfect for a bot and boring for humans. The bot never forgets that the team decided all money values are integers in minor units.
Where it doesn't: naming debates, API design, anything where "correct" depends on product context the bot doesn't have. We tried; the comments were technically defensible and practically useless, which is the worst combination - they burn trust exactly like noise does.
Two things we got wrong
Reviewing generated code too harshly. Migrations, boilerplate, codegen output - the bot flagged style issues in files humans never read line by line. Wasted comments, wasted trust. The fix was a simple heuristic: generated files get a reduced rule set.
Trusting the diff boundary. A change that looks local can have cross-service effects - a renamed field here, a consumer over there. The bot initially reviewed each repo in isolation and missed exactly the class of bugs it was most valuable for. Cross-repo context is on the roadmap precisely because of that gap.
What I'd tell a team adopting AI review
Start with read-only mode: the bot comments, nobody is required to act. Tune for two weeks until the signal-to-noise ratio is something a senior engineer respects. Only then make it blocking. Rolling it out the other way - strict from day one - is how you get a muted bot and a cynical team.
And keep a human in the loop for anything that touches money, auth, or data deletion. The bot is a filter, not a gate. The day you treat it as the latter is the day it silently approves something expensive.
Want AI review wired into your team's workflow? Get in touch - or try CodeCritic on your next PR.