AI & Engineering
AI Can Write Your Code. Can It Review It?
A practical look at AI code review: real bugs, false alarms, missed issues, and the evidence to ask for before merging.
Rommel Clarino 7 min read
At a glance
An AI reviewer can help you find problems, but its comments need checking—and its silence needs questioning. GitHub’s new ReviewBench provides a timely starting point. A hypothetical refund feature shows how to turn an AI finding into a decision you can defend.
A confident comment is the beginning of a review
Imagine opening a pull request with an AI-generated implementation and an AI-generated review. The code looks tidy. The reviewer has left three comments. You fix them, run the checks, and the conversation goes quiet. Is the change ready?
That quiet ending is easy to mistake for evidence. It tells you the reviewer has nothing more to say. It does not tell you whether the feature handles the business rules, the failure paths, or the assumptions nobody wrote down.
AI review becomes useful when it helps you investigate a change. Every finding should lead to a concrete question: what input triggers this, what happens, and why is that behavior wrong?

Why ReviewBench is a useful starting point
GitHub introduced ReviewBench on October 5, 2026. Its research preview evaluates review agents on 219 public pull requests from 187 repositories across 19 languages. Those are the evaluation cases; the 103.9 million pull requests GitHub analyzed informed the dataset’s composition.
GitHub describes a reference set assembled from several sources, evaluated with an LLM judge and audited by senior engineers. Results distinguish findings that match known issues from additional findings the judge considers valid. Severity and category breakdowns help readers look beyond a single overall score.
These are GitHub’s reported methodology and findings, not measurements from an experiment conducted for this article. The benchmark offers a useful comparison under defined conditions. It cannot establish that a reviewer understands your application’s unwritten rules or will catch every important bug in your next change.
Three review outcomes, one hypothetical refund feature
Consider a simplified refund endpoint. This is an invented teaching example, not a real incident or an evaluation of a named AI tool. A customer submits an order ID and a refund amount. Authentication identifies the customer, and the request schema already rejects non-integer, zero, and negative amounts. Amounts are expressed in cents.
The agreed rules are straightforward: a customer can refund only their own order, and the requested amount cannot exceed the refundable balance. For this walkthrough, assume the payment integration handles duplicate requests and concurrent refunds elsewhere. We are isolating two mistakes in this endpoint.
The implementation loads an order, checks that the requested amount does not exceed the original purchase total, and calls the refund operation. It never checks the order’s owner. It also ignores previous refunds when calculating the remaining balance.
The real bug: signed in does not mean authorized
The reviewer flags the missing ownership check. That is a useful lead, but “authorization issue” is still too vague to act on confidently.
A stronger finding explains the path: customer A is signed in, submits customer B’s order ID, and reaches the refund call because the endpoint never compares the order’s owner with the authenticated customer. The expected behavior is rejection before any refund is attempted.
Verification should demonstrate that boundary with two customers and a controlled payment stub. Then the fix can enforce ownership in the lookup or before the refund call, with a check that the unauthorized request never reaches the payment operation.
- Trigger: a valid session submits another customer’s order ID.
- Evidence: no ownership restriction exists along the request path in this example.
- Required outcome: reject the request before causing a payment side effect.
The false alarm: validation exists outside the diff
The imagined reviewer also says the endpoint accepts negative amounts. Taken in isolation, the refund function appears to allow them. But this example’s request schema already rejects those values before the function runs.
To dismiss the warning, trace the actual route and confirm the schema is applied. A validation helper sitting unused in the repository would prove nothing. An alternate call path that bypasses it could turn the warning into a real issue.
Once the protected path is confirmed, record the reason for dismissal. Adding another check may be a reasonable design choice, but it should follow an explicit boundary decision. A plausible comment alone is not a reason to scatter duplicate rules around the application.
The missed issue: the remaining balance is smaller
The reviewer says nothing about previous refunds. Yet the code compares the new refund with the original total, not the refundable balance.
Take an order worth 10,000 cents with 8,000 cents already refunded. A new request for 5,000 cents passes the original-total check even though only 2,000 cents remain. Whether the payment provider rejects it later is a separate question; the application has failed to enforce its own agreed rule.
This is why reviewing comments and reviewing the change are different activities. You can correctly resolve every comment and still miss a requirement. Walk through the acceptance criteria independently, including changes in state over time.
| Quantity | Amount in cents |
|---|---|
| Original purchase | 10,000 |
| Already refunded | 8,000 |
| Actually available | 2,000 |
| New refund request | 5,000 — must be rejected |
Measure useful findings and missed problems separately
In our invented example, the AI raised two warnings and one was valid. Its precision is therefore 1 out of 2, or 50%. We deliberately placed two real bugs in the example and it found one, so its recall is also 50%.
The identical percentages describe different things. Precision asks how much of the review deserves action. Recall asks how much of the known problem set the reviewer found. In a real repository, you rarely know every existing bug, so recall against known cases is an estimate with a limited reference set.
Severity adds another dimension. Catching a cosmetic inconsistency and catching an unauthorized refund should not carry the same practical weight. A reviewer that leaves more comments may simply create more work to sort through.
| Measure | Useful question |
|---|---|
| Valid findings | How many comments describe a reproducible problem? |
| Missed known bugs | Which historical failures does the reviewer overlook? |
| Severity | Does it identify failures with meaningful user impact? |
| Review effort | How long do people spend validating or dismissing comments? |
| Total cost | What do the review run, follow-up investigation, and rework cost? |
Give the reviewer context it can actually use
Start with the intended behavior, the diff, and the relevant surrounding code. For the refund example, that includes the request schema, authentication and authorization boundaries, the order model, and the payment integration’s guarantees. Keep unresolved assumptions visible.
A useful review request is: “Review this refund change against the attached acceptance criteria. Follow the request path through validation and authorization. For each finding, identify the relevant code, a triggering input or state, the observed consequence, and the expected behavior. Mark missing context as a question. Prioritize correctness and user impact.”
Ask the reviewer to distinguish a demonstrated failure from a possibility that needs investigation. That makes the output easier to evaluate and keeps uncertainty attached to the claim.
A review workflow worth repeating
Use AI review as another source of evidence alongside the implementation, the specification, and the checks you can run. The person deciding to merge should be able to explain why the important behavior is covered.
- State the intended behavior and important constraints before reviewing.
- Run the existing checks and inspect the diff and relevant call paths.
- Ask AI for specific findings, with triggers and consequences.
- Reproduce credible issues; resolve false alarms with concrete evidence.
- Check important requirements even when the reviewer says nothing about them.
- After a fix, inspect the updated change and rerun the relevant checks.
Make the merge decision explainable
AI can help review code. The useful outcome is a clearer understanding of the change: which failures were found, how they were verified, what was fixed, and what remains uncertain.
For your next pull request, try a small habit: take the most important review comment and turn it into a reproducible scenario. Then choose one important requirement the reviewer never mentioned and check that yourself. Those two actions make the review stronger than accepting either confident warnings or reassuring silence at face value.