Skip to content

Assisted code review πŸ‘οΈ

Code review is an exercise in judgment, not speed. An agent can read a diff and flag issues in seconds β€” but it cannot understand the context in which that code lives. It can tell you “this function is too long” but it cannot tell you “this function is long because the business required that specific logic three releases ago.” The agent is a second pair of eyes, not a second brain.

The agent sees the code. It does not see the context. Real review requires both.

When to use the agent for this case 🎯

  • PR review before merge β€” automatic check for security, performance, readability.
  • Review of code generated by other agents β€” self-evaluation with limits.
  • Reusable review checklist β€” standard prompt to apply to every PR.
  • Anti-pattern identification β€” duplicated code, overly long functions, circular dependencies.

When NOT to use the agent β›”

  • The review requires domain knowledge β€” if the code implements complex business logic, the agent cannot evaluate functional correctness.
  • The PR is huge β€” more than 500 lines. Break it down first, then delegate review by parts.
  • You need to decide whether to approve or not β€” the final judgment is always human. The agent flags, it does not decide.

Opening prompt πŸ“

“Review this PR/diff. Checklist: [security, performance, readability, AGENTS.md coherence, edge cases]. Flag with severity [blocks/suggests]. Do not modify files.”

Concrete example:

“Review the diff in the feature/payment branch. Checklist: security (injection, hardcoded credentials), performance (N+1 queries, memory leaks), readability (naming, long functions), AGENTS.md coherence (conventions, patterns). Flag with severity: blocks / suggests. Do not modify files.”

Context setup πŸ”§

  1. Provide the diff β€” the agent must see the modified code, not the entire file.
  2. Define the checklist β€” security? Performance? Readability? Coherence? All together is too much: prioritize.
  3. Specify severity β€” blocks (must be fixed before merge) vs suggests (optional improvement).

Workflow πŸ’‘

  1. Prepare the diff β€” only relevant changes, not the entire branch history.
  2. Send with checklist β€” the agent analyzes the diff against the defined criteria.
  3. Review findings β€” every finding must be verified: is it a real problem or a false positive?
  4. Apply necessary corrections β€” only those you agree with, not everything the agent flags.
  5. Re-run the review after corrections β€” to verify that fixes did not introduce new problems.

Findings should be treated as suggestions, not verdicts. The agent can flag suspicious patterns, but the decision is always yours.

For review of agent-generated code, use the same approach: the agent that generated the code is not the best one to evaluate it. Delegate the review to a different context.

Pathologies of AI-generated code 🦠

Code generated by models presents failure patterns different from human code. Human code has obvious syntax errors, inconsistent formatting, incomplete logic when rushed. AI-generated code maintains uniform syntax, professional naming, and plausible structures β€” even when the underlying logic is fundamentally flawed.

Invented APIs

The primary failure mode. The agent invokes methods that do not exist on valid objects, references deprecated arguments, imports packages that were renamed or never existed. In multi-tenant environments, this enables supply chain vulnerabilities like slopsquatting: attackers register hallucinated package names on public registries (PyPI, npm) to inject malicious payloads.

The tell is that generated code reads more naturally than the real API β€” because it was generated to be plausible, not recalled from documentation.

Tautological tests

When the agent generates tests for its own code, it creates suites that pass consistently and inflate coverage metrics, but do not validate real operational behavior. They often assert that a function’s return value equals an identical call to the implementation β€” testing the mock configuration, not the business logic.

Cargo-cult abstractions

Trained on vast enterprise repositories, agents default to complex architectural abstractions for basic features. A file parser becomes an abstract factory interface with strategy pattern and multi-layered dependency injection. Unnecessary structural complexity that complicates debugging and inflates the maintenance surface.

Quantified security review πŸ“Š

Veracode 2026 data (100+ models tested, 4 longitudinal snapshots):

  • Average security pass rate: 56% β€” virtually unchanged from the first report.
  • ~44% of generation tasks introduce an OWASP Top 10 vulnerability.
Language Security pass rate Primary vulnerabilities
Java 29% SQL injection (CWE-89), output encoding (CWE-80), logging (CWE-117)
C# (.NET) 55% Cryptography (CWE-327), Entity Framework raw queries, path traversal
JavaScript/TypeScript 57% XSS (CWE-80), prototype pollution, unvalidated redirects
Python 62% Command injection (subprocess), deserialization (pickle), access control

Coding-specialized models (51%) perform no better than general-purpose (52%). Model size (>100B parameters: 53%) does not meaningfully improve pass rate.

The 1Password study: of 6,000+ AI-generated security patches, only 26% fully corrected the vulnerability without unintended effects. Even asking the agent to fix a known vulnerability is unreliable without re-verification.

The SAST blind spot

Formal research using Z3 solvers on AI-generated artifacts reveals that traditional SAST tools miss up to 97.8% of verified security flaws. Rule-based scanners rely on pattern matching and miss deeper semantic vulnerabilities.

Auto-review: structural limits 🚫

When the same model generates and reviews code, both agents draw on the same training corpus, the same pattern vocabulary, and the same assumptions about what “correct” code looks like β€” a closed loop with no exit point touching the original spec.

Empirical data:

  • Legacy modernization study (1,980 calls, 11 LLMs): when the model silently changed program behavior, 31.7% were silently endorsed as correct by the very model that produced them.
  • GPT-3.5 self-review on vulnerabilities: 43.6% accuracy β€” close to random guessing.
  • GPT-4 self-review: 74.6% β€” better but not sufficient.
  • Adversarial peer review experiment: when two reviewers converge on the same false critique, the worker agent exhibits sycophancy and modifies correct code, introducing errors.

The practical rule: AI review is a fast first pass for mechanical/pattern-based issues, but it is not independent verification. It does not replace a human or a genuinely separate check (deterministic static analysis, specification tests the agent cannot touch).

Severity classification and SARIF πŸ“‹

Classifying findings by severity is essential to prevent alert fatigue. A 4-tier system:

Severity Definition Pipeline Action
Critical Exploitable vulnerabilities (SQLi, Auth Bypass), immediate data exposure Hard Block β€” prevents merge
High N+1 queries in core paths, missing authorization, unhandled exceptions Conditional Block β€” requires sign-off
Medium Missing input validation, unbounded collection growth, non-standard error handling Advisory β€” inline comment, merge with approval
Low / Advisory Style mismatches, redundant variables, simplification opportunities Informational β€” collapsed summary

Findings must be exported in SARIF v2.1.0 format for integration with GitHub Code Scanning, GitLab Security Dashboard, or Azure DevOps. This standard enables direct mapping of severities to pipeline gates.

Anti-patterns in code review 🚫

Anti-pattern Description Data
Review fatigue Too many findings generate noise; engineers stop treating “critical” as urgent Cubic.dev: up to 40% of alerts ignored once alert fatigue sets in
Trust paradox The more syntactically polished the generated code, the more the human reviewer defaults to superficial inspection 96% skeptical of AI code safety, fewer than 48% verify logic before merging
Unbounded self-correction Without iteration limits, agent modifies code randomly and exhausts token budgets Max 3-5 attempts; escalate to human when limit is reached
Unvalidated auto-merge Automatic merge of dependency bumps just because tests pass Unit tests miss integration edge cases, performance regressions, memory leaks

GitClear 2026 report (623M code changes): cross-file function calls (proxy for genuine reuse) are down 35%; refactoring line moves are down 70%; long-term legacy maintenance is down 74% versus 2022. The default AI workflow incentivizes atomic code β€” a happy path, a passing test, a closed ticket β€” while taxing the invisible work of reuse, consolidation, and error-surfacing.

Acceptance criteria βœ…

  • Every finding is verifiable and localized (file:line).
  • No systematic false positives (if the agent always flags the same things unnecessarily, adjust the checklist).
  • No code modifications during review (read-only).
  • Applied corrections have been re-evaluated.

Further reading πŸ“š

Last updated on