Skip to slide
Chapter 13 · Code Review and Security Agents
92 / 142

CHAPTER 13 · Code Review and Security Agents · 3 / 9

The lesson: an agent cannot mark its own homework

Here is the reliability principle that this chapter exists to deliver, and it generalizes far beyond security. LLM reviewers have two failure modes that are both expensive: false positives (confidently flagging a parameterized query as "critical SQL injection" because the model misread the data flow) and false negatives (missing a real bug because attention drifted across a large codebase). When your detection layer is entirely probabilistic, you accept both risks.

Snyk's principle: "the agent cannot mark its own homework." You need an independent validation layer. The recommended architecture is two-tier: the LLM agent is the researcher (creative, finds novel cross-file logic bugs that rule-based tools miss), and a deterministic engine (static analysis, SAST) is the peer reviewer (catches known patterns with mechanical precision and confirms the LLM's findings are real). You want both, because each catches what the other misses. This is the same idea as Chapter 11's verifier subagent and Chapter 10's deterministic computation, now stated as a law: never let probabilistic reasoning be the only check on something that matters.

Two more details reinforce it. First, Cursor's own review agent prompt explicitly ends with "do not push changes or open fix PRs from this workflow." Even Cursor's security team keeps a human in the loop for their own tooling; the agent finds and reports, a human decides. Second, the BaxBench benchmark found that 62% of solutions from even the best models are incorrect or contain vulnerabilities, which is precisely why layered, independent validation is essential rather than optional.

← → arrow keys work too