Skip to slide
Chapter 13 · Code Review and Security Agents
96 / 142

CHAPTER 13 · Code Review and Security Agents · 7 / 9

Code MVP: a two-tier PR review agent

"""
chapter 13: a two-tier review agent.
Tier 1: an LLM reviewer (here, a focused prompt + fake model) finds candidate
issues. Tier 2: a DETERMINISTIC checker independently confirms them. Only
confirmed findings are reported, and fixes are never auto-pushed.
"""
import re

# The actual 15-line-style security prompt (abbreviated).
SECURITY_PROMPT = """You are a security reviewer for pull requests.
Trace attacker-controlled input to a real sink. Verify existing controls do
not already block it. Report only medium/high/critical findings with a concrete
attack path and code evidence. Do not push changes or open fix PRs."""

def llm_reviewer(diff_lines: list, model_fn) -> list:
    """Tier 1: probabilistic. Returns candidate findings (each with a line)."""
    return model_fn(SECURITY_PROMPT, diff_lines)

# Tier 2: deterministic confirmation, LINE-AWARE. The agent cannot mark its
# own homework, so we independently check the exact line it flagged.
DANGEROUS_PATTERNS = {
    # String-formatted SQL ("...%s..." % var) is injectable; a parameterized
    # query ("... ?", (var,)) is not, so this pattern matches one but not the other.
    "sql_injection": r"\"(SELECT|INSERT|UPDATE|DELETE)[^\"]*%s?[^\"]*\"\s*%",
    "dynamic_exec":  r"\b(eval|exec)\s*\(",
    "hardcoded_secret": r"(api_key|password|secret)\s*=\s*['\"][^'\"]+['\"]",
}

def deterministic_confirm(line: str, finding: dict) -> bool:
    """Independently check whether the claimed issue really appears on its line."""
    pattern = DANGEROUS_PATTERNS.get(finding["type"])
    return bool(pattern and re.search(pattern, line))

def review_pr(diff_lines: list, model_fn) -> dict:
    candidates = llm_reviewer(diff_lines, model_fn)        # tier 1
    confirmed, dismissed = [], []
    for f in candidates:
        line = diff_lines[f["line"]]
        (confirmed if deterministic_confirm(line, f) else dismissed).append(f)
    return {
        "confirmed": confirmed,        # report these (still to a human)
        "dismissed": dismissed,        # likely false positives; suppressed
        "auto_fixed": False,           # never auto-push: human decides
    }

if __name__ == "__main__":
    diff_lines = [
        'query = "SELECT * FROM users WHERE id = %s" % user_input',     # 0: injectable
        'cursor.execute(query)',                                         # 1
        'safe = cursor.execute("SELECT * FROM t WHERE id = ?", (uid,))', # 2: safe
    ]
    def fake_model(prompt, lines):
        # The LLM flags TWO candidates: line 0 (real) and line 2 (false positive).
        return [{"type": "sql_injection", "line": 0, "severity": "critical"},
                {"type": "sql_injection", "line": 2, "severity": "critical"}]
    from pprint import pprint
    pprint(review_pr(diff_lines, fake_model))

Run it: the LLM flags two SQL-injection candidates, but the deterministic check confirms only the string-formatted one and dismisses the parameterized query (a false positive). Only the confirmed finding survives, and nothing is auto-fixed. That two-tier structure, probabilistic researcher plus deterministic peer review plus human gate, is the chapter in a nutshell.

← → arrow keys work too