BlackTree Security · Infrastructure · Automation · AI

Google Put AI on Every Code Change and Cut False Positives to 3%

Google says AI security agents now scan every code change across hundreds of millions of lines used in its infrastructure. The company claims the system prevents hundreds of vulnerabilities each month from reaching its codebase or production. The interesting part is not the scale. It is how Google tries to keep the agents from drowning developers in plausible but wrong findings.

In an official 18 September engineering post, Google describes pre-submit agents that review individual code changes rather than running only large, periodic scans. A smaller change supplies less irrelevant context, allowing the system to reason about a narrower security decision.

Threat models became live inputs

Google evolved its open-source Mantis multi-agent review harness and paired specialised agents with local threat models. Instead of relying on a static document, the system uses codebase metadata and dependency call graphs to identify which vulnerable paths an attacker could actually reach.

The company says this approach reduced false-positive rates to 3 per cent in some cases. That is not a universal benchmark for every codebase or vulnerability class. It is Google’s reported result for parts of its own system, with access to internal metadata that many organisations do not possess.

A second agent has to prove the path

Findings pass through a specialised triage agent that uses abstract syntax trees, call-graph traversal and pre-indexed safety rules to test whether an attacker can reach the vulnerable code. Google reports more than 92 per cent precision and a response in under a minute. Nightly post-submit analysis provides another layer for defects that appear only across multiple changes.

A separate fixing agent uses the finding and its proof to propose a patch consistent with internal coding standards. The change still goes to human review. That design keeps generation, validation and approval separate rather than asking one model to find a problem, judge itself and merge its own answer.

What other organisations can copy

  • Scan changes before merge. Smaller diffs reduce context and shorten the feedback loop.
  • Keep agents separate. Use different tools and context for generation, scanning and triage.
  • Ground findings structurally. Require a reachable path or reproducible proof instead of accepting a model’s explanation alone.
  • Maintain live threat models. Connect architecture, dependency and ownership data to the review.
  • Measure precision. Adoption will collapse if developers learn that security-agent findings are mostly noise.
  • Keep humans in the merge path. Automated fixes should be reviewed like other code changes.

Google’s report is a provider account, not an independent audit, and it does not publish raw vulnerability samples. The operational model is still valuable: agentic security becomes credible when its reasoning is constrained by live architecture, tested by deterministic analysis and separated from the authority to deploy.

Sources

Leave a Reply

Your email address will not be published. Required fields are marked *