BlackTree Security · Infrastructure · Automation · AI

Why None of 12 AI Models Passed AWS’s Code-Review Trust Test

The most expensive security alert may be the one everybody eventually learns to ignore. AI code review tools promise to find bugs faster, but a tool that calls safe code vulnerable creates its own workload. A tool that stays quiet to avoid false alarms may miss the flaw that matters. AWS has now put both mistakes on the same scoreboard, and none of the 12 general-purpose models it tested cleared both of its chosen error bars.

In a 9 September research release, AWS introduced Deception Benchmark, a collection of 14,822 code samples across 16 programming languages and more than 70 weakness categories. The primary page gives a publication date but no time. Of those samples, 9,695 were in the scored set. The tests asked models to decide whether a sample was exploitable or safe, including examples in which a familiar-looking vulnerability was neutralised by a small code change or an environmental control.

AWS evaluated 12 models from five providers. Its proposed minimum for production use was a false-positive rate below 10 per cent and a false-negative rate below 10 per cent at the same time. The company reports that no tested model and prompting configuration met both thresholds. That is AWS’s benchmark criterion, not a universal certification standard, and the experiment does not measure every AI security product in the market.

Why safe code deceives AI code review

Imagine a web endpoint taking user input into a database query. A model may recognise a SQL injection pattern and raise an alarm. But if the query is parameterised correctly, the input does not become SQL code. The pattern is present; the exploit path is not. Other benchmark samples ask whether deployment controls, such as a network policy or an identity boundary, block a path that would otherwise be dangerous.

Those distinctions are daily work for an application-security team. A finding needs more than a recognisable code shape. It needs a traceable route from attacker-controlled input to a real consequence, with the actual safeguards evaluated rather than assumed away. Equally, the reviewer must not mistake a cosmetic safeguard for an effective one. The challenge is not to be sceptical of every alert; it is to be accurate about why the code is or is not exploitable.

AWS says direct prompting often found real vulnerabilities but labelled large amounts of safe code vulnerable. Asking models to seek proof of exploitability reduced false positives, yet increased missed vulnerabilities in its tests. These are different failure modes with different costs. The first consumes engineers’ time and confidence. The second leaves defects undetected. An overall accuracy number can hide the trade-off: on a roughly balanced dataset, a model that declares everything vulnerable would catch every real bug while being wrong about every safe sample.

What the result does not prove

Deception Benchmark is an intentionally difficult, adversarially constructed test. It is not a random sample of all code flowing through a typical development team. AWS used single-turn model decisions to expose the base model’s reasoning. A purpose-built review system may retrieve context, run tests, ask follow-up questions and involve a human before reaching a conclusion. Those systems were not the subject of the headline result. BlackTree has also examined why enforced boundaries around AI agents matter beyond a model’s own judgement. Conversely, a vendor should not claim that extra tooling automatically solves the problem without demonstrating it.

The dataset is also designed to resist easy score-chasing. AWS says it withheld 5,127 samples from scoring and does not release the answer labels, while making the samples and evaluation workflow available. That makes the research more useful as a repeatable challenge, though independent teams will still need to assess how closely it resembles their own codebases and deployment environments.

Questions to put to an AI code review vendor

For AI code review, buyers should ask for two numbers, not one: how often does the system miss a confirmed vulnerability, and how often does it flag safe code? Ask how the figures were measured, what languages and frameworks were included, whether the test covered deployment context, and whether the results describe a raw model or the complete workflow being sold. A demonstration on a known vulnerable snippet is not enough to establish that a tool will avoid burying your team in plausible but incorrect findings.

Inside an organisation, record what happens after an AI finding is raised. Can reviewers reproduce the alleged exploit? Does the tool identify the exact data flow and the mitigation it considered? How much analyst time does each false alarm consume? Keep human verification for high-impact findings and for declarations that code is safe. Use local test sets containing both real defects and safe lookalikes; otherwise a tool can look successful merely by saying yes to everything suspicious.

The valuable question AWS has put on the table is not whether AI should be allowed near security review. It is what level of evidence would make its verdict trustworthy. A system that finds more possible bugs but cannot tell an effective defence from a decorative one may move the bottleneck from detection to adjudication. Teams need to measure that cost before calling the workflow faster.

Sources

Leave a Reply

Your email address will not be published. Required fields are marked *