That AI-generated code contains more security flaws than human-written code is, at this point, well established, independent studies have converged on roughly 40% of AI-generated code containing a vulnerability, a figure that's held up across different models, different research teams, and several years of model improvement. What's less commonly explained is how this actually happens at a mechanical level. Understanding the mechanism matters, because the fix looks very different depending on whether the problem is a knowledge gap, a training artifact, or something else entirely.
Here's what the research actually shows about how vulnerabilities get introduced, and why some of the most intuitive assumptions about the cause turn out to be wrong.
-
Mechanism 1: The Training Data Isn't Curated for Security
AI coding models learn from enormous volumes of real, public code, and that code was written by real developers, with all the mistakes real developers make. If a meaningful share of the login examples a model was trained on handled passwords insecurely, the model has no independent signal that those examples were wrong. It's not reasoning about security from first principles; it's reproducing patterns weighted by how often it saw them, and popular doesn't always mean correct. Researchers examining this directly have found this isn't a marginal effect, insecure patterns showing up frequently enough in training data measurably shift how often a model reproduces them, sometimes tracking closely with the mix of good and bad examples it learned from in the first place.
-
Mechanism 2: The Model Is Optimized for "Works," Not "Safe"
This is the deeper structural issue. AI coding assistants are trained to produce code that satisfies the prompt, functionally correct, plausible, aligned with what the developer asked for. Security isn't the thing being optimized for, and it isn't consistently represented as a requirement unless a developer explicitly asks for it. Left to its own defaults, a model tends toward the simplest version of a solution that technically works, which is very often the least secure version, because secure implementations usually require more code, more edge-case handling, and more explicit constraints than the minimal path does.
-
Mechanism 3: The Surprising "Format-Reliability Gap"
Here's the part that upends the simplest explanation for all of this. If the problem were purely a knowledge gap, the model just doesn't know what secure code looks like, then asking the model directly about a vulnerability should fail the same way generating it does. Recent mechanistic research on this found the opposite: the same models that generate insecure code will correctly identify and explain that exact vulnerability when asked about it directly afterward. Researchers have termed this the format-reliability gap, the model's internal representation of "this is insecure" exists early in its processing, but during code generation, the competing pressure to produce well-formatted, prompt-satisfying output crowds it out before it ever surfaces in the actual code.
In plainer terms: the model often "knows better" in some internal sense, and generates the vulnerability anyway, because generating secure code and generating plausible-looking, prompt-satisfying code are handled differently by the underlying process, and the second one wins by default. That's a strange and important distinction, because it means simply having the model "know about" the OWASP Top 10 doesn't reliably stop it from producing an OWASP Top 10 vulnerability.
-
Mechanism 4: Fixes Are Often Local, Not Contextual
Even when a developer explicitly asks a model to secure a piece of code, the fix tends to be narrowly scoped to exactly what was flagged. Researchers testing this directly found a telling example: asked to convert a vulnerable SQL query into a parameterized one, a model successfully fixed the query itself, and left security issues in the surrounding logic completely untouched. The model solved the specific instruction with precision and missed the broader context the instruction was embedded in, because it doesn't hold a full, persistent model of the application's security posture the way a developer familiar with the whole codebase would.
This is a direct consequence of how these tools process code: they reason over what's in front of them, not the full system that code lives inside. A fix that's technically correct in isolation can still leave an application vulnerable, because the vulnerability was never really about that one line, it was about how that line interacts with everything around it.
-
Mechanism 5: Under-Specified Prompts Default to the Least Secure Path
Tie the above mechanisms together and a pattern emerges: security in AI-generated code is largely opt-in, triggered by explicit instruction rather than applied by default. A prompt that says "write a login function" and a prompt that says "write a login function following secure authentication practices" can produce meaningfully different code from the same model, not because the model didn't know better, but because the first prompt never asked it to prioritize that dimension, and its default optimization target doesn't include it automatically.
-
What This Means in Practice
The format-reliability gap is the finding that should reshape how organizations think about this problem. If the failure mode were purely "the model doesn't know," then more security training data or better prompting would fully solve it. Since the evidence points to something closer to an internal prioritization failure during generation, one that persists even when the model can correctly explain the same vulnerability moments later, the practical implication is that neither trusting the model's judgment nor trusting a developer's general security knowledge is sufficient on its own. What's verifiable is the actual output: whether the specific application, as built, holds up against real testing.
That's a meaningfully different posture than "review the code carefully", it means treating the shipped application itself as the thing that needs to be tested, not just the reasoning that produced it. Hands-on environments where developers or security teams actively try to find and exploit real vulnerabilities, rather than just reading code for red flags, catch exactly the class of issue these mechanisms predict: a fix that looks complete, passes review, and still leaves the surrounding system exposed.



