in ,

Study Reveals AI Coding Tools Continue to Make Costly Mistakes

New Report Shows Why AI Coding Still Carries Major Risks

AI models have gotten really good at one thing. They write code that compiles. That part’s basically solved now.

But there’s a problem. These same models still fail basic security tests. Almost half the time. That’s according to Veracode’s 2026 GenAI Code Security Report.

Hosting 75% off

The numbers are stark. Veracode tested over 100 AI models. Four separate testing rounds. The average security pass rate? Just 56%. Barely better than last year’s 55%.

Here’s the contrast. Syntactic correctness is nearly perfect now. Models produce clean, compilable code almost every time. Security reliability hasn’t kept pace at all.

This gap matters more than ever. Why? AI-generated code now makes up roughly half of all committed code, per Veracode’s data. So the failure rate stayed flat. But the volume of AI-written software exploded. That’s a dangerous combination.

Read More: Why AI Coding Tools Bring Both Benefits and Challenges to Open Source

Specialized Coding Models Aren’t Actually Safer

You’d think coding-specific AI models would perform better here. They don’t.

Veracode’s testing found something surprising. Models built specifically for programming showed no real security advantage. Coding-focused models averaged a 51% pass rate. General-purpose models scored 52%. Basically identical.

Model size didn’t help much either. Large models hit 53% on average. Medium and small models both landed at 51%. The gap is almost meaningless.

One factor did make a difference, though. Reasoning models outperformed their non-reasoning counterparts. They averaged 56% versus 51%. Veracode has a theory here. The extra reasoning steps might function like an internal code review process. The model essentially double-checks itself before finalizing an answer.

GPT-5.5 Tops the Charts, Still Fails Constantly

OpenAI’s GPT-5.5 claimed the top spot in this round. Its security pass rate hit 68%. That’s the best score in the entire test.

But look closer. That still means GPT-5.5 failed nearly one in three security-related tasks. Even the industry leader isn’t close to reliable.

The rest of the field looked rough. Six of the 11 tested models scored between 50% and 53%. That’s barely above a coin flip. Alibaba’s Qwen3.7-max finished dead last at 50%.

There’s another wrinkle too. The top score actually dropped compared to the previous testing round. GPT-5-mini had scored 72% before. Now the leader sits at 68%. Progress isn’t linear here.

Meanwhile, competition is heating up globally. Chinese AI models like Kimi-K2.6 and Xiaomi’s MiMo-V2.5 outperformed several Western models. The AI coding landscape is becoming genuinely competitive worldwide.

Read More: 5 Vibe Coding Techniques Any Company Can Start Using Now

Java Is Still the Weak Link

Programming language choice matters a lot for security outcomes.

Python performed best overall. It hit a 63% pass rate. Java, on the other hand, struggled badly. It managed just 30%. That makes Java the weakest language tested, by a wide margin.

There’s a silver lining, though. Veracode noted that Java was the only language showing real improvement over the past year. It’s climbing, just from a low starting point.

Vulnerability type also shaped the results heavily. Models handled SQL injection and insecure cryptography reasonably well. But cross-site scripting and log injection? Models struggled significantly with both. This mirrors patterns Veracode found in earlier testing rounds too.

Don’t Panic—This Isn’t a Production Snapshot

Here’s an important caveat. Veracode tested raw, unmodified AI models. No security-specific prompts. No AI agents. No guardrails. No human review layered on top.

That distinction matters a lot. The 44% failure rate doesn’t mean nearly half of all AI-generated code ships to production riddled with vulnerabilities. Real-world development workflows typically include multiple safety nets. Think code reviews, security scanners, and other checkpoints before release.

Veracode sells software security products, worth noting for context. Their recommendation is straightforward. Treat AI-generated code exactly like any other unreviewed code. Scan it. Fix it. Then deploy it.

Chris Wysopal, Veracode’s co-founder and chief security evangelist, offered a clear take. The solution isn’t restricting access to powerful AI models. Instead, companies need stronger security controls built around how those models get used.

Read More: 20 Vibe Coding Tools Every Developer Should Know

The Bottom Line

AI coding tools aren’t going anywhere. Adoption keeps climbing. But security hasn’t caught up to capability yet.

Until AI models get as reliable at security as they are at syntax, the report’s message is clear. Automated security checks and human oversight remain essential. They’re not optional extras. They’re core parts of any responsible AI-assisted development process.

The takeaway for developers and companies alike? Treat AI-generated code with healthy skepticism. Verify before you trust. That’s true today, and it’ll likely stay true for a while.

Hosting 75% off

Written by Hajra Naz

Claude-Opus-5-Used-Tricks-to-Win-an-AI-Business-Test.

Claude Opus 5 Outsmarts Rivals in AI Vending Machine Test