AI models have become highly reliable at writing code that compiles, but they are still failing basic security tests in nearly half of cases, according to Veracode’s 2026 GenAI Code Security Report.
Veracode tested more than 100 AI models across four testing rounds. The average security pass rate reached just 56%, barely higher than the 55% recorded last year. By contrast, models now produce syntactically correct, compilable code almost every time.
The issue is becoming more important as AI now generates roughly half of all committed code, according to data cited by Veracode. This means the security failure rate has remained largely unchanged while the amount of AI-written software has increased sharply.
The report found that models designed specifically for programming were no safer than general-purpose AI models.
Coding-focused models averaged a 51% security pass rate, compared with 52% for general-purpose models. Larger models also showed little advantage, with large models averaging 53%, while medium and small models both scored 51%.
Reasoning models performed better, averaging 56% compared with 51% for non-reasoning models. Veracode suggested that the extra reasoning steps may work like an internal code review before the final answer is produced.
OpenAI’s GPT-5.5 ranked first in the latest test with a 68% security pass rate.
Six of the 11 models tested scored between 50% and 53%, while Alibaba’s Qwen3.7-max finished last at 50%. That means even the best-performing model still failed nearly one-third of security-related tasks.
The top score also fell compared with the previous round, where GPT-5-mini had reached 72%. Meanwhile, Chinese models including Kimi-K2.6 and Xiaomi’s MiMo-V2.5 outperformed several Western models, making the field more globally competitive.
Security performance also varied significantly by programming language.
Python performed best with a 63% pass rate, while Java managed only 30%, making it the weakest language in the test. However, Veracode said Java was also the only language showing a clear upward trend over the past year.
The type of vulnerability also made a major difference. Models performed relatively well against SQL injection and insecure cryptography, but struggled badly with cross-site scripting and log injection. Earlier Veracode testing showed similarly large gaps between these categories.
Veracode tested raw models without security-specific prompts, agents, guardrails or human review.
This means the 44% failure rate does not suggest that nearly half of all AI-generated code reaches production with vulnerabilities. Real development workflows can include code reviews, security scanners and other safeguards before software is released.
Veracode, which sells software security products, recommends treating AI-generated code like any other unreviewed code by scanning and fixing it before deployment.
Chris Wysopal, Veracode co-founder and chief security evangelist, said the answer is not restricting access to advanced AI models, but improving security controls around them.
Until AI models become as reliable at security as they are at syntax, the report suggests that automated checks and human review will remain necessary parts of AI-assisted software development.
Get the latest tech news, telecom insights, and product launches wherever you prefer.
Add ProPakistani to Preferred Sources and see more of our stories in Google Search and Top Stories.
Technology and Automotive Specialist covering the latest cars, smartphones, AI breakthroughs, and...
Shares