option
Home
News
AI Code Performance Overstated in Real-World Tests, Study Finds

AI Code Performance Overstated in Real-World Tests, Study Finds

May 9, 2026
207

Research from the METR institute indicates that the widely used SWE-bench Verified benchmark for evaluating AI programming capabilities may significantly overstate AI agents' performance in real-world software development. The study found that approximately half of the AI-generated code solutions marked as "passed" by the benchmark would likely be rejected by actual project maintainers during code review, highlighting a substantial gap between automated assessment results and real-world code quality.

SWE-bench Verified has long been regarded as a key standard for evaluating AI-assisted software engineering, testing whether models can solve genuine programming tasks in open-source projects and verifying if code changes pass the project's automated test suite. Several AI companies, including Anthropic and OpenAI, frequently cite results from this benchmark to demonstrate model advancements.

QQ20260312-093454.jpg

In this study, the METR team enlisted four experienced developers who maintain the open-source projects scikit-learn, Sphinx, and pytest to manually review 296 pieces of AI-generated code. These code samples were produced by five different models: Claude 3.5 Sonnet, Claude 3.7 Sonnet, Claude 4 Opus, Claude 4.5 Sonnet, and GPT-5. The results revealed that the actual acceptance rate by maintainers was, on average, about 24 percentage points lower than the automated scores from SWE-bench—a statistically significant difference.

The study also determined that the rejected AI code was not primarily due to stylistic issues but rather more substantial engineering flaws. Maintainers categorized the problems into three main types: code quality failing to meet project specifications, disruption of the existing code structure, and fundamental functional errors. A significant portion of the cases involved functional errors where, despite passing automated tests, the code did not correctly solve the intended problem.

Regarding model comparison, the research found that upgrading from Claude 3.5 Sonnet to Claude 3.7 Sonnet significantly improved the benchmark pass rate, but the number of functional errors flagged by maintainers also increased. The transition from Claude 3.7 Sonnet to Claude 4 Opus saw a shift toward more code quality issues, while Claude 4.5 Sonnet showed improvements in code quality. In contrast, GPT-5 performed notably worse overall than the Anthropic model series in this manual assessment.

Artificial intelligence brain, large model

The research team also performed an estimated analysis of "task completion time": according to SWE-bench automated evaluation results, it would take Claude 4.5 Sonnet approximately 50 minutes of human effort to complete tasks with a 50% success rate. However, based on the maintainers' scores, the estimated time drops to only about 8 minutes, suggesting the benchmark may overestimate capabilities by up to sevenfold.

However, the researchers also emphasized that this study does not imply a fundamental limit on AI programming agents' capabilities. With improved prompting strategies, more human feedback, or multiple iteration cycles, the gap between automated evaluation and manual review could be reduced. Additionally, the experimental setup differs from real development processes—for instance, AI agents had only one submission attempt, whereas human developers can typically iteratively modify code based on feedback.

In summary, the study concludes that relying solely on benchmark scores to assess the practical utility of AI programming agents may introduce systematic bias. As AI coding models evolve rapidly, developing evaluation systems that better reflect real-world development environments has become a crucial research direction in AI software engineering.

Related article
South Korea Breaks Ground on National AI Computing Center, Investing 2.5 Trillion Won with 2028 Target South Korea Breaks Ground on National AI Computing Center, Investing 2.5 Trillion Won with 2028 Target South Korean outlet EtNews reports that groundbreaking for the Korea AI Computing Center (KOACC) took place on August 3 at the Solar City data center park in Sunan, Jeollanam-do. Backed by a total investment of 2.5 trillion KRW (roughly 11.838 billio
Six Tech Giants Back Linux Foundation With $12.5M to Tackle AI Vulnerability Noise Six Tech Giants Back Linux Foundation With $12.5M to Tackle AI Vulnerability Noise To tackle the flood of low-quality security reports produced by AI automation tools, six major tech companies—Anthropic, Amazon (AWS), GitHub, Google, Microsoft, and OpenAI—have collectively contributed $12.5 million in funding to Linux Foundation in
Musk Considered Leaving OpenAI to His Kids as Altman Testifies Musk Considered Leaving OpenAI to His Kids as Altman Testifies This morning, OpenAI CEO Sam Altman took the stand to address former co-founder Elon Musk’s lawsuit challenging the company’s corporate structure.When asked about Musk’s claim that other founders “stole a charity” by launching a for-profit subsidiary
Related Special Topic Recommendations
Music composition AI Vocal Demo Tools for Songwriters, Hooks, Toplines, and Multilingual Draft Sessions
AI Vocal Demo Tools for Songwriters, Hooks, Toplines, and Multilingual Draft Sessions

2026 Latest Best AI Vocal Demo Tools for Songwriters, Hook Creators, and Multi-Language Content Teams! XIX.AI has curated a top-rated list of powerful game-changing tools that go through rigorous real-world tests. You’ll find detailed free vs paid comparison data, comprehensive rankings, and must-try options to help you boost writing efficiency and unlock your creative potential. Explore now to discover your perfect tool for all your content needs!

9 tools
xix.ai
Business Best AI Competitive Research Tools for Small Businesses
Best AI Competitive Research Tools for Small Businesses

2026 Latest Best Top-rated AI Competitive Research Tools for Small Businesses! XIX.AI has curated a highly powerful game-changing collection, updated weekly with rigorous real-world tests and detailed rankings. You can find a comprehensive free vs paid comparison to help you identify the must-try tools that boost your productivity and give you a competitive edge. Explore now to discover your perfect tool!

9 tools
xix.ai
Image editing Photoshop AI Retouch Tools for Ecommerce Apparel, Skin Cleanup, and Color Consistency
Photoshop AI Retouch Tools for Ecommerce Apparel, Skin Cleanup, and Color Consistency

2026 Latest Best Photoshop AI retouch tools for ecommerce apparel, skin cleanup, and color consistency! This top-rated curated list features powerful game-changing solutions that help you boost writing efficiency, streamline content creation, and achieve perfect visual results effortlessly. Each tool has undergone real-world tests through weekly updated rankings, complete with free vs paid comparison details. Backed by XIX.AI, it’s the must-try guide for anyone aiming to unlock your AI edge. Explore now!

10 tools
xix.ai
Prompt Best AI Prompt Libraries for ChatGPT Workflows
Best AI Prompt Libraries for ChatGPT Workflows

2026 Latest Best Top-Rated AI Prompt Libraries for optimizing all types of ChatGPT workflows. XIX.AI has curated a powerful, game-changing collection that goes through rigorous real-world tests to ensure top performance. You can find detailed free vs paid comparisons and expert rankings to help you choose the must-try tools that boost your productivity and unlock your AI edge. Explore now!

11 tools
xix.ai
Education and Learning AI Quiz Builder Platforms for Teachers, Tutors, and Cohort-Based Learning Programs
AI Quiz Builder Platforms for Teachers, Tutors, and Cohort-Based Learning Programs

2026 Latest Best AI Quiz Builder Platforms for Teachers, Tutors, and Cohort-Based Learning Programs! XIX.AI has curated a top-rated list of powerful game-changing tools that go through real-world tests to deliver accurate rankings. These must-try platforms help boost writing efficiency, streamline content creation, and simplify quiz design across all learning scenarios. Explore now to discover your perfect tool for unlocking your AI edge in teaching!

13 tools
xix.ai
code AI Pull Request Review Tools for GitHub Teams Handling Refactors, Bugs, and Security Gaps
AI Pull Request Review Tools for GitHub Teams Handling Refactors, Bugs, and Security Gaps

2026 Latest Best AI Pull Request Review Tools for GitHub Teams are here on XIX.AI! This top-rated curated list showcases powerful game-changing solutions that streamline refactoring, bug fixing, and security gap detection across all team workflows. Enjoy a free vs paid comparison along with real-world tests and detailed rankings to help you find the perfect tool that boosts productivity significantly. Explore now to unlock your AI edge!

12 tools
xix.ai
Comments (2)
0/500
MichaelMartinez
MichaelMartinez June 27, 2026 at 2:00:18 PM EDT

Interesting findings! I've always suspected those benchmarks were too good to be true. Real-world coding is messy, and it's no surprise that AI agents struggle with edge cases. 😅

GregoryRamirez
GregoryRamirez May 16, 2026 at 8:00:16 AM EDT

Interessant, aber irgendwie auch nicht überraschend. Benchmarks sind oft zu optimistisch, weil sie in einer kontrollierten Umgebung laufen. In der echten Welt mit Legacy-Code, unklaren Anforderungen und Teamarbeit sieht es dann anders aus. 🤔 Vielleicht sollten wir weniger auf die Marketing-Hypes hören und mehr auf praktische Tests setzen. Wer hat schon Erfahrung mit AI-Coding-Tools im Alltag gemacht?

OR