AI Code Performance Overstated in Real-World Tests, Study Finds
Research from the METR institute indicates that the widely used SWE-bench Verified benchmark for evaluating AI programming capabilities may significantly overstate AI agents' performance in real-world software development. The study found that approximately half of the AI-generated code solutions marked as "passed" by the benchmark would likely be rejected by actual project maintainers during code review, highlighting a substantial gap between automated assessment results and real-world code quality.
SWE-bench Verified has long been regarded as a key standard for evaluating AI-assisted software engineering, testing whether models can solve genuine programming tasks in open-source projects and verifying if code changes pass the project's automated test suite. Several AI companies, including Anthropic and OpenAI, frequently cite results from this benchmark to demonstrate model advancements.

In this study, the METR team enlisted four experienced developers who maintain the open-source projects scikit-learn, Sphinx, and pytest to manually review 296 pieces of AI-generated code. These code samples were produced by five different models: Claude 3.5 Sonnet, Claude 3.7 Sonnet, Claude 4 Opus, Claude 4.5 Sonnet, and GPT-5. The results revealed that the actual acceptance rate by maintainers was, on average, about 24 percentage points lower than the automated scores from SWE-bench—a statistically significant difference.
The study also determined that the rejected AI code was not primarily due to stylistic issues but rather more substantial engineering flaws. Maintainers categorized the problems into three main types: code quality failing to meet project specifications, disruption of the existing code structure, and fundamental functional errors. A significant portion of the cases involved functional errors where, despite passing automated tests, the code did not correctly solve the intended problem.
Regarding model comparison, the research found that upgrading from Claude 3.5 Sonnet to Claude 3.7 Sonnet significantly improved the benchmark pass rate, but the number of functional errors flagged by maintainers also increased. The transition from Claude 3.7 Sonnet to Claude 4 Opus saw a shift toward more code quality issues, while Claude 4.5 Sonnet showed improvements in code quality. In contrast, GPT-5 performed notably worse overall than the Anthropic model series in this manual assessment.

The research team also performed an estimated analysis of "task completion time": according to SWE-bench automated evaluation results, it would take Claude 4.5 Sonnet approximately 50 minutes of human effort to complete tasks with a 50% success rate. However, based on the maintainers' scores, the estimated time drops to only about 8 minutes, suggesting the benchmark may overestimate capabilities by up to sevenfold.
However, the researchers also emphasized that this study does not imply a fundamental limit on AI programming agents' capabilities. With improved prompting strategies, more human feedback, or multiple iteration cycles, the gap between automated evaluation and manual review could be reduced. Additionally, the experimental setup differs from real development processes—for instance, AI agents had only one submission attempt, whereas human developers can typically iteratively modify code based on feedback.
In summary, the study concludes that relying solely on benchmark scores to assess the practical utility of AI programming agents may introduce systematic bias. As AI coding models evolve rapidly, developing evaluation systems that better reflect real-world development environments has become a crucial research direction in AI software engineering.
Related article
South Korea Breaks Ground on National AI Computing Center, Investing 2.5 Trillion Won with 2028 Target
South Korean outlet EtNews reports that groundbreaking for the Korea AI Computing Center (KOACC) took place on August 3 at the Solar City data center park in Sunan, Jeollanam-do. Backed by a total investment of 2.5 trillion KRW (roughly 11.838 billio
Six Tech Giants Back Linux Foundation With $12.5M to Tackle AI Vulnerability Noise
To tackle the flood of low-quality security reports produced by AI automation tools, six major tech companies—Anthropic, Amazon (AWS), GitHub, Google, Microsoft, and OpenAI—have collectively contributed $12.5 million in funding to Linux Foundation in
Musk Considered Leaving OpenAI to His Kids as Altman Testifies
This morning, OpenAI CEO Sam Altman took the stand to address former co-founder Elon Musk’s lawsuit challenging the company’s corporate structure.When asked about Musk’s claim that other founders “stole a charity” by launching a for-profit subsidiary
Related Special Topic Recommendations
Comments (2)
0/500
Interesting findings! I've always suspected those benchmarks were too good to be true. Real-world coding is messy, and it's no surprise that AI agents struggle with edge cases. 😅
Interessant, aber irgendwie auch nicht überraschend. Benchmarks sind oft zu optimistisch, weil sie in einer kontrollierten Umgebung laufen. In der echten Welt mit Legacy-Code, unklaren Anforderungen und Teamarbeit sieht es dann anders aus. 🤔 Vielleicht sollten wir weniger auf die Marketing-Hypes hören und mehr auf praktische Tests setzen. Wer hat schon Erfahrung mit AI-Coding-Tools im Alltag gemacht?
Research from the METR institute indicates that the widely used SWE-bench Verified benchmark for evaluating AI programming capabilities may significantly overstate AI agents' performance in real-world software development. The study found that approximately half of the AI-generated code solutions marked as "passed" by the benchmark would likely be rejected by actual project maintainers during code review, highlighting a substantial gap between automated assessment results and real-world code quality.
SWE-bench Verified has long been regarded as a key standard for evaluating AI-assisted software engineering, testing whether models can solve genuine programming tasks in open-source projects and verifying if code changes pass the project's automated test suite. Several AI companies, including Anthropic and OpenAI, frequently cite results from this benchmark to demonstrate model advancements.

In this study, the METR team enlisted four experienced developers who maintain the open-source projects scikit-learn, Sphinx, and pytest to manually review 296 pieces of AI-generated code. These code samples were produced by five different models: Claude 3.5 Sonnet, Claude 3.7 Sonnet, Claude 4 Opus, Claude 4.5 Sonnet, and GPT-5. The results revealed that the actual acceptance rate by maintainers was, on average, about 24 percentage points lower than the automated scores from SWE-bench—a statistically significant difference.
The study also determined that the rejected AI code was not primarily due to stylistic issues but rather more substantial engineering flaws. Maintainers categorized the problems into three main types: code quality failing to meet project specifications, disruption of the existing code structure, and fundamental functional errors. A significant portion of the cases involved functional errors where, despite passing automated tests, the code did not correctly solve the intended problem.
Regarding model comparison, the research found that upgrading from Claude 3.5 Sonnet to Claude 3.7 Sonnet significantly improved the benchmark pass rate, but the number of functional errors flagged by maintainers also increased. The transition from Claude 3.7 Sonnet to Claude 4 Opus saw a shift toward more code quality issues, while Claude 4.5 Sonnet showed improvements in code quality. In contrast, GPT-5 performed notably worse overall than the Anthropic model series in this manual assessment.

The research team also performed an estimated analysis of "task completion time": according to SWE-bench automated evaluation results, it would take Claude 4.5 Sonnet approximately 50 minutes of human effort to complete tasks with a 50% success rate. However, based on the maintainers' scores, the estimated time drops to only about 8 minutes, suggesting the benchmark may overestimate capabilities by up to sevenfold.
However, the researchers also emphasized that this study does not imply a fundamental limit on AI programming agents' capabilities. With improved prompting strategies, more human feedback, or multiple iteration cycles, the gap between automated evaluation and manual review could be reduced. Additionally, the experimental setup differs from real development processes—for instance, AI agents had only one submission attempt, whereas human developers can typically iteratively modify code based on feedback.
In summary, the study concludes that relying solely on benchmark scores to assess the practical utility of AI programming agents may introduce systematic bias. As AI coding models evolve rapidly, developing evaluation systems that better reflect real-world development environments has become a crucial research direction in AI software engineering.
South Korea Breaks Ground on National AI Computing Center, Investing 2.5 Trillion Won with 2028 Target
South Korean outlet EtNews reports that groundbreaking for the Korea AI Computing Center (KOACC) took place on August 3 at the Solar City data center park in Sunan, Jeollanam-do. Backed by a total investment of 2.5 trillion KRW (roughly 11.838 billio
Six Tech Giants Back Linux Foundation With $12.5M to Tackle AI Vulnerability Noise
To tackle the flood of low-quality security reports produced by AI automation tools, six major tech companies—Anthropic, Amazon (AWS), GitHub, Google, Microsoft, and OpenAI—have collectively contributed $12.5 million in funding to Linux Foundation in
Musk Considered Leaving OpenAI to His Kids as Altman Testifies
This morning, OpenAI CEO Sam Altman took the stand to address former co-founder Elon Musk’s lawsuit challenging the company’s corporate structure.When asked about Musk’s claim that other founders “stole a charity” by launching a for-profit subsidiary
Interesting findings! I've always suspected those benchmarks were too good to be true. Real-world coding is messy, and it's no surprise that AI agents struggle with edge cases. 😅
Interessant, aber irgendwie auch nicht überraschend. Benchmarks sind oft zu optimistisch, weil sie in einer kontrollierten Umgebung laufen. In der echten Welt mit Legacy-Code, unklaren Anforderungen und Teamarbeit sieht es dann anders aus. 🤔 Vielleicht sollten wir weniger auf die Marketing-Hypes hören und mehr auf praktische Tests setzen. Wer hat schon Erfahrung mit AI-Coding-Tools im Alltag gemacht?





Home






