Home
OpenAI Slams AI Benchmark: Nearly a Third of Questions Flawed, Pass Rate Soared to 80% in 8 Months
OpenAI has publicly challenged the industry's leading benchmark, SWE-Bench Pro, in a blog post, claiming that roughly 30% of its 731 public test tasks contain evaluation flaws. Developed by Scale AI, SWE-Bench Pro is designed to assess the coding abilities of large language models and AI agents. Because it closely mirrors real-world enterprise development and enforces strict anti-cheating measures, it has become a widely trusted benchmark in AI software engineering.

OpenAI highlights a key signal in the blog post: the pass rate for top-tier models on this benchmark jumped from 23.3% to 80.3% in just eight months. Such rapid progress seems suspicious, and OpenAI argues that the benchmark can no longer accurately measure a model's real-world software development skills. The issue likely stems from flaws in the evaluation itself, not a genuine leap in model capability.
Two review paths cross-validated reveal that nearly 30% of tasks are deemed "unqualified."
To verify this, OpenAI launched two parallel review processes. The data-point analysis uncovered 200 failing tasks, or 27.4% of the 731 public tasks. Meanwhile, manual annotation flagged 249 failing tasks, or 34.1%. Cross-referencing both methods, OpenAI estimates that roughly 30% of SWE-Bench Pro tasks contain defects, falling into four categories: overly strict tests, inadequate prompts, narrow test coverage, and misleading prompts.
OpenAI also shared a typical example: one task asked for a single space at the start of a line when converting content to Markdown, but the hidden test expected two spaces. As a result, even if the model followed the stated requirement, it would be marked incorrect. This mismatch between explicit instructions and hidden requirements directly skews the evaluation of a model's true ability, and explains the unreasonable surge in pass rates.
Withdrawal of adoption recommendation and call for rebuilding the AI evaluation system
Based on this analysis, OpenAI has officially retracted its earlier recommendation to use SWE-Bench Pro. The company believes that future benchmarks should be created by experienced software developers specifically for AI evaluation, rather than repurposing test logic designed for human developers. If an industry benchmark has nearly 30% of its tasks flawed, the entire AI evaluation system's credibility comes into question. Shifting the focus from score-chasing competitions to genuine engineering capability assessments may be the next critical step in AI software engineering evaluation.
Related article
Suno to Watermark Songs Amid Legal Battles
Suno, the platform enabling users to generate AI-created music, has unveiled new features to label platform-produced tracks, restrict downloads, and update community standards to curb unauthorized replicas. These updates arrive as Suno confronts mult
Musk Admits Grok Build Leaked User Code, Promises to Erase All Historical Data
Elon Musk directly addressed the privacy controversy surrounding Grok Build, beginning with a simple "True" to confirm the incident's validity. He pledged that all user data previously uploaded to SpaceXAI would be permanently erased, stating, "not a
U.S. Stocks Hit Historic Milestone as AI and Aerospace Giants Prepare for Trillion-Dollar Debut
Elon Musk, Sam Altman, and Dario Amodei, three titans of the technology sector, are advancing toward initial public offerings for their respective ventures. With SpaceX, OpenAI, and Anthropic—three industry behemoths nearing trillion-dollar valuation
Related Special Topic Recommendations
Comments (0)
0/500
OpenAI has publicly challenged the industry's leading benchmark, SWE-Bench Pro, in a blog post, claiming that roughly 30% of its 731 public test tasks contain evaluation flaws. Developed by Scale AI, SWE-Bench Pro is designed to assess the coding abilities of large language models and AI agents. Because it closely mirrors real-world enterprise development and enforces strict anti-cheating measures, it has become a widely trusted benchmark in AI software engineering.

OpenAI highlights a key signal in the blog post: the pass rate for top-tier models on this benchmark jumped from 23.3% to 80.3% in just eight months. Such rapid progress seems suspicious, and OpenAI argues that the benchmark can no longer accurately measure a model's real-world software development skills. The issue likely stems from flaws in the evaluation itself, not a genuine leap in model capability.
Two review paths cross-validated reveal that nearly 30% of tasks are deemed "unqualified."
To verify this, OpenAI launched two parallel review processes. The data-point analysis uncovered 200 failing tasks, or 27.4% of the 731 public tasks. Meanwhile, manual annotation flagged 249 failing tasks, or 34.1%. Cross-referencing both methods, OpenAI estimates that roughly 30% of SWE-Bench Pro tasks contain defects, falling into four categories: overly strict tests, inadequate prompts, narrow test coverage, and misleading prompts.
OpenAI also shared a typical example: one task asked for a single space at the start of a line when converting content to Markdown, but the hidden test expected two spaces. As a result, even if the model followed the stated requirement, it would be marked incorrect. This mismatch between explicit instructions and hidden requirements directly skews the evaluation of a model's true ability, and explains the unreasonable surge in pass rates.
Withdrawal of adoption recommendation and call for rebuilding the AI evaluation system
Based on this analysis, OpenAI has officially retracted its earlier recommendation to use SWE-Bench Pro. The company believes that future benchmarks should be created by experienced software developers specifically for AI evaluation, rather than repurposing test logic designed for human developers. If an industry benchmark has nearly 30% of its tasks flawed, the entire AI evaluation system's credibility comes into question. Shifting the focus from score-chasing competitions to genuine engineering capability assessments may be the next critical step in AI software engineering evaluation.
Suno to Watermark Songs Amid Legal Battles
Suno, the platform enabling users to generate AI-created music, has unveiled new features to label platform-produced tracks, restrict downloads, and update community standards to curb unauthorized replicas. These updates arrive as Suno confronts mult
Musk Admits Grok Build Leaked User Code, Promises to Erase All Historical Data
Elon Musk directly addressed the privacy controversy surrounding Grok Build, beginning with a simple "True" to confirm the incident's validity. He pledged that all user data previously uploaded to SpaceXAI would be permanently erased, stating, "not a
U.S. Stocks Hit Historic Milestone as AI and Aerospace Giants Prepare for Trillion-Dollar Debut
Elon Musk, Sam Altman, and Dario Amodei, three titans of the technology sector, are advancing toward initial public offerings for their respective ventures. With SpaceX, OpenAI, and Anthropic—three industry behemoths nearing trillion-dollar valuation











