Home
Washington State University Study Reveals ChatGPT's Major Contradictions in Complex Scientific Judgments

A recent study from Washington State University (WSU) reveals that while ChatGPT responds with confidence, its handling of complex scientific statements resembles random guessing. The research highlights not only limited accuracy but also frequent contradictory answers to the same question.
Professor Mesut Cicek and his team extracted 719 research hypotheses from business journals published since 2021 and repeatedly submitted them to the model for truth verification.
Although ChatGPT's surface accuracy appears around 80%, after accounting for random guessing, its actual performance is only about 60% better than a 50% coin flip. Researchers rated this as a "low D-grade score." The model performed especially poorly at identifying false statements, correctly judging only 16.4% of false propositions.
The researchers submitted each hypothesis to the model ten times and found it struggled to maintain consistent positions:
Answer Fluctuation: In about 73% of cases, the model maintained consistent conclusions across ten repetitions.
Extreme Contradictions: In some cases, the model alternated between "true" and "false" answers, with extreme instances where half were true and half false, despite using the exact same prompt.
The study notes that users are easily misled by AI's fluent and persuasive language, but this does not indicate genuine reasoning ability.
No True Understanding: The model relies on memory and pattern matching, unlike humans who genuinely understand the world and what they are saying.
Limited Version Progress: Testing showed that the updated ChatGPT‑5 mini (tested in 2025) performed similarly to earlier versions on this specific task, with no significant improvements.
Based on these findings, Cicek advises business managers to maintain a high degree of skepticism when making complex decisions. They should not regard generative AI as an "authority" that can replace professional judgment and must manually verify all outputs. Organizations should enhance training to help employees understand both the strengths and limitations of AI tools, preventing decision biases from blind trust.
This study serves as another reminder that despite rapid AI development, the technology's deep logical reasoning and evidence evaluation capabilities still require improvement.
Related article
Swedish AI Startup Lovable Eyes $13.2 Billion Valuation After Major Funding Round
As AI-driven coding tools gain traction, Swedish startup Lovable has secured a major funding round. The company aims to raise $3 billion, potentially boosting its valuation to $13.2 billion—double the $6.6 billion recorded last December. Menlo Ventur
Google Tests Remy AI Agent for Gemini as Focus Shifts to User Control
According to Business Insider, Google is testing Remy, a new AI personal agent for Gemini. This tool aims to execute tasks on behalf of users, streamlining both professional workflows and daily routines.Currently, Remy is undergoing testing in an int
How to fix Core Web Vitals for better SEO rankings
Streamline Report Card Comments with AI ToolsIntroductionAI Tools for Generating Report Card CommentsMagic SchoolAlmanac AIChat GPTUsing Magic School to Generate Report Card CommentsLogging into Magic SchoolSelecting the Report Card Comments ToolCust
Related Special Topic Recommendations
Comments (0)
0/500

A recent study from Washington State University (WSU) reveals that while ChatGPT responds with confidence, its handling of complex scientific statements resembles random guessing. The research highlights not only limited accuracy but also frequent contradictory answers to the same question.
Professor Mesut Cicek and his team extracted 719 research hypotheses from business journals published since 2021 and repeatedly submitted them to the model for truth verification.
Although ChatGPT's surface accuracy appears around 80%, after accounting for random guessing, its actual performance is only about 60% better than a 50% coin flip. Researchers rated this as a "low D-grade score." The model performed especially poorly at identifying false statements, correctly judging only 16.4% of false propositions.
The researchers submitted each hypothesis to the model ten times and found it struggled to maintain consistent positions:
Answer Fluctuation: In about 73% of cases, the model maintained consistent conclusions across ten repetitions.
Extreme Contradictions: In some cases, the model alternated between "true" and "false" answers, with extreme instances where half were true and half false, despite using the exact same prompt.
The study notes that users are easily misled by AI's fluent and persuasive language, but this does not indicate genuine reasoning ability.
No True Understanding: The model relies on memory and pattern matching, unlike humans who genuinely understand the world and what they are saying.
Limited Version Progress: Testing showed that the updated ChatGPT‑5 mini (tested in 2025) performed similarly to earlier versions on this specific task, with no significant improvements.
Based on these findings, Cicek advises business managers to maintain a high degree of skepticism when making complex decisions. They should not regard generative AI as an "authority" that can replace professional judgment and must manually verify all outputs. Organizations should enhance training to help employees understand both the strengths and limitations of AI tools, preventing decision biases from blind trust.
This study serves as another reminder that despite rapid AI development, the technology's deep logical reasoning and evidence evaluation capabilities still require improvement.
Swedish AI Startup Lovable Eyes $13.2 Billion Valuation After Major Funding Round
As AI-driven coding tools gain traction, Swedish startup Lovable has secured a major funding round. The company aims to raise $3 billion, potentially boosting its valuation to $13.2 billion—double the $6.6 billion recorded last December. Menlo Ventur
How to fix Core Web Vitals for better SEO rankings
Streamline Report Card Comments with AI ToolsIntroductionAI Tools for Generating Report Card CommentsMagic SchoolAlmanac AIChat GPTUsing Magic School to Generate Report Card CommentsLogging into Magic SchoolSelecting the Report Card Comments ToolCust











