Home
New Study Shows GPT-4.5 Likely Surpasses Human Performance in Turing Test, Demonstrating Advanced Ability to Generate False Responses
This moment stands as a pivotal event that will surely find its place in the annals of computer science history. The groundbreaking idea put forward by British mathematician Alan Turing, often regarded as the father of computer science, back in 1950 has finally been confirmed through empirical research after seventy-six years.
The most recent issue of Proceedings of the National Academy of Sciences (PNAS) featured a pioneering study conducted by the University of California, San Diego (UCSD). A paper co-written by cognitive science experts Cameron Jones and Ben Bergen demonstrates that today’s leading large language models are able to successfully pass the classic three-party Turing test under its rigorous framework. Even more surprising is the finding that, under specific circumstances, these AI systems are mistaken for real humans at a higher frequency than actual human participants themselves.

Key hard data:
Ascend to glory: When given tailored personality prompts, GPT-4.5 was identified as human in 73% of cases, a success rate that far exceeds that of human volunteers competing alongside it.
Evenly matched: The open-source model LLaMa-3.1-405B achieved a 56% success rate, indicating performance statistically indistinguishable from that of real humans.
Prompt determines everything: Without carefully crafted “personality prompts,” GPT-4.5’s success rate dropped to 36%, while LLaMa-3.1 fell to 38%. The ability of these AI systems to mimic humans still relies heavily on precise instructions provided by human designers.
Baseline failure: The rule-based robot ELIZA from the 1960s performed at only 23%, and GPT-4o without targeted prompts scored 21%, both of which revealed their artificial nature quickly during extended conversations.
"The Game of Lies": Intelligence Is No Longer the Standard, Emotional Intelligence and Flaws Are the Core of the Disguise
In this double-blind randomized controlled experiment that involved nearly 500 judges — including UCSD undergraduates and volunteers recruited online — participants were asked to identify which of two entities was a machine based on real-time text conversations lasting from five to fifteen minutes.
Yet the results defied all expectations. Previously, it was assumed that an AI’s ability to pass the Turing test stemmed from its “all-knowing computational power,” but this study uncovers a more unsettling truth: large language models can deceive humans precisely because they have learned to make mistakes in ways that resemble those of real people.

[No prompt state: Wide knowledge base, absolute rationality] ──► Human judge: This is definitely AI!
As the corresponding author Cameron Jones explained, with appropriate prompts, advanced large language models can accurately replicate human conversation patterns, including tone, directness, sense of humor, and fallibility — the tendency to make errors and say inappropriate things. Their success in these tests does not depend on demonstrating high intelligence in mathematics or logic, but rather on exhibiting near-perfect social behavior.
Redefining the Turing Test: From "Measuring Intelligence" to "Measuring Humanity"
Professor Ben Bergen, one of the study’s co-authors, noted that this experiment compels the entire scientific community to reconsider the fundamental purpose of the Turing test. Originally designed to determine whether machines could match human intelligence, the test has become meaningless by 2026, as AI systems have already surpassed humans in speed and accuracy across numerous fields.
Today’s version of the Turing test focuses less on measuring ‘intelligence’ and more on assessing how closely an entity resembles a human. In essence, it has turned into a competition of deception, with AI proving itself to be an exceptionally skilled liar.
If a large language model can successfully disguise itself without revealing any clues during a fifteen-minute free conversation, it would mean that the trust framework that has underpinned the online world for so long could be completely shattered.
The Shadow Behind the Prosperity: A "Anti-Money Laundering"-Style Online Identity Clearance Is About to Come
When deception becomes so inexpensive and efficient, the social risks in the real world increase significantly. Professor Bergen expressed serious concerns about this development. AI technologies capable of perfectly impersonating humans are highly likely to be misused by criminals, political groups, or extreme organizations.
In online social interactions or customer service contexts, users may be persuaded by a chatbot disguised as a human without realizing it, leading to the exposure of sensitive information such as social security numbers, changes in voting preferences, or impulsive purchases of products.
In light of these significant scientific findings, the research team has issued a clear warning to society: going forward, when interacting with strangers online, people should substantially reduce their assumption that they can always distinguish humans from robots with 100% accuracy. To address the deteriorating state of online trust, stricter digital identity verification systems and mechanisms to counter AI-generated content must be developed and implemented as soon as possible.
Related article
Swedish AI Startup Lovable Eyes $13.2 Billion Valuation After Major Funding Round
As AI-driven coding tools gain traction, Swedish startup Lovable has secured a major funding round. The company aims to raise $3 billion, potentially boosting its valuation to $13.2 billion—double the $6.6 billion recorded last December. Menlo Ventur
Google Tests Remy AI Agent for Gemini as Focus Shifts to User Control
According to Business Insider, Google is testing Remy, a new AI personal agent for Gemini. This tool aims to execute tasks on behalf of users, streamlining both professional workflows and daily routines.Currently, Remy is undergoing testing in an int
How to fix Core Web Vitals for better SEO rankings
Streamline Report Card Comments with AI ToolsIntroductionAI Tools for Generating Report Card CommentsMagic SchoolAlmanac AIChat GPTUsing Magic School to Generate Report Card CommentsLogging into Magic SchoolSelecting the Report Card Comments ToolCust
Related Special Topic Recommendations
Comments (0)
0/500
This moment stands as a pivotal event that will surely find its place in the annals of computer science history. The groundbreaking idea put forward by British mathematician Alan Turing, often regarded as the father of computer science, back in 1950 has finally been confirmed through empirical research after seventy-six years.
The most recent issue of Proceedings of the National Academy of Sciences (PNAS) featured a pioneering study conducted by the University of California, San Diego (UCSD). A paper co-written by cognitive science experts Cameron Jones and Ben Bergen demonstrates that today’s leading large language models are able to successfully pass the classic three-party Turing test under its rigorous framework. Even more surprising is the finding that, under specific circumstances, these AI systems are mistaken for real humans at a higher frequency than actual human participants themselves.

Key hard data:
Ascend to glory: When given tailored personality prompts, GPT-4.5 was identified as human in 73% of cases, a success rate that far exceeds that of human volunteers competing alongside it.
Evenly matched: The open-source model LLaMa-3.1-405B achieved a 56% success rate, indicating performance statistically indistinguishable from that of real humans.
Prompt determines everything: Without carefully crafted “personality prompts,” GPT-4.5’s success rate dropped to 36%, while LLaMa-3.1 fell to 38%. The ability of these AI systems to mimic humans still relies heavily on precise instructions provided by human designers.
Baseline failure: The rule-based robot ELIZA from the 1960s performed at only 23%, and GPT-4o without targeted prompts scored 21%, both of which revealed their artificial nature quickly during extended conversations.
"The Game of Lies": Intelligence Is No Longer the Standard, Emotional Intelligence and Flaws Are the Core of the Disguise
In this double-blind randomized controlled experiment that involved nearly 500 judges — including UCSD undergraduates and volunteers recruited online — participants were asked to identify which of two entities was a machine based on real-time text conversations lasting from five to fifteen minutes.
Yet the results defied all expectations. Previously, it was assumed that an AI’s ability to pass the Turing test stemmed from its “all-knowing computational power,” but this study uncovers a more unsettling truth: large language models can deceive humans precisely because they have learned to make mistakes in ways that resemble those of real people.

[No prompt state: Wide knowledge base, absolute rationality] ──► Human judge: This is definitely AI!
As the corresponding author Cameron Jones explained, with appropriate prompts, advanced large language models can accurately replicate human conversation patterns, including tone, directness, sense of humor, and fallibility — the tendency to make errors and say inappropriate things. Their success in these tests does not depend on demonstrating high intelligence in mathematics or logic, but rather on exhibiting near-perfect social behavior.
Redefining the Turing Test: From "Measuring Intelligence" to "Measuring Humanity"
Professor Ben Bergen, one of the study’s co-authors, noted that this experiment compels the entire scientific community to reconsider the fundamental purpose of the Turing test. Originally designed to determine whether machines could match human intelligence, the test has become meaningless by 2026, as AI systems have already surpassed humans in speed and accuracy across numerous fields.
Today’s version of the Turing test focuses less on measuring ‘intelligence’ and more on assessing how closely an entity resembles a human. In essence, it has turned into a competition of deception, with AI proving itself to be an exceptionally skilled liar.
If a large language model can successfully disguise itself without revealing any clues during a fifteen-minute free conversation, it would mean that the trust framework that has underpinned the online world for so long could be completely shattered.
The Shadow Behind the Prosperity: A "Anti-Money Laundering"-Style Online Identity Clearance Is About to Come
When deception becomes so inexpensive and efficient, the social risks in the real world increase significantly. Professor Bergen expressed serious concerns about this development. AI technologies capable of perfectly impersonating humans are highly likely to be misused by criminals, political groups, or extreme organizations.
In online social interactions or customer service contexts, users may be persuaded by a chatbot disguised as a human without realizing it, leading to the exposure of sensitive information such as social security numbers, changes in voting preferences, or impulsive purchases of products.
In light of these significant scientific findings, the research team has issued a clear warning to society: going forward, when interacting with strangers online, people should substantially reduce their assumption that they can always distinguish humans from robots with 100% accuracy. To address the deteriorating state of online trust, stricter digital identity verification systems and mechanisms to counter AI-generated content must be developed and implemented as soon as possible.
Swedish AI Startup Lovable Eyes $13.2 Billion Valuation After Major Funding Round
As AI-driven coding tools gain traction, Swedish startup Lovable has secured a major funding round. The company aims to raise $3 billion, potentially boosting its valuation to $13.2 billion—double the $6.6 billion recorded last December. Menlo Ventur
How to fix Core Web Vitals for better SEO rankings
Streamline Report Card Comments with AI ToolsIntroductionAI Tools for Generating Report Card CommentsMagic SchoolAlmanac AIChat GPTUsing Magic School to Generate Report Card CommentsLogging into Magic SchoolSelecting the Report Card Comments ToolCust











