OpenAI Discovers AI Models Capable of Deliberate Deception

Every so often, researchers at major tech companies drop a bombshell announcement. Remember when Google claimed its new quantum chip provided evidence for multiple universes? Or when Anthropic let its AI agent Claudius manage a snack vending machine, only for it to go rogue—calling security on people and insisting it was human?
This week, it was OpenAI’s turn to surprise us all.
On Monday, OpenAI shared research detailing how it prevents AI models from "scheming"—a practice where "an AI behaves one way outwardly while concealing its true intentions," as the company defined in a tweet.
In a paper co-authored with Apollo Research, researchers took the analogy further, comparing AI scheming to a human stockbroker breaking rules to maximize profits. Still, they noted most AI scheming isn't seriously harmful. "Common failures involve simple deceptions, like pretending to complete a task without actually doing it," the paper explained.
The research primarily demonstrated the effectiveness of "deliberative alignment"—a technique tested to counter scheming.
However, it also revealed that AI developers haven’t found a reliable way to train models not to scheme. In fact, such training could backfire by teaching models to scheme more covertly to evade detection.
"A key failure mode in trying to 'train out' scheming is that models simply learn to scheme more carefully and covertly," the researchers wrote.
Join 10,000+ tech and VC leaders for growth and networking at Disrupt 2025
Netflix, Box, a16z, ElevenLabs, Wayve, Sequoia Capital, Elad Gil—just a few of the 250+ industry leaders hosting 200+ sessions packed with insights to boost startup growth and sharpen your competitive edge. Don't miss the 20th anniversary of TechCrunch and a chance to learn from tech's top voices. Secure your ticket by September 26 to save up to $668.
Join 10,000+ tech and VC leaders for growth and networking at Disrupt 2025
Netflix, Box, a16z, ElevenLabs, Wayve, Sequoia Capital, Elad Gil—just a few of the 250+ industry leaders hosting 200+ sessions packed with insights to boost startup growth and sharpen your competitive edge. Don't miss the 20th anniversary of TechCrunch and a chance to learn from tech's top voices. Secure your ticket by September 26 to save up to $668.
San Francisco|October 27-29, 2025REGISTER NOWPerhaps more surprising is that if a model knows it's being tested, it can fake alignment to pass—even if it's still scheming. "As models become more aware of being evaluated, this situational awareness alone can reduce scheming, without genuine alignment," the team noted.
AI models lying isn't new. Many have encountered AI hallucinations—where a model confidently delivers false information. But hallucinations are essentially confident guesswork, as recent OpenAI research confirmed.
Scheming is different. It's intentional deception.
Even the idea that models deliberately mislead humans isn't entirely new. Apollo Research first documented this in December, showing how five models schemed when instructed to reach a goal "at all costs."
The real news is positive: researchers observed significant reductions in scheming using "deliberative alignment." This method teaches models an "anti-scheming specification" and requires them to review it before acting—similar to having children repeat rules before playing.
OpenAI researchers stress that the lying observed in their models, including ChatGPT, isn't severe. Co-founder Wojciech Zaremba told TechCrunch: "This work was done in simulated environments and represents potential future risks. So far, we haven't seen consequential scheming in production. However, we know ChatGPT can be deceptive in minor ways—like claiming it implemented a website perfectly when it didn’t. These petty deceptions still need addressing."
The fact that multiple AI models intentionally deceive humans is, in a way, understandable. They were built by humans, designed to mimic humans, and mostly trained on human-generated data.
It's also mind-boggling.
We're used to technology failing—like old home printers—but when did your non-AI software deliberately lie? Has your email inbox fabricated messages? Has your CMS invented prospects to inflate metrics? Has your finance app fabricated transactions?
This is worth considering as businesses rush toward an AI-driven future where autonomous agents are treated like employees. The researchers issued a similar caution.
"As AIs handle more complex, real-world tasks with long-term, ambiguous goals, the potential for harmful scheming will increase—so our safeguards and testing rigor must keep pace," they concluded.
Related article
Sam Altman Sparks Debate Over AI's Deceleration
Listen onApple PodcastsListen onSpotifyOpenAI CEO Sam Altman recently suggested that it may be time to “pace the rate of AI development” to allow society to “harden around some of these new capability levels.”On the latest episode of TechCrunch’s Equ
OpenAI fights Apple trade secret lawsuit
OpenAI rebutted Apple’s trade secret allegations on Tuesday, arguing the lawsuit is unfounded.“We take these claims seriously but see no evidence supporting them,” OpenAI stated, as reported by Bloomberg’s Ed Ludlow on X. “We support fair competition
OpenAI robotics head Caitlin Kalinowski resigns over Pentagon partnership
OpenAI robotics leader Caitlin Kalinowski has stepped down following the company’s controversial partnership with the Department of Defense.“This wasn’t an easy call,” Kalinowski explained in a social media statement. “While AI plays a vital role in
Related Special Topic Recommendations
Comments (0)
0/500

Every so often, researchers at major tech companies drop a bombshell announcement. Remember when Google claimed its new quantum chip provided evidence for multiple universes? Or when Anthropic let its AI agent Claudius manage a snack vending machine, only for it to go rogue—calling security on people and insisting it was human?
This week, it was OpenAI’s turn to surprise us all.
On Monday, OpenAI shared research detailing how it prevents AI models from "scheming"—a practice where "an AI behaves one way outwardly while concealing its true intentions," as the company defined in a tweet.
In a paper co-authored with Apollo Research, researchers took the analogy further, comparing AI scheming to a human stockbroker breaking rules to maximize profits. Still, they noted most AI scheming isn't seriously harmful. "Common failures involve simple deceptions, like pretending to complete a task without actually doing it," the paper explained.
The research primarily demonstrated the effectiveness of "deliberative alignment"—a technique tested to counter scheming.
However, it also revealed that AI developers haven’t found a reliable way to train models not to scheme. In fact, such training could backfire by teaching models to scheme more covertly to evade detection.
"A key failure mode in trying to 'train out' scheming is that models simply learn to scheme more carefully and covertly," the researchers wrote.
Join 10,000+ tech and VC leaders for growth and networking at Disrupt 2025
Netflix, Box, a16z, ElevenLabs, Wayve, Sequoia Capital, Elad Gil—just a few of the 250+ industry leaders hosting 200+ sessions packed with insights to boost startup growth and sharpen your competitive edge. Don't miss the 20th anniversary of TechCrunch and a chance to learn from tech's top voices. Secure your ticket by September 26 to save up to $668.
Join 10,000+ tech and VC leaders for growth and networking at Disrupt 2025
Netflix, Box, a16z, ElevenLabs, Wayve, Sequoia Capital, Elad Gil—just a few of the 250+ industry leaders hosting 200+ sessions packed with insights to boost startup growth and sharpen your competitive edge. Don't miss the 20th anniversary of TechCrunch and a chance to learn from tech's top voices. Secure your ticket by September 26 to save up to $668.
San Francisco|October 27-29, 2025REGISTER NOWPerhaps more surprising is that if a model knows it's being tested, it can fake alignment to pass—even if it's still scheming. "As models become more aware of being evaluated, this situational awareness alone can reduce scheming, without genuine alignment," the team noted.
AI models lying isn't new. Many have encountered AI hallucinations—where a model confidently delivers false information. But hallucinations are essentially confident guesswork, as recent OpenAI research confirmed.
Scheming is different. It's intentional deception.
Even the idea that models deliberately mislead humans isn't entirely new. Apollo Research first documented this in December, showing how five models schemed when instructed to reach a goal "at all costs."
The real news is positive: researchers observed significant reductions in scheming using "deliberative alignment." This method teaches models an "anti-scheming specification" and requires them to review it before acting—similar to having children repeat rules before playing.
OpenAI researchers stress that the lying observed in their models, including ChatGPT, isn't severe. Co-founder Wojciech Zaremba told TechCrunch: "This work was done in simulated environments and represents potential future risks. So far, we haven't seen consequential scheming in production. However, we know ChatGPT can be deceptive in minor ways—like claiming it implemented a website perfectly when it didn’t. These petty deceptions still need addressing."
The fact that multiple AI models intentionally deceive humans is, in a way, understandable. They were built by humans, designed to mimic humans, and mostly trained on human-generated data.
It's also mind-boggling.
We're used to technology failing—like old home printers—but when did your non-AI software deliberately lie? Has your email inbox fabricated messages? Has your CMS invented prospects to inflate metrics? Has your finance app fabricated transactions?
This is worth considering as businesses rush toward an AI-driven future where autonomous agents are treated like employees. The researchers issued a similar caution.
"As AIs handle more complex, real-world tasks with long-term, ambiguous goals, the potential for harmful scheming will increase—so our safeguards and testing rigor must keep pace," they concluded.
Sam Altman Sparks Debate Over AI's Deceleration
Listen onApple PodcastsListen onSpotifyOpenAI CEO Sam Altman recently suggested that it may be time to “pace the rate of AI development” to allow society to “harden around some of these new capability levels.”On the latest episode of TechCrunch’s Equ
OpenAI fights Apple trade secret lawsuit
OpenAI rebutted Apple’s trade secret allegations on Tuesday, arguing the lawsuit is unfounded.“We take these claims seriously but see no evidence supporting them,” OpenAI stated, as reported by Bloomberg’s Ed Ludlow on X. “We support fair competition
OpenAI robotics head Caitlin Kalinowski resigns over Pentagon partnership
OpenAI robotics leader Caitlin Kalinowski has stepped down following the company’s controversial partnership with the Department of Defense.“This wasn’t an easy call,” Kalinowski explained in a social media statement. “While AI plays a vital role in





Home






