AI Models Make Breakthrough in Solving Advanced Math

Over the weekend, software engineer, former quantitative researcher, and startup founder Neel Somani was evaluating the mathematical capabilities of OpenAI's new model when he stumbled upon a surprising result. After inputting a problem into ChatGPT and allowing it to process for 15 minutes, he returned to find a complete solution. He reviewed the proof and verified it using a tool called Harmonic—and everything was correct.
"I wanted to set a baseline for understanding when large language models can effectively solve open mathematical problems versus where they still face challenges," Somani explained. The unexpected finding was that the latest model had pushed the boundary of what's possible.
ChatGPT's chain of reasoning was particularly striking, seamlessly applying mathematical principles like Legendre's formula, Bertrand's postulate, and the Star of David theorem. The model eventually referenced a 2013 Math Overflow post where Harvard mathematician Noam Elkies had provided an elegant solution to a related problem. However, ChatGPT's final proof differed from Elkies' work in key aspects, offering a more comprehensive solution to a problem variant originally posed by the legendary mathematician Paul Erdős. Erdős's extensive collection of unsolved problems has become a testing ground for AI capabilities.
For those skeptical of machine intelligence, this outcome is remarkable—and it's not an isolated case. AI tools have become pervasive in mathematics, ranging from formalization-focused LLMs like Harmonic's Aristotle to literature-review systems like OpenAI's deep research. Since the release of GPT-5.2—which Somani notes is "anecdotally more skilled at mathematical reasoning than previous versions"—the number of solved problems has grown significantly, raising new questions about the ability of large language models to expand the frontiers of human knowledge.
Somani was examining the Erdős problems, a set of over 1,000 conjectures by the Hungarian mathematician that are cataloged online. These problems, which vary widely in topic and difficulty, have become a prime target for AI-driven mathematical exploration. The first autonomous solutions emerged in November from a Gemini-powered model called AlphaEvolve, but more recently, Somani and others have observed that GPT-5.2 demonstrates exceptional proficiency with advanced mathematics.
Since Christmas, 15 problems have been updated from "open" to "solved" on the Erdős problem website—with 11 of those solutions explicitly crediting the involvement of AI models.
The esteemed mathematician Terence Tao offers a more detailed perspective on this progress on his GitHub page, noting eight distinct problems where AI models made substantial autonomous contributions, along with six other cases where progress was achieved by identifying and building upon prior research. While fully autonomous AI mathematical reasoning remains a distant goal, it is evident that large models are beginning to play a significant role.
Join the Disrupt 2026 Waitlist
Add yourself to the Disrupt 2026 waitlist to secure early access when Early Bird tickets are released. Past Disrupt events have featured leaders from Google Cloud, Netflix, Microsoft, Box, Phia, a16z, ElevenLabs, Wayve, Hugging Face, Elad Gil, and Vinod Khosla across our stages—part of over 250 industry experts leading more than 200 sessions designed to accelerate your growth and sharpen your competitive edge. Additionally, connect with hundreds of startups driving innovation across every industry.
Join the Disrupt 2026 Waitlist
Add yourself to the Disrupt 2026 waitlist to secure early access when Early Bird tickets are released. Past Disrupt events have featured leaders from Google Cloud, Netflix, Microsoft, Box, Phia, a16z, ElevenLabs, Wayve, Hugging Face, Elad Gil, and Vinod Khosla across our stages—part of over 250 industry experts leading more than 200 sessions designed to accelerate your growth and sharpen your competitive edge. Additionally, connect with hundreds of startups driving innovation across every industry.
On Mastodon, Tao suggested that the scalable nature of AI systems makes them "particularly well-suited for tackling the 'long tail' of lesser-known Erdős problems, many of which actually have straightforward solutions."
"Consequently, many of these more accessible Erdős problems are now more likely to be solved through purely AI-based methods than by human or hybrid approaches," Tao added.
Another contributing factor is a recent shift toward formalization—a detailed process that makes mathematical reasoning easier to verify and extend. While formalization doesn't inherently require AI or computers, a new generation of automated tools has significantly streamlined the workflow. The open-source "proof assistant" Lean, developed at Microsoft Research in 2013, has gained widespread adoption in the field for formalizing proofs. AI tools like Harmonic's Aristotle now promise to automate much of this formalization work.
For Harmonic founder Tudor Achim, the sudden increase in solved Erdős problems is less significant than the fact that leading mathematicians are beginning to take these tools seriously. "I'm more interested in the fact that mathematics and computer science professors are using [AI tools]," Achim stated. "These individuals have reputations to uphold, so when they confirm they use Aristotle or ChatGPT, that's meaningful validation."
Related article
ByteDance Boosts Core AI Incentives as Doubao Surges 14.6%
ByteDance recently convened a DouBao equity briefing to unveil fresh incentive policies for staff involved in the DouBao division. The strike price for DouBao shares has been lifted from $14.85 in June 2026 to $17.02, marking an approximate 14.6% inc
MiniMax Unveils 10x Team Program to Incentivize Global AI Experts
MiniMax (Xiyu Technology), the General Artificial Intelligence Lab, has officially launched "10x Team," a global talent collaboration initiative. This program aims to recruit top experts across industries to explore the deep application of large mode
South Korea Breaks Ground on National AI Computing Center, Investing 2.5 Trillion Won with 2028 Target
South Korean outlet EtNews reports that groundbreaking for the Korea AI Computing Center (KOACC) took place on August 3 at the Solar City data center park in Sunan, Jeollanam-do. Backed by a total investment of 2.5 trillion KRW (roughly 11.838 billio
Related Special Topic Recommendations
Comments (0)
0/500

Over the weekend, software engineer, former quantitative researcher, and startup founder Neel Somani was evaluating the mathematical capabilities of OpenAI's new model when he stumbled upon a surprising result. After inputting a problem into ChatGPT and allowing it to process for 15 minutes, he returned to find a complete solution. He reviewed the proof and verified it using a tool called Harmonic—and everything was correct.
"I wanted to set a baseline for understanding when large language models can effectively solve open mathematical problems versus where they still face challenges," Somani explained. The unexpected finding was that the latest model had pushed the boundary of what's possible.
ChatGPT's chain of reasoning was particularly striking, seamlessly applying mathematical principles like Legendre's formula, Bertrand's postulate, and the Star of David theorem. The model eventually referenced a 2013 Math Overflow post where Harvard mathematician Noam Elkies had provided an elegant solution to a related problem. However, ChatGPT's final proof differed from Elkies' work in key aspects, offering a more comprehensive solution to a problem variant originally posed by the legendary mathematician Paul Erdős. Erdős's extensive collection of unsolved problems has become a testing ground for AI capabilities.
For those skeptical of machine intelligence, this outcome is remarkable—and it's not an isolated case. AI tools have become pervasive in mathematics, ranging from formalization-focused LLMs like Harmonic's Aristotle to literature-review systems like OpenAI's deep research. Since the release of GPT-5.2—which Somani notes is "anecdotally more skilled at mathematical reasoning than previous versions"—the number of solved problems has grown significantly, raising new questions about the ability of large language models to expand the frontiers of human knowledge.
Somani was examining the Erdős problems, a set of over 1,000 conjectures by the Hungarian mathematician that are cataloged online. These problems, which vary widely in topic and difficulty, have become a prime target for AI-driven mathematical exploration. The first autonomous solutions emerged in November from a Gemini-powered model called AlphaEvolve, but more recently, Somani and others have observed that GPT-5.2 demonstrates exceptional proficiency with advanced mathematics.
Since Christmas, 15 problems have been updated from "open" to "solved" on the Erdős problem website—with 11 of those solutions explicitly crediting the involvement of AI models.
The esteemed mathematician Terence Tao offers a more detailed perspective on this progress on his GitHub page, noting eight distinct problems where AI models made substantial autonomous contributions, along with six other cases where progress was achieved by identifying and building upon prior research. While fully autonomous AI mathematical reasoning remains a distant goal, it is evident that large models are beginning to play a significant role.
Join the Disrupt 2026 Waitlist
Add yourself to the Disrupt 2026 waitlist to secure early access when Early Bird tickets are released. Past Disrupt events have featured leaders from Google Cloud, Netflix, Microsoft, Box, Phia, a16z, ElevenLabs, Wayve, Hugging Face, Elad Gil, and Vinod Khosla across our stages—part of over 250 industry experts leading more than 200 sessions designed to accelerate your growth and sharpen your competitive edge. Additionally, connect with hundreds of startups driving innovation across every industry.
Join the Disrupt 2026 Waitlist
Add yourself to the Disrupt 2026 waitlist to secure early access when Early Bird tickets are released. Past Disrupt events have featured leaders from Google Cloud, Netflix, Microsoft, Box, Phia, a16z, ElevenLabs, Wayve, Hugging Face, Elad Gil, and Vinod Khosla across our stages—part of over 250 industry experts leading more than 200 sessions designed to accelerate your growth and sharpen your competitive edge. Additionally, connect with hundreds of startups driving innovation across every industry.
On Mastodon, Tao suggested that the scalable nature of AI systems makes them "particularly well-suited for tackling the 'long tail' of lesser-known Erdős problems, many of which actually have straightforward solutions."
"Consequently, many of these more accessible Erdős problems are now more likely to be solved through purely AI-based methods than by human or hybrid approaches," Tao added.
Another contributing factor is a recent shift toward formalization—a detailed process that makes mathematical reasoning easier to verify and extend. While formalization doesn't inherently require AI or computers, a new generation of automated tools has significantly streamlined the workflow. The open-source "proof assistant" Lean, developed at Microsoft Research in 2013, has gained widespread adoption in the field for formalizing proofs. AI tools like Harmonic's Aristotle now promise to automate much of this formalization work.
For Harmonic founder Tudor Achim, the sudden increase in solved Erdős problems is less significant than the fact that leading mathematicians are beginning to take these tools seriously. "I'm more interested in the fact that mathematics and computer science professors are using [AI tools]," Achim stated. "These individuals have reputations to uphold, so when they confirm they use Aristotle or ChatGPT, that's meaningful validation."
ByteDance Boosts Core AI Incentives as Doubao Surges 14.6%
ByteDance recently convened a DouBao equity briefing to unveil fresh incentive policies for staff involved in the DouBao division. The strike price for DouBao shares has been lifted from $14.85 in June 2026 to $17.02, marking an approximate 14.6% inc
MiniMax Unveils 10x Team Program to Incentivize Global AI Experts
MiniMax (Xiyu Technology), the General Artificial Intelligence Lab, has officially launched "10x Team," a global talent collaboration initiative. This program aims to recruit top experts across industries to explore the deep application of large mode
South Korea Breaks Ground on National AI Computing Center, Investing 2.5 Trillion Won with 2028 Target
South Korean outlet EtNews reports that groundbreaking for the Korea AI Computing Center (KOACC) took place on August 3 at the Solar City data center park in Sunan, Jeollanam-do. Backed by a total investment of 2.5 trillion KRW (roughly 11.838 billio





Home






