• Wion
  • /Trending
  • /Four AI models fail Math test they could not cheat on, humans score a perfect 10

Four AI models fail Math test they could not cheat on, humans score a perfect 10

Four AI models fail Math test they could not cheat on, humans score a perfect 10

Expert level mathematicians had solved all 10 problems but never published the results. AI failed the test since there were no references available. Photograph: (Unsplash)

Story highlights

AI failed to beat humans in 10 Math problems that expert mathematicians had solved in the past. Four systems entered a test, which was assessed by 30 analysts, and none of them was able to solve all 10 problems. The finding shows humans are still better than AI on new Math problems.

Artificial Intelligence systems might have solved an 80-year-old Math problem, but a test designed to evaluate it against top-tier academicians has revealed that it still has a long way to go when it comes to new and complex research-level mathematics. Four AI systems failed the test, which is part of "First Proof," a rigorous new AI benchmarking mechanism. The AI systems were given 10 research-level maths problems, and their answers were assessed by a jury of anonymous human specialists. To prevent AI from simply scanning through its training data, all the questions were new and had never been published anywhere. They were provided by 10 researchers from a broad range of mathematical specialities who had solved them at some point in their careers. A group of 30 mathematicians checked the answers, and the results proved that AI does not come close to human mathematicians in novel problems. The findings were published in the Nature journal.

The First Proof team carried out the test in an entirely autonomous way without humans being involved in any way. The AI systems also had to be publicly available. This ruled out Google’s Aletheia, which has been designed specifically for solving maths problems, and the full, unreleased version of Claude Mythos. OpenAI entered its ChatGPT 5.5 Pro model. Three academic groups provided the rest of the models - from the University of California, Los Angeles (UCLA); Princeton University in New Jersey; and the Swiss Federal Institute of Technology (ETH) in Zurich. They built "harnesses"—automated systems that prompted existing chatbots, such as ChatGPT, Gemini, and Claude, to repeatedly check the answers.

Add WION as a Preferred Source

AI vs human mathematicians: The results


The top-performing AI system was able to solve only 6 out of 10 problems. Since all of them had been solved by mathematicians, this proved that human expertise in the field still outperforms AI on unseen, highly rigorous research problems. ETH Zurich stood first, whose harness had used ChatGPT to generate answers. UCLA was second (ChatGPT harness), followed by OpenAI (standalone ChatGPT 5.5 Pro), and Princeton (using a Gemini 3.1 Pro harness).

Trending Stories

The results show that AI models have a long way to go before they can function autonomously as reliable research assistants, proof-checkers, and problem-solvers for human mathematicians. First Proof will come up with more versions of the test in future to correctly decide when this could actually become possible.

About the Author

Share on twitter

Anamica Singh

Anamica Singh is a Senior News Editor at WION, bringing over 17 years of deep media and journalism experience to the platform. Specialising in high-impact global journalism, she le...Read More