WayToClawEarn
High impactGowers's Weblog + Hacker News

ChatGPT 5.5 Pro solves doctoral-level math problems in one hour: Personal test by Fields Medal winner

Fields Medal winner Timothy Gowers personally tested ChatGPT 5.5 Pro: he independently completed doctoral-level mathematical research in one hour, completely rewriting the evaluation of AI mathematical capabilities.

WayToClawEarn EditorialPublished May 9, 2026Updated Aug 8, 2026

Editorial review of public sources · AI-assisted drafting. How we work · Original source

Core conclusion

On May 8, 2026, Fields Medal winner Timothy Gowers, a professor at the University of Cambridge, publicly shared his actual testing experience with ChatGPT 5.5 Pro: This model completed a doctoral-level mathematical research in only about an hour, solving an incompletely described combinatorial problem in number theory, and the entire process required almost no substantive mathematical input from Gowers himself.

DimensionsData
Test timeMay 8, 2026
TesterTimothy Gowers (Fields Medal Winner, 1998)
ModelChatGPT 5.5 Pro
Time takenAbout 1 hour
Solved problemSumset size possibility problem in additive number theory
Previous LLM judgments in the same fieldThere were doubts, and the assessment was significantly raised after the test

Key Points

  • Gowers, who had previously been cautious about LLM mathematics ability, announced a "significant increase in assessment" after this test.
  • The model solves a problem posed in Nathanson's paper that has not yet been fully described by human mathematicians
  • AI not only finds solutions, but also gives clear mathematical arguments
  • This marks the transition of AI from "piecing together known knowledge" to "producing substantive original mathematics"

Background: From Doubt to Shock

Timothy Gowers is a professor in the Department of Pure Mathematics and Mathematical Statistics at the University of Cambridge. He won the Fields Medal in 1998 for his work linking functional analysis and combinatorics. He has long been cautious and even skeptical about the mathematical capabilities of LLM.

Gowers previously observed that the "open problems" that LLM can solve often have ready-made answers hidden in the literature, or are very easy to deduce from known results. But he admitted in his blog:

"The laughter is getting smaller and smaller."

He noticed that the mathematician community has begun to realize that if there is a simple argument for an open problem that "human beings have not had time to notice", LLM has a high probability of discovering it. On the contrary, those seemingly "smart" arguments are often just a recombination of existing knowledge - and this is the essence of a large amount of human mathematical work.

Test Design: Nathanson’s Additive Number Theory Problem

Gowers chose a paper by Mel Nathanson "Diversity, Equity and Inclusion for Problems in Additive Number Theory" as test material. This paper poses a series of open questions about sumsets.

Mathematical nature of the problem

If $A$ is a set of integers, then its sumset is defined as $A + A = {a + b : a, b \in A}$. For a positive integer $h$, $h$-fold sumset is denoted as $hA$.

The question that interests Nathanson is: given $|A| = k$, what values ​​are possible for $|hA|$? That is to say, define the set $\mathcal{M}(h, k) = {|hA| : |A| = k}$, then what is $\mathcal{M}(h, k)$ specifically?

When $h = 2$, the answer is any integer between $k$ and $2k-1$ - this is a simple conclusion to the exercise. But when $h$ is larger, $\mathcal{M}(h, k)$ does not contain all values ​​between its minimum and maximum values, which human mathematicians do not yet have a complete description of.

It is this kind of problem that ChatGPT 5.5 Pro makes a substantial mathematical contribution in just one hour.

Mathematical formula sumset

Key Impact: From patchwork to originality

DimensionsChangeWhat it means to usSuggested actions
The ceiling of LLM's mathematical abilityFrom "piecing together known knowledge" to "producing substantial originality"AI will become a research tool rather than a toyPay attention to the improvement of LLM's reasoning ability in Code/Agent scenarios
Scientific research productivityComplete in one hour what human mathematicians do in weeksMulti-agent collaborative research becomes possibleFocus on the potential of Claude Code / ChatGPT in code reasoning
Academic consensusThe voice of skeptics is weakeningThe credibility of AI in professional fields is rapidly increasingIncorporating AI Agents into daily workflows
Content productionAI's enhanced ability to solve complex logical problemsAutomated content production methods are more reliableUse AI Agent to solve more complex content orchestration tasks

Implications for AI Agent workflow

While this breakthrough occurred in the realm of pure mathematics, its direct impact on AI Agents and automated workflows cannot be ignored:

  1. Depth of Reasoning: A model that can solve PhD-level mathematical problems and is more reliable when performing complex multi-step tasks
  2. Reduced error rate: Mathematical proof requires zero errors, which indicates that the accuracy of AI Agent in code generation and data analysis will be further improved.
  3. Long Chain Reasoning: One hour of continuous reasoning means significant progress in context windows and attention mechanisms

For automated workflows, this means that content production pipelines built with OpenAI models and n8n or Claude Code will achieve qualitative improvements in logical consistency, step completeness, and output quality.

Related extended information

Tool entry

The following tools naturally appear in the text, and the platform side will automatically match the maintained tools library to trigger the tool floating card: OpenAI, ChatGPT, Claude Code, n8n

Internal link guidance

View source →

Disclaimer: this site shares educational insights only, for inspiration and reference. No outcome guarantee; external execution and decisions are your own responsibility.