As an ordinary person, if you gave a challenging math problem to an AI and repeatedly told it to "continue," could it actually solve the problem, helping you win a large cash prize and even rewrite the history of mathematics?
Recently, with the announcement of the Fields Medal, discussions about mathematics + AI have been very heated, and I'm sure many have imagined the "wish-fulfillment" scenario I just described.
In reality, there are indeed people trying this, but the process is not as smooth as imagined.
Recently, top expert in reversible computing and computational physics, Michael P. Frank, posted that he set a high-difficulty research goal for GPT-5.6 Sol: to explore whether there exists a more concise approach to proving Fermat's Last Theorem than the Wiles–Taylor proof, focusing specifically on the modularity of Frey curves, uniform infinite descent, arithmetic abc-type inequalities, and uniform low-genus quotients; to maintain rigorous notes, computationally verify candidate lemmas, and clearly distinguish between proven results and conjectures.

This task ran continuously in the background for about 33 hours, consuming significant computational resources. But in the end, it was forcibly blocked by the OpenAI system.
In its own post-hoc analysis, GPT-5.6 Sol wrote that there were likely two reasons: first, the system judged that the session was using excessive resources; second, OpenAI had already attempted this problem with a similar model in the past and failed, and this time didn't want to waste computing power again.

Michael P. Frank seemed to agree with this analysis by GPT-5.6 Sol.

Furthermore, GPT-5.6 Sol candidly reported what it did during those 33 hours. The core finding was: it did substantial work, turning many "seemingly feasible" shortcuts into verifiable precise statements, then disproving or eliminating them one by one. The output was a map of "dead ends," not a proof itself. It recommended stopping the effort.
Regarding OpenAI's approach, some offered other possible interpretations: OpenAI might be holding back, not letting the model casually solve super-impressive math problems, fearing others would steal the spotlight, or wanting to claim the credit themselves. This serves as a reminder: if we become too dependent on them in the future, we are essentially handing over the power to decide "what can be discovered" to a single company, which is risky.


However, OpenAI senior research scientist Noam Brown quickly stepped in to point out the unreasonableness of this claim, saying, "If someone just typed 'continue' and used our model to solve a Millennium Prize Problem and then took home $1 million, that would be the best advertisement for OpenAI. Nothing would be better."

That sounds reasonable, but rebuttals argue that while opening top-tier capabilities to users might bring short-term publicity benefits, in the long run, it could weaken OpenAI's leading edge in the AI race. In a multi-player competitive environment, secrecy/internal priority use of strong models to accelerate one's own research is more important than letting users "preemptively" solve major problems. Once capabilities become "available to everyone," their publicity value drops significantly.
Furthermore, recent regulatory scrutiny gives OpenAI reason to hold back. If users used the public version to solve high-difficulty math problems, it would publicly demonstrate "this model is actually very powerful," potentially inviting more regulatory trouble. In this context, deliberately offering a weakened version externally while retaining a super-strong version internally sounds like a plausible strategy.

However, some also pointed out another possibility from a technical perspective: in codex-cli, "goal blocked" means the model repeatedly hit the same limit/blocker for 3 consecutive reasoning turns. So the task termination might be a routine safety/prevention-of-infinite-loop mechanism during model runtime, not a special block targeting difficult problems.

Others said this might simply be caused by a bug. Specifically, version 5.6 Sol has an annoying bug when running long-duration tasks. The temporary workaround is: switch back to version 5.5, compress the context, run it for a few minutes with 5.5, then switch back to version 5.6 Sol. "If you want to keep the same session, you'll get stuck and can't proceed."

What do you think? Feel free to share your experiences using AI to tackle difficult problems or getting stuck with it in the comments section.
References:
https://x.com/lu_sichu/status/2081367506468360495
This article is from the WeChat public account "Almost Human" (ID: almosthuman2014), author: Zhang Qian






