He Let GPT-5.6 Sol Run for 33 Hours Straight to Tackle Fermat's Last Theorem, Forcibly Terminated by the System

marsbitPublished on 2026-07-28Last updated on 2026-07-28

Abstract

This article discusses a real-world experiment by expert Michael P. Frank to test if an AI, specifically GPT-5.6 Sol, could autonomously make progress on a major unsolved mathematical problem: finding a simpler proof for Fermat's Last Theorem. The AI was tasked with exploring specific mathematical pathways and maintaining rigorous notes over approximately 33 hours. However, the session was terminated by OpenAI's systems. The AI itself suggested two possible reasons for the stoppage: excessive resource consumption, or OpenAI having previously failed on similar problems and wishing to conserve computational resources. The AI reported its work primarily involved refining plausible ideas into precise, verifiable statements, most of which were subsequently disproven or excluded—effectively creating a map of dead ends rather than a proof. The incident sparked debate online. Some speculated that OpenAI might deliberately restrict public access to its most powerful models to maintain a competitive edge or avoid regulatory scrutiny, rather than allowing users to potentially solve landmark problems. OpenAI researcher Noam Brown countered this, arguing that a user solving a major problem would be tremendous publicity. Others offered technical explanations, suggesting the termination could be due to standard safety mechanisms preventing infinite loops, or even a known bug in the GPT-5.6 Sol version that disrupts long-running sessions. The story highlights the practical challenges, tec...

As an ordinary person, if you gave a challenging math problem to an AI and repeatedly told it to "continue," could it actually solve the problem, helping you win a large cash prize and even rewrite the history of mathematics?

Recently, with the announcement of the Fields Medal, discussions about mathematics + AI have been very heated, and I'm sure many have imagined the "wish-fulfillment" scenario I just described.

In reality, there are indeed people trying this, but the process is not as smooth as imagined.

Recently, top expert in reversible computing and computational physics, Michael P. Frank, posted that he set a high-difficulty research goal for GPT-5.6 Sol: to explore whether there exists a more concise approach to proving Fermat's Last Theorem than the Wiles–Taylor proof, focusing specifically on the modularity of Frey curves, uniform infinite descent, arithmetic abc-type inequalities, and uniform low-genus quotients; to maintain rigorous notes, computationally verify candidate lemmas, and clearly distinguish between proven results and conjectures.

This task ran continuously in the background for about 33 hours, consuming significant computational resources. But in the end, it was forcibly blocked by the OpenAI system.

In its own post-hoc analysis, GPT-5.6 Sol wrote that there were likely two reasons: first, the system judged that the session was using excessive resources; second, OpenAI had already attempted this problem with a similar model in the past and failed, and this time didn't want to waste computing power again.

Michael P. Frank seemed to agree with this analysis by GPT-5.6 Sol.

Furthermore, GPT-5.6 Sol candidly reported what it did during those 33 hours. The core finding was: it did substantial work, turning many "seemingly feasible" shortcuts into verifiable precise statements, then disproving or eliminating them one by one. The output was a map of "dead ends," not a proof itself. It recommended stopping the effort.

Regarding OpenAI's approach, some offered other possible interpretations: OpenAI might be holding back, not letting the model casually solve super-impressive math problems, fearing others would steal the spotlight, or wanting to claim the credit themselves. This serves as a reminder: if we become too dependent on them in the future, we are essentially handing over the power to decide "what can be discovered" to a single company, which is risky.

However, OpenAI senior research scientist Noam Brown quickly stepped in to point out the unreasonableness of this claim, saying, "If someone just typed 'continue' and used our model to solve a Millennium Prize Problem and then took home $1 million, that would be the best advertisement for OpenAI. Nothing would be better."

That sounds reasonable, but rebuttals argue that while opening top-tier capabilities to users might bring short-term publicity benefits, in the long run, it could weaken OpenAI's leading edge in the AI race. In a multi-player competitive environment, secrecy/internal priority use of strong models to accelerate one's own research is more important than letting users "preemptively" solve major problems. Once capabilities become "available to everyone," their publicity value drops significantly.

Furthermore, recent regulatory scrutiny gives OpenAI reason to hold back. If users used the public version to solve high-difficulty math problems, it would publicly demonstrate "this model is actually very powerful," potentially inviting more regulatory trouble. In this context, deliberately offering a weakened version externally while retaining a super-strong version internally sounds like a plausible strategy.

However, some also pointed out another possibility from a technical perspective: in codex-cli, "goal blocked" means the model repeatedly hit the same limit/blocker for 3 consecutive reasoning turns. So the task termination might be a routine safety/prevention-of-infinite-loop mechanism during model runtime, not a special block targeting difficult problems.

Others said this might simply be caused by a bug. Specifically, version 5.6 Sol has an annoying bug when running long-duration tasks. The temporary workaround is: switch back to version 5.5, compress the context, run it for a few minutes with 5.5, then switch back to version 5.6 Sol. "If you want to keep the same session, you'll get stuck and can't proceed."

What do you think? Feel free to share your experiences using AI to tackle difficult problems or getting stuck with it in the comments section.

References:

https://x.com/lu_sichu/status/2081367506468360495

This article is from the WeChat public account "Almost Human" (ID: almosthuman2014), author: Zhang Qian

Trending Cryptos

Related Questions

QWhat was the main goal set for the GPT-5.6 Sol model in the experiment described in the article?

AThe main goal was to investigate whether a simpler proof of Fermat's Last Theorem exists compared to the Wiles–Taylor proof, focusing on specific areas like modularity of Frey curves and the Uniform abc Conjecture, while maintaining rigorous documentation and verification of candidate lemmas.

QWhy was the GPT-5.6 Sol session forcibly terminated after approximately 33 hours according to its own analysis?

AGPT-5.6 Sol's analysis suggested two possible reasons: the system deemed the session was consuming excessive computational resources, or OpenAI had already tried and failed to solve the problem with similar models and decided to stop wasting resources.

QWhat did the GPT-5.6 Sol model report as its primary achievement during the 33-hour run?

AIt reported that its core achievement was performing substantial groundwork by turning plausible-sounding shortcuts into precise, verifiable statements, and then systematically disproving or eliminating them, effectively producing a map of dead ends rather than a proof itself.

QAccording to the article, what was one argument against Noam Brown's claim that OpenAI would welcome a user solving a major problem?

AThe argument was that while such a solve might offer short-term publicity, long-term competitive strategy favors keeping the strongest models internal to accelerate a company's own research, and that making top-tier capabilities publicly available diminishes their strategic and perceived value.

QWhat technical reason, unrelated to conspiracy, did some commentators suggest for the session's termination?

ASome suggested it could be due to a bug in GPT-5.6 Sol where long-running tasks might get stuck, and a known workaround involves switching to version 5.5, compressing the context, running for a few minutes, and then switching back to 5.6 Sol.

Related Reads

3% Cashback, 6% Annual Interest, and a Debit Card: Elon Musk Launches X Money in the US

On July 27, 2026, the social media platform X (formerly Twitter) launched its integrated financial service, X Money, initially available only to US-based Premium and Premium+ subscribers. Founder Elon Musk emphasized the launch's significance for the platform. US residents over 18 with a verified X account can use the service. X Payments LLC is not a bank; user funds are held in FDIC-insured accounts at Cross River Bank, with insurance potentially extended up to $10 million through a cash sweep program across partner banks. The core of the service is the X Card, a virtual or physical metal Visa debit card. Key features include: 3% cashback on most purchases (with some exclusions), free instant peer-to-peer payments, worldwide ATM fee reimbursement, no foreign transaction fees, and support for Apple/Google Pay. The service offers interest on account balances, up to 6.00% APY. Premium+ users get this rate immediately, while Premium users must meet a qualifying direct deposit requirement. Additional features include early direct deposit, wire transfers, bill pay, and paper check ordering. Security features include passkey login, customizable transaction limits, and Visa's Zero Liability Policy for fraud. The public launch followed a limited beta test with select Premium+ users in late June 2026. The launch marks X's continued evolution from a social network into a broader digital platform combining communication with everyday financial operations.

cryptonews.ru47m ago

3% Cashback, 6% Annual Interest, and a Debit Card: Elon Musk Launches X Money in the US

cryptonews.ru47m ago

Trading

Spot

Hot Articles

Discussions

Welcome to the HTX Community. Here, you can stay informed about the latest platform developments and gain access to professional market insights. Users' opinions on the price of S (S) are presented below.

活动图片