He Let GPT-5.6 Sol Run for 33 Hours Straight to Tackle Fermat's Last Theorem, Forcibly Terminated by the System

marsbitPublished on 2026-07-28Last updated on 2026-07-28

Abstract

This article discusses a real-world experiment by expert Michael P. Frank to test if an AI, specifically GPT-5.6 Sol, could autonomously make progress on a major unsolved mathematical problem: finding a simpler proof for Fermat's Last Theorem. The AI was tasked with exploring specific mathematical pathways and maintaining rigorous notes over approximately 33 hours. However, the session was terminated by OpenAI's systems. The AI itself suggested two possible reasons for the stoppage: excessive resource consumption, or OpenAI having previously failed on similar problems and wishing to conserve computational resources. The AI reported its work primarily involved refining plausible ideas into precise, verifiable statements, most of which were subsequently disproven or excluded—effectively creating a map of dead ends rather than a proof. The incident sparked debate online. Some speculated that OpenAI might deliberately restrict public access to its most powerful models to maintain a competitive edge or avoid regulatory scrutiny, rather than allowing users to potentially solve landmark problems. OpenAI researcher Noam Brown countered this, arguing that a user solving a major problem would be tremendous publicity. Others offered technical explanations, suggesting the termination could be due to standard safety mechanisms preventing infinite loops, or even a known bug in the GPT-5.6 Sol version that disrupts long-running sessions. The story highlights the practical challenges, tec...

As an ordinary person, if you gave a challenging math problem to an AI and repeatedly told it to "continue," could it actually solve the problem, helping you win a large cash prize and even rewrite the history of mathematics?

Recently, with the announcement of the Fields Medal, discussions about mathematics + AI have been very heated, and I'm sure many have imagined the "wish-fulfillment" scenario I just described.

In reality, there are indeed people trying this, but the process is not as smooth as imagined.

Recently, top expert in reversible computing and computational physics, Michael P. Frank, posted that he set a high-difficulty research goal for GPT-5.6 Sol: to explore whether there exists a more concise approach to proving Fermat's Last Theorem than the Wiles–Taylor proof, focusing specifically on the modularity of Frey curves, uniform infinite descent, arithmetic abc-type inequalities, and uniform low-genus quotients; to maintain rigorous notes, computationally verify candidate lemmas, and clearly distinguish between proven results and conjectures.

This task ran continuously in the background for about 33 hours, consuming significant computational resources. But in the end, it was forcibly blocked by the OpenAI system.

In its own post-hoc analysis, GPT-5.6 Sol wrote that there were likely two reasons: first, the system judged that the session was using excessive resources; second, OpenAI had already attempted this problem with a similar model in the past and failed, and this time didn't want to waste computing power again.

Michael P. Frank seemed to agree with this analysis by GPT-5.6 Sol.

Furthermore, GPT-5.6 Sol candidly reported what it did during those 33 hours. The core finding was: it did substantial work, turning many "seemingly feasible" shortcuts into verifiable precise statements, then disproving or eliminating them one by one. The output was a map of "dead ends," not a proof itself. It recommended stopping the effort.

Regarding OpenAI's approach, some offered other possible interpretations: OpenAI might be holding back, not letting the model casually solve super-impressive math problems, fearing others would steal the spotlight, or wanting to claim the credit themselves. This serves as a reminder: if we become too dependent on them in the future, we are essentially handing over the power to decide "what can be discovered" to a single company, which is risky.

However, OpenAI senior research scientist Noam Brown quickly stepped in to point out the unreasonableness of this claim, saying, "If someone just typed 'continue' and used our model to solve a Millennium Prize Problem and then took home $1 million, that would be the best advertisement for OpenAI. Nothing would be better."

That sounds reasonable, but rebuttals argue that while opening top-tier capabilities to users might bring short-term publicity benefits, in the long run, it could weaken OpenAI's leading edge in the AI race. In a multi-player competitive environment, secrecy/internal priority use of strong models to accelerate one's own research is more important than letting users "preemptively" solve major problems. Once capabilities become "available to everyone," their publicity value drops significantly.

Furthermore, recent regulatory scrutiny gives OpenAI reason to hold back. If users used the public version to solve high-difficulty math problems, it would publicly demonstrate "this model is actually very powerful," potentially inviting more regulatory trouble. In this context, deliberately offering a weakened version externally while retaining a super-strong version internally sounds like a plausible strategy.

However, some also pointed out another possibility from a technical perspective: in codex-cli, "goal blocked" means the model repeatedly hit the same limit/blocker for 3 consecutive reasoning turns. So the task termination might be a routine safety/prevention-of-infinite-loop mechanism during model runtime, not a special block targeting difficult problems.

Others said this might simply be caused by a bug. Specifically, version 5.6 Sol has an annoying bug when running long-duration tasks. The temporary workaround is: switch back to version 5.5, compress the context, run it for a few minutes with 5.5, then switch back to version 5.6 Sol. "If you want to keep the same session, you'll get stuck and can't proceed."

What do you think? Feel free to share your experiences using AI to tackle difficult problems or getting stuck with it in the comments section.

References:

https://x.com/lu_sichu/status/2081367506468360495

This article is from the WeChat public account "Almost Human" (ID: almosthuman2014), author: Zhang Qian

Trending Cryptos

Related Questions

QWhat was the main goal set for the GPT-5.6 Sol model in the experiment described in the article?

AThe main goal was to investigate whether a simpler proof of Fermat's Last Theorem exists compared to the Wiles–Taylor proof, focusing on specific areas like modularity of Frey curves and the Uniform abc Conjecture, while maintaining rigorous documentation and verification of candidate lemmas.

QWhy was the GPT-5.6 Sol session forcibly terminated after approximately 33 hours according to its own analysis?

AGPT-5.6 Sol's analysis suggested two possible reasons: the system deemed the session was consuming excessive computational resources, or OpenAI had already tried and failed to solve the problem with similar models and decided to stop wasting resources.

QWhat did the GPT-5.6 Sol model report as its primary achievement during the 33-hour run?

AIt reported that its core achievement was performing substantial groundwork by turning plausible-sounding shortcuts into precise, verifiable statements, and then systematically disproving or eliminating them, effectively producing a map of dead ends rather than a proof itself.

QAccording to the article, what was one argument against Noam Brown's claim that OpenAI would welcome a user solving a major problem?

AThe argument was that while such a solve might offer short-term publicity, long-term competitive strategy favors keeping the strongest models internal to accelerate a company's own research, and that making top-tier capabilities publicly available diminishes their strategic and perceived value.

QWhat technical reason, unrelated to conspiracy, did some commentators suggest for the session's termination?

ASome suggested it could be due to a bug in GPT-5.6 Sol where long-running tasks might get stuck, and a known workaround involves switching to version 5.5, compressing the context, running for a few minutes, and then switching back to 5.6 Sol.

Related Reads

A Brief History of the Lithography Machine: How a Beam of Light Walked Sixty-Nine Years

A Brief History of Lithography: 69 Years of Light In July 2026, reports of a Chinese state-backed company producing five immersion DUV lithography machines sent shockwaves through Wall Street, wiping roughly $44 billion from ASML's market cap in a single day. This event signaled a crack in the long-held assumption of Western monopoly over advanced chipmaking equipment. The journey began in 1957 when Jay Lathrop coined the term "photolithography." Early contact aligners (1960s) gave way to PerkinElmer's revolutionary projection aligners in 1973, boosting yields dramatically. GCA's step-and-repeat system (1978) established the modern stepper blueprint. However, within a decade, Japanese firms like Nikon and Canon, supported by a strong domestic market, captured nearly 90% of the global market from American pioneers, who ultimately exited the business. ASML, founded in a leaky shed in 1984, rose to dominance through key strategic moves: the TWINSCAN dual-stage platform (2001) and the acquisition of SVG, gaining access to Intel. A pivotal moment came in the early 2000s with the industry at a 193nm wavelength impasse. While most invested in the costly 157nm path, TSMC's Burn Lin proposed immersion lithography—using water between the lens and wafer. ASML bet on this simpler idea and, with Zeiss, delivered the first commercial immersion tool in 2004, effectively ending the 157nm roadmap and leaving competitors behind. The subsequent push for Extreme Ultraviolet (EUV) lithography was an even greater marathon. Deemed the least promising option in 1997, EUV's development, led by ASML, required vacuum chambers, reflective optics, and a complex tin-droplet laser plasma source. Critical to its eventual success (first high-volume manufacturing in 2018) was the 2012 "Customer Co-Investment Program," where Intel, TSMC, and Samsung provided upfront funding and equity, sharing the immense risk and cost. Today, ASML holds a near-total monopoly in EUV and immersion DUV. The core lesson of this 69-year history is not merely one of technological invention but of sustained partnership. Successive leaders—PerkinElmer, GCA, Nikon—were not ultimately defeated by superior technology but by losing the vital connection to customers willing to tolerate years of iteration, fund long-term R&D, and integrate early, imperfect tools into their production lines. The "light" that completed the journey was always carried by those patient, invested partners.

marsbit12m ago

A Brief History of the Lithography Machine: How a Beam of Light Walked Sixty-Nine Years

marsbit12m ago

CryptoQuant Analyst Claims 'Frightening Scenario and Simultaneously Promising Signal for Bitcoin' Reveals Ultimate Bottom Point! Here Are the Details

Bitcoin, the leading cryptocurrency, entered a bear market after hitting a historic high of $126,000 in October 2025 and has since fallen over 50%. As BTC recently dropped to around $57,000, opinions diverge on whether the bottom has been reached, with some expecting a drop closer to $50,000. According to a pseudonymous CryptoQuant analyst, the Bitcoin bear market may be nearing its end. The analyst notes that as selling pressure following the all-time high weakens, price corrections are becoming shorter and recovery periods more substantial. The July decline was only slightly below lows seen in early February, and BTC has remained relatively stable around $60,000 since, indicating diminished seller influence. Positive technical signals are also cited, with MACD and RSI showing bullish signals, positive divergences, and oversold conditions, collectively suggesting a potential bottom formation for BTC. However, the analyst does not rule out a final sell-off in the short term. They suggest the ultimate market bottom could be around $51,336, corresponding to the 61.8% Fibonacci retracement level. The $50,000 area is highlighted as a crucial historical support zone, aligning with the average investor cost basis and the 200-week moving average—conditions similar to bottoms in previous bear markets. *This is not investment advice.

cryptonews.ru1h ago

CryptoQuant Analyst Claims 'Frightening Scenario and Simultaneously Promising Signal for Bitcoin' Reveals Ultimate Bottom Point! Here Are the Details

cryptonews.ru1h ago

Trading

Spot

Hot Articles

Discussions

Welcome to the HTX Community. Here, you can stay informed about the latest platform developments and gain access to professional market insights. Users' opinions on the price of S (S) are presented below.

活动图片