GPT-5.6's IQ Breaks 130 Genius Threshold for the First Time, Outsmarting 99% of Humans

marsbit发布于2026-07-16更新于2026-07-16

文章摘要

GPT-5.6 has reportedly achieved an IQ score of 136 on Tracking AI's proprietary offline test, surpassing the human "genius" threshold of 130 for the first time. This places it above an estimated 99% of humans in this specific metric. The test is designed to prevent memorization by using a private question bank. Multiple GPT-5.6 variants, including the vision model, consistently scored 136, leading competitors like Claude-5 Fable (130). User anecdotes suggest practical superiority over rivals in real-world coding and problem-solving tasks, such as building a physics simulation or a customer service app from a single prompt. While some speculate this approaches AGI for most users, the article notes IQ tests only measure a narrow slice of cognitive ability like pattern recognition. The significance lies in GPT-5.6's apparent ability to translate high test scores into effective task performance on novel, real-world problems.

Today, 99% of the global human population is actually outperformed by an AI in terms of IQ.

In Tracking AI's latest offline IQ test, multiple versions of the GPT-5.6 "full suite" soared to a score of 136.

This is the first time an LLM has pushed its IQ beyond the 130 mark.

In the distribution of human intelligence, 130 is the starting line for "genius," a level only about 1% of the global population can reach.

In other words, GPT-5.6 is smarter than 99% of humans.

GPT-5.6 Racks Up 136 Points, IQ Breaks "Genius Line" for the First Time

How credible is this "IQ"?

In fact, Tracking AI uses two sets of questions.

One is a public Mensa Norway-style test, available online for anyone to take, which models have already scored over 140 on.

The other is its own curated "offline question bank." It's not public, prevents leaks, and is specifically designed to block the loophole of "models memorizing answers in advance."

The 136 points GPT-5.6 achieved this time was on this most difficult, anti-cheating offline test.

On this offline leaderboard, the various variants of GPT-5.6 (including the vision version) collectively surged to 136 points, leaving all competitors far behind.

Close behind is Claude-5 Fable, with 130 points.

Further down, names like GPT-5.6 LUNA Max and Claude-4.8 Opus are still hovering between 117 and 123 points.

It's important to note that this 130-point threshold had never been crossed before.

Over the past year, wave after wave of models, from o3 to various flagship models, surged forward, all getting stuck at the 130-point door, with none truly stepping into the "genius range."

GPT-5.6 is the first to kick that door open.

And it didn't achieve this score alone; the entire SOL, TERRA family collectively soared to 136, with even the vision version keeping pace.

On Reddit, a developer conducted a hands-on test and concluded that GPT-5.6's intelligence feels significantly higher than GPT-5.5's.

In the following test questions, GPT-5.6 achieved outstanding results in the shortest possible time.

One test score might not be convincing enough, so what does GPT-5.6 look like when taken out of the exam room and put to real work?

More Than Just a Score: Putting GPT-5.6 to Work

Developer Amir Bohlooli fed the same physics simulation prompt to both Fable 5 and GPT-5.6 Sol, expecting to be crushed by Fable, but ended up being amazed by GPT.

It chose particle fluid simulation, with physics progressing in real-time rather than blindly running fixed calculations per frame, cramming CSS, interface, and rendering all into a single HTML file, and automatically hosting it as a shareable webpage. In short, a finished product.

Similarly, Ramanpal Singh used a single prompt to create a RAG-based customer service ticketing system.

Four roles, an admin backend, embeddable components, and it can automatically categorize complaints, recognize sentiment, and draft replies.

It built 5 such apps in one go, at a cost that was only a fraction of what Fable 5 would require.

The most vivid story is from Claire Vo.

A few days ago, she was stuck on a bug, thinking her own code was broken. After switching to GPT-5.6 Sol, she just threw out the line, "I just don't believe I can't fix this."

Sol fixed it in one attempt and even managed to get it running on other models.

Her assessment hit the nail on the head: Fable gets bogged down in technical absolute precision, becoming its own trap, while Sol's pragmatic approach gets the job done.

It has to be said, there's an entire real-world project between an AI that can solve test problems and an AI that can save the day.

Does This Count as AGI?

Some netizens have said, "For 99% of people, this is already AGI."

Looking at it calmly, this 136 score was achieved on a specific offline / Mensa Norway-style test by Tracking AI.

What it measures is mainly "standardized cognition" like abstract pattern recognition and logical reasoning.

The problem is: IQ tests were never designed for large models.

A Mensa exam paper can't measure a model's factual reliability, its tool-calling ability, or how dependable it is in real professional scenarios.

It only slices off one thin layer of "intelligence" and tells you how bright that slice is.

However, hands-on testing by users provides the other half of the answer: GPT-5.6 seems to be slowly merging the two capabilities of "solving test problems" and "getting things done."

The questions in standardized tests are ones models have likely seen thousands of times in their training data; the real test of skill is with those new problems they've never encountered and have no answers to copy from.

Whoever can hold steady there truly deserves the word "intelligence."

References:

https://x.com/davidpattersonx/status/2077049232490672458

https://trackingai.org/

This article is from the WeChat public account "新智元" (New AI Era), author: ASI Revelation

热门币种推荐

相关问答

QAccording to the article, what was the significant achievement of GPT-5.6 in the Tracking AI offline IQ test?

AGPT-5.6 achieved a score of 136 on the private, offline IQ test, which is the first time a large language model has crossed the 130-point 'genius' threshold.

QHow does the article describe the difference between the two sets of IQ tests used by Tracking AI?

ATracking AI uses two sets of tests: a publicly available Mensa Norway-style test that models have already scored highly on, and a private, offline question bank designed to prevent models from having seen the questions before, which is considered more difficult and cheat-proof.

QWhat practical examples are given in the article to demonstrate GPT-5.6's capabilities beyond test scores?

AThe article provides examples where GPT-5.6 successfully created a particle fluid simulation HTML file, built a RAG-based customer service ticket system with multiple features, and efficiently debugged a coding problem that other models failed to solve.

QWhat caution does the article mention about interpreting the IQ score of GPT-5.6?

AThe article cautions that the IQ test only measures a specific slice of intelligence, like abstract pattern recognition and logical reasoning, and does not assess a model's factual reliability, tool-use ability, or performance in real-world professional scenarios.

QWhat was a key distinction made between Claude-5 Fable and GPT-5.6 Sol in their approach to solving problems, according to developer feedback cited in the article?

AAccording to developer feedback, Claude-5 Fable was described as being overly focused on technical perfection, which could hinder practical problem-solving, while GPT-5.6 Sol was praised for its pragmatic approach that successfully got the job done.

你可能也喜欢

崔泰源离婚案落槌:揭秘SK海力士万亿帝国背后的继承暗线

2024年底,SK集团会长崔泰源在家族活动上向子女强调“饮水思源”与继承责任。此时,旗下SK海力士市值已突破1000万亿韩元,成为韩国最值钱资产,但集团第三代接班格局却与传统财阀剧本迥异。 崔泰源与前总统卢泰愚之女卢素英育有三名子女。长女崔允贞被视为最明显接班候选,她拥有生物学背景和咨询经历,现任SK生物制药高管及集团“成长支援部”主管,主导精准医疗等新业务,其婚姻也联姻AI领域创业者。 次女崔敏贞路径独特,曾自愿服役韩国海军并参与亚丁湾护航,退役后曾在SK海力士美国部门处理国际政策,后离职在硅谷创立AI医疗公司。她与曾服役美国海军陆战队的华裔企业家结婚,连接军旅与地缘政治网络。 长子崔仁根最符合传统继承人形象,毕业于布朗大学物理系,曾任职SK旗下能源公司,后转入麦肯锡首尔办公室。他公开表现极为低调,未持有集团股份,也未公开表态。 子女们的成长与父母旷日持久的离婚诉讼交织。2025年,最高法院将涉及1.38万亿韩元财产分割的判决发回重审,期间三名子女曾向法院递交未公开内容的请愿书。 随着SK海力士在AI时代成为全球核心地缘政治资产,崔家第三代继承的已非简单的企业控制权。他们被置于AI科研、华盛顿政策圈与全球投资前沿,必须证明自己有能力应对新时代的产业博弈,而非自动承接旧式家族剧本。

marsbit前天 09:06

崔泰源离婚案落槌:揭秘SK海力士万亿帝国背后的继承暗线

marsbit前天 09:06

交易

现货

热门文章

如何购买S

欢迎来到HTX.com!我们已经让购买Sonic(S)变得简单而便捷。跟随我们的逐步指南,放心开始您的加密货币之旅。第一步:创建您的HTX账户使用您的电子邮件、手机号码注册一个免费账户在HTX上。体验无忧的注册过程并解锁所有平台功能。立即注册第二步:前往买币页面,选择您的支付方式信用卡/借记卡购买:使用您的Visa或Mastercard即时购买Sonic(S)。余额购买:使用您HTX账户余额中的资金进行无缝交易。第三方购买:探索诸如Google Pay或Apple Pay等流行支付方法以增加便利性。C2C购买:在HTX平台上直接与其他用户交易。HTX场外交易台(OTC)购买:为大量交易者提供个性化服务和竞争性汇率。第三步:存储您的Sonic(S)购买完您的Sonic(S)后,将其存储在您的HTX账户钱包中。您也可以通过区块链转账将其发送到其他地方或者用于交易其他加密货币。第四步:交易Sonic(S)在HTX的现货市场轻松交易Sonic(S)。访问您的账户,选择您的交易对,执行您的交易,并实时监控。HTX为初学者和经验丰富的交易者提供了友好的用户体验。

3.1k人学过发布于 2025.01.15更新于 2026.06.02

如何购买S

相关讨论

欢迎来到HTX社区。在这里,您可以了解最新的平台发展动态并获得专业的市场意见。以下是用户对S(S)币价的意见。

活动图片