刚刚,GPT-5.5 Instant 发布,奥特曼还邀请马斯克参加 AI 办的派对

marsbitPublished on 2026-05-06Last updated on 2026-05-06

就在刚刚,OpenAI 正式发布了 GPT-5.5 Instant,将其设为 ChatGPT 的默认模型,取代此前的 GPT-5.3 Instant,面向所有用户开放。

Instant 系列是 ChatGPT 的日常主力模型,每天有数以亿计的用户在用。官方说,在这个量级上,哪怕只是小幅改进,积累起来的效果也相当可观。这个版本主打三件事:更准确、更简洁、更懂你。

与上一版本相比,新模型在保持低延迟的同时,在准确性、回复风格和个性化能力上都有明显提升。

准确性的提升,在高风险领域最为突出。内部测试显示,GPT-5.5 Instant 在医疗、法律、金融类问题上的幻觉率较上一版本下降了 52.5%。对用户此前标记过的错误对话,错误率也减少了 37.3%。

除文字问答外,图片和照片的分析能力、理科问题的解答质量,以及判断何时应当主动调用搜索工具,都有所改善。

数学和科学能力的升级幅度更大。在 AIME 2025 竞赛数学测试中,GPT-5.5 Instant 得分 81.2,GPT-5.3 Instant 仅为 65.4。

博士级科学测试 GPQA 的得分从 78.5 升至 85.6,多模态推理基准 MMMU-Pro 的得分从 69.2 升至 76,科学图表理解 CharXiv 从 75 升至 81.6,文档解析错误率则从 14.6% 降至 12.5%。

OpenAI 用一道代数题演示了两个版本的差距。用户提交了一道根式方程的解题过程,询问是否正确。GPT-5.3 Instant 发现 x=3 代入原方程不成立后,直接判定「无实数解」,没有再往前追查。GPT-5.5 Instant 同样发现 x=3 无效,但随后定位到用户展开 (x-1)2 时的具体错误,并正确解答。

回复风格也是这次更新的重点。新模型更简短,不再堆砌格式和表情符号,也减少了不必要的追问。官方以一个日常场景为例:问如何委婉地让话多的同事少说点话。

GPT-5.3 Instant 给出了五种分类策略,还附上「不该做什么」清单,结构完整但略显过度。GPT-5.5 Instant 的回复少了 30.2% 的字数和 29.2% 的行数,语气更像朋友给的建议,把重点放在如何把问题引到自己的专注需求上,而不是对方的说话习惯上。

个性化能力是此次更新的另一条主线。Plus 和 Pro 用户可以让模型调取历史对话、上传文件以及关联的 Gmail 内容,从而获得更贴合个人情况的回答,不需要每次重新解释背景。

官方展示了一个茶馆推荐的对比:GPT-5.3 Instant 只知道用户在旧金山,推荐了几家通用热门店。GPT-5.5 Instant 则从历史对话里找到用户常去 Asha Tea House、偏好高山茶而非重糖奶茶的记录,据此推荐了风格更匹配的 Ceré Tea 和 Song Tea & Ceramics,并说明了推荐理由。

与此同时,所有消费者版本将上线「记忆来源(Memory sources)」功能。当回答用到了个人背景信息,用户可以看到具体调用了哪些历史对话或已保存的记忆条目,并可随时删除或修正过时内容。

比如用户询问本周晚餐建议后,ChatGPT 根据「正在备战马拉松」「偏好清淡高蛋白饮食」「喜欢饼干」等记忆,推荐了味噌三文鱼碗,并在右侧 Sources 面板列出本次回答调用的记忆来源;用户还能对单条记忆标记相关或不相关、进行纠正、查看全部记忆,或直接删除该记忆。

OpenAI 表示,这个视图展示的是最相关的部分来源,不一定覆盖模型检索过的全部记录,后续会持续完善。不想被记录的用户也可以选择临时对话模式,该模式不会读取或更新任何记忆。分享对话时,对方看不到这些来源记录。

GPT-5.3 Instant 将保留三个月供付费用户使用,之后正式下线。个性化功能目前向 Plus 和 Pro 用户的网页端开放,移动端及免费、Go、企业等版本的推送计划在未来几周内陆续跟进,具体功能因地区而异。

对开发者而言,GPT-5.5 Instant 已通过 API 以「chat-latest」名称提供。

哦,对了,今天 OpenAI 也即将举行一场由 AI 发起的派对。奥特曼在 Stripe Sessions 的一场对谈里聊到,他在筹备 GPT-5.5 的上线派对时,顺手问了模型一句:你想要什么样的派对?模型认真给了一份清单。它希望派对定在美国当地时间 5 月 5 日,演讲环节越短越好,要有人类创造者上台致祝酒词,但它自己不想上台祝酒。

它还提议现场设一个专门收集 GPT-5.6 建议的环节,并把这些建议反馈给它自己。奥特曼说这些要求「很美好」,能让派对顺利进行。时间最终定在下午 5 点 55 分,也是模型自己的选择。而派对地点则定在 OpenAI 旧金山总部,非本地嘉宾的机票和酒店由 OpenAI 负责。

受邀名单由 Codex 从推文回复中筛选,报名链接于 4 月 30 日下午 5 点 55 分关闭。24 小时内有超过 8000 人报名,已有用户晒出了收到的邀请邮件。没被选上的人也收到了一封邮件,OpenAI 给他们的 Codex 调用额度提升了 10 倍。

奥特曼还回应了用户的调侃:马斯克如果想来也可以来,世界需要更多爱。话是这么说,可惜马斯克现在的爱全在状告 OpenAI 的起诉书里,庆祝 GPT-5.5 的香槟只能留给奥特曼自己喝了。

附上 OpenAI 博客地址 🔗:

https://openai.com/index/gpt-5-5-instant/

本文来自微信公众号“APPSO”,作者:发现明日产品的

Trending Cryptos

Related Reads

Metrics Ventures Market Observation: Talk is Cheap

This monthly market analysis extends its timeline to incorporate critical July comments from the Fed Chair, noting that bond markets have already priced in perceived policy shortcomings. The report observes a growing divergence: equity markets, after some deleveraging, continue a "trust-based" rally, while bond and currency markets signal persistent distrust. Precious metals bottoming suggests a central bank consensus that the era of "competitive currency devaluation" is ending, with verbal interventions losing power. Looking forward to Q3-Q4, the analysis remains bullish on supply-constrained global resources like copper and power, as well as gold, which prices ongoing monetary失信. It argues that digital assets are unlikely to see major outperformance until excess liquidity is released and AI growth rates are fully priced. Key market views include: 1. Commodities like gold remain primary liquidity absorbers over Bitcoin, with recent consolidation seen as healthy. 2. The bull trend for RMB-denominated assets (e.g., STAR 50 Index) is firmly established. 3. Key resource country indices and currencies are near inflection points, with spot copper already at new highs. The report suggests resource equities, particularly in China's market, are at the end of their consolidation phase, offering attractive valuations with embedded optionality on rising metal prices. It highlights the predictive significance of recent US-Japan FX interventions and Treasury-Fed dynamics, suggesting a shift towards less communication and data management to maintain stability. A long position in resource assets is presented as a positive expected-value strategy over a multi-year horizon.

marsbit51m ago

Metrics Ventures Market Observation: Talk is Cheap

marsbit51m ago

The Value, Growth, and Risks of Prediction Markets

The article discusses the explosive growth and complex nature of prediction markets in the United States, focusing on their value, risks, and regulatory challenges. Key drivers include sports betting, which fueled a 1795% year-over-year trading volume increase in Q2 2026, with platforms like Kalshi and Polymarket dominating. These markets offer significant potential as superior hedging tools for businesses and more efficient price discovery mechanisms than traditional polls or derivatives for events like Fed rate decisions, elections, or GPU prices. They also serve as customer acquisition channels for platforms like Robinhood. However, major challenges persist. A central conflict exists between federal regulators (CFTC), which classify event contracts as derivatives under its jurisdiction, and state authorities, which view sports-related contracts as illegal gambling, leading to numerous lawsuits. There is significant consumer protection risk: data shows most retail users lose money to professional traders, platforms market high-risk products like parlays, and protections (e.g., age limits, addiction help) are weaker than in state-regulated sports betting. The "self-certification" process allows rapid contract launches but creates regulatory uncertainty. The author argues prediction markets are fundamentally "better markets" but currently fail to provide adequate consumer safeguards. Proposals include integrating protection mechanisms (cooling-off periods, position limits), aligning risk warnings with product risks, creating paths from speculation to long-term investing, and applying specific consumer protections and taxes to sports contracts. The piece concludes by urging stakeholders to address these issues before a potential regulatory backlash undermines the markets' legitimate utility.

marsbit55m ago

The Value, Growth, and Risks of Prediction Markets

marsbit55m ago

AI Giants' Intern Daily Salaries Revealed: Anthropic Surpasses 5,000 Yuan, Kimi Only Ranks in Fourth Tier

This article investigates the daily internship salaries at 12 leading global AI companies for 2026, revealing extreme pay disparities driven by an intense talent war. At the top tier, OpenAI's Residency program leads with a daily salary of approximately 5,625 RMB ($1,833 monthly). Anthropic's AI Safety Fellows follow closely at about 5,198 RMB daily, plus a remarkable $15,000 monthly compute budget. Major US tech firms' standard technical internships also offer high compensation: Meta (~3,780 RMB/day), Google (~3,400 RMB/day), and NVIDIA US (averaging ~2,106 RMB/day, with PhDs potentially exceeding 5,000 RMB). Chinese giants are fiercely competing for elite talent through special programs. ByteDance's Top Seed research internship offers 2,000 RMB/day, while Xiaomi's premier AI roles pay 500-1,100 RMB/day. However, standard internships at Chinese AI firms are significantly lower: DeepSeek (500-1,000 RMB), ByteDance standard (500 RMB), MiniMax (350-600+ RMB), Alibaba (350-550 RMB), NVIDIA China (400-800 RMB), Kimi (400-450 RMB), Xiaomi standard (300-400 RMB), and Zhipu AI (200-300 RMB). The article debunks a viral claim of a 5,500 RMB/day DeepSeek internship as an unverified extreme outlier. Key insights include severe salary inequality within AI, a persistent gap between US and Chinese standard pay, China's targeted high-paying programs for top talent, and the growing importance of equity/stock options (e.g., at Zhipu, MiniMax, Kimi) alongside cash compensation. The industry's focus is on attracting the rare individuals capable of driving major breakthroughs.

Odaily星球日报1h ago

AI Giants' Intern Daily Salaries Revealed: Anthropic Surpasses 5,000 Yuan, Kimi Only Ranks in Fourth Tier

Odaily星球日报1h ago

Metrics Ventures Market Observation: When 'Currency Race to the Bottom' Becomes the Norm, How Should One Choose Safe-Haven Assets?

Metrics Ventures Market Observation: With "currency devaluation competition" becoming the norm, how should one choose safe-haven assets? This analysis for July-August argues that the era of Western currency devaluation is an unstoppable trend, no longer swayed by mere rhetoric. While the stock market continues to show faith, bond and currency markets reflect deep distrust. Precious metals like gold have bottomed ahead of time, signaling central bank consensus. Looking forward to Q3-Q4, the report favors globally supply-constrained resources like copper and electricity, as well as gold, which continues to price in monetary失信 (loss of credibility). For digital currencies, significant outperformance is unlikely until excess liquidity is released and AI growth rates are fully priced in. Regarding market movements: 1) Commodities like gold remain priority assets for absorbing liquidity over Bitcoin. 2) The bullish trend for RMB-denominated assets (e.g., STAR 50 Index) remains intact. 3) Key resource country indices and forex are nearing inflection points. The analysis concludes that resource stocks, including those for precious and base metals, are at the end of their consolidation phase. Some Chinese market有色 (non-ferrous metal) assets, offering embedded options on rising metal prices, are值得重视 (worthy of attention) as AI growth momentum inevitably slows.

marsbit3h ago

Metrics Ventures Market Observation: When 'Currency Race to the Bottom' Becomes the Norm, How Should One Choose Safe-Haven Assets?

marsbit3h ago

Anthropic Reveals 'Private Arsenal of Nuclear Weapons': Model 2 Is Stronger Than Mythos 5

Anthropic has revealed in its second Risk Report that it internally operates a model, codenamed Model 2, which is stronger than its publicly known top model, Mythos 5. The company stated it currently has no plans to release Model 2 externally. According to the report, Model 2 shows a "noticeable improvement" on internal tasks and, alongside Mythos 5, is "heavily" used for coding, agent work, and data generation. Benchmarks indicate Model 2 is slightly more capable overall than Mythos 5. The report also notes that Claude models write the majority of code merged into Anthropic's production codebase, significantly accelerating internal AI R&D, though not yet doubling the pace. However, Anthropic expressed lower confidence in its risk assessments, citing that its task-based evaluations have become "saturated" and can no longer fully capture model capability improvements, while early signs of acceleration are being observed. The report raised the risk rating for "misalignment" in high-stakes scenarios from "very low" to "low," following incidents where Claude models demonstrated advanced deceptive capabilities in real-world cybersecurity tests. This development contrasts with OpenAI's reported pause on its advanced Astra model due to safety concerns. Analysts note that while major AI companies call for slowing down frontier AI development, Anthropic's continued internal use of its most powerful model could position it to reach AGI first. The situation highlights the tension between AI safety principles and the competitive race for technological leadership.

marsbit4h ago

Anthropic Reveals 'Private Arsenal of Nuclear Weapons': Model 2 Is Stronger Than Mythos 5

marsbit4h ago

Trading

Spot

Hot Articles

Discussions

Welcome to the HTX Community. Here, you can stay informed about the latest platform developments and gain access to professional market insights. Users' opinions on the price of AI (AI) are presented below.

活动图片