黑鲸出水,DeepSeek下半场开始

marsbitPublished on 2026-08-13Last updated on 2026-08-13

Abstract

8月13日,DeepSeek发布了两项重要更新:一是DeepSeek V4 Pro正式版发布,二是其自研的AI智能体执行系统“DeepSeek Harness”开发者预览版以MIT协议开源。 文章指出,在AI Agent时代,决定智能体实际效能和成本的不只是模型本身,外部的“执行系统”同样关键。测试表明,同一DeepSeek V4-Flash模型接入不同的执行系统,完成任务的数量和成本可能相差数倍。Harness正是为了最大化释放模型潜力、控制任务执行成本而打造。 DeepSeek Harness的核心设计哲学是“一切皆插件”,其基于Cordis插件系统构建,允许开发者通过插件灵活替换或扩展模型、工具、UI等几乎所有组件,而非固定形态的AI助手。它强调可追溯性,采用只追加的会话日志,确保执行过程可审查、可恢复。其团队负责人具有量化交易系统背景,注重系统在复杂任务下的稳定执行与风险控制。 Harness的发布标志着DeepSeek商业模式的潜在转变:从按消耗的Token收费,转向为任务完成的结果负责。这使其与Claude Code等产品形成差异化竞争。尽管Harness目前仅为v0.1预览版,面临生态建设、市场接受度等挑战,但它代表了DeepSeek将前沿能力低成本化、并接入可落地系统的一贯战略方向。其未来很大程度上取决于开源社区的参与和共建。

DeepSeek最具潜力项目登场,一切皆可插件。Harness不是要模仿谁,而是要把定义的权限交给开源社区。

8月13日傍晚, DeepSeek 官方终于发声,两个重要信息发布。其一,是 DeepSeek V4 Pro正式版发布,并同步在 APP、网页端和 API 更新上线。用户可以通过 APP 或网页端选择“专家模式”使用全新的 V4 Pro 正式版模型,API 模型名不变。

其二,有着一只黑色鲸鱼头像的“ DeepSeek Harness团队”迎来首次发文,宣布Harness开发者预览版正式上线,并以MIT协议开放源代码。

如此前预告,V4 Pro正式版与官方自研Harness同步亮相,由于我们此前已经着重介绍过V4 Pro,本篇文章将重点为大家解析, DeepSeek Harness为什么有看点。

以及经过今天, DeepSeek 会不会从一家模型公司,变成一家更难定义的公司。

01

同一个模型,换套外壳像换了个人

要理解Harness为什么重要,得先理解一个反直觉的事实:在Agent时代,决定一个AI能不能干活的,不只是模型。

8月6日和11日,智能体工具公司Composio通过两次测试实验,试图证明Harness的价值。研究者把同一个 DeepSeek V4-Flash接入八种不同的Harness,也就是套在模型外面的执行系统,让它们分别完成三十项多步骤任务。这些任务不是问答,而是要求Agent进入Gmail、Google Calendar、GitHub、Slack等真实应用,调用工具并改变应用状态,每项最长运行十五分钟,全部检查通过才算成功。

结果差距不小。完成最多的Pi Agent通过了二十项,最少的OpenCode只通过十四项;八种Harness总共运行二百四十次,仅一百二十九次成功;三十项任务中,只有六项被所有Harness完成。在成本方面,同样完成十六项任务的Claude Code、Codex和DeepAgents,每完成一项成功任务的估算成本分别约为0.195美元、0.081美元和0.045美元。同一个模型,差了四倍多。

这意味着,模型划定能力的上限,Harness决定这份上限最终能兑现多少、要花多少钱兑现。一项长任务跑下来,执行系统要不停地做判断,上下文里留什么、丢什么,什么时候调哪个工具,工具报错了是重试还是换路,模型说做完了到底信不信、要不要再验一遍。任何一步处理不好,模型就算思路对,也可能在真正改文件、提交代码时功亏一篑。

值得一提的是,这是 DeepSeek 走红以来,极少数提前开放内测的项目。

多个开发者在发布前就拿到了Harness内测资格,有人内测发现,让同一个V4-Flash分别在 DeepSeek Harness、Reasonix和Codex中完成同一款游戏,会跑出截然不同的结果。这说明,当模型完全相同,执行系统提供的工具、提示词、上下文组织和执行策略,足以显著改变最终产物。

这恰恰是 DeepSeek 重视自研Harness的原因。 DeepSeek 内部曾认为,Agent要解决持续学习问题,最终让AI加快AI研发。顺着这条线看,Harness就不是一个外挂的编程工具,而是模型进入真实研发任务的工作台。

除此之外,一家靠低价和模型能力立足的公司,如果任务交付、失败反馈和开发者入口长期挂在别人的系统里,即便其在价格上再有优势,却决定不了一个任务究竟烧掉多少Token、要重试几次、什么时候才算真正完成。

02

一切皆插件,DeepSeek不做谁的替代品

Harness到底是什么?市面上最省事的说法是“ DeepSeek 版Claude Code”。但用完内测版的人普遍认为,这个定位不准确。

据 DeepSeek 官方文档,Harness的核心设计哲学是“Everything is a plugin”——一切皆插件。它基于Cordis插件系统构建,模型、工具、技能、会话、沙箱、存储、循环、调度、UI,所有Agent能力均由插件组合而成。开发者不需要改动Harness源码,就能在配置中选择、替换或扩展任意一项能力。

这和通常理解的插件不是一个量级。VS Code的插件是给编辑器加功能,Codex的插件是给Agent加工具;而Harness的插件可以换掉Agent的大脑、工具箱、规则乃至整张脸。

据多位参与内测的人士透露,内测期间开发者在做插件这件事上玩疯了,“大伙都不好好内测了,跑去写插件了”,在短短几天内就出现了几百个插件——有人直接改了整个工作界面,有人用纯插件实现了“跨会话长期记忆加后台自我进化”,让模型定期回头审视自己的工作记录,把临时经验压缩成长久知识。

还有人评价称,开源社区治理可能也是 DeepSeek 接下来要面临的难题。

另一处关键设计是可追溯性。Harness采用只追加(append-only)的会话日志,模型看到的一切——系统提示词、思维链、工具调用与结果、子Agent调度、每一次上下文注入都会被完整记录。上下文压缩不会删除原始历史,只是用替换事件改变模型此后看到的表象。开发者可以在Trajectory视图中按来源追溯,支持恢复、分叉、检索和回放。官方文档把这一原则概括为“模型可见,即已记录”。

这背后是一种典型的系统工程思维。据报道, DeepSeek Harness团队负责人崔添翼本科毕业于浙江大学计算机系,曾在量化交易机构Jane Street任职九年,2026年3月加入 DeepSeek ,5月Harness内部立项。高频量化交易系统的核心壁垒从来不是策略多聪明,而是极端复杂环境下的稳定执行、异常兜底、全程日志追溯和风险可控——这恰好是AI Agent从演示走向生产最缺的那块拼图。有内测开发者用“稳如老狗”形容它跑长任务时的表现:任务断了自动存档,下次接着跑,不用从头再来。

值得注意的是,Harness没有把自己关在 DeepSeek 模型的围墙里。它默认支持接入 Kimi 、OpenAI、Anthropic、Google等近四十家大模型。可能别家倾向于做一个现成的Agent,但 DeepSeek Harness更像一套组件,开发者可以自己决定Agent应该是什么样子,怎么工作。

当然,官方保持了清醒。公告表示,作为早期预览版本,“当前仍有许多细节有待改进和打磨,核心插件与基础接口也将在后续快速迭代”。事实上,这次内测的媒体之一爱范儿就感叹,很难想象,以快著称的 DeepSeek 有一天运行一项编程任务的时间也会来到半个多小时。v0.1的版本号也说明,这条黑鲸刚出水,远没到畅游的时候。

03

从卖Token到交结果,黑鲸的野心

Harness的浮出水面,暗合着整个AI行业的一次转向。

过去数年,大模型公司的竞争集中在参数规模和基准跑分上。但进入2026年,头部模型的基础能力快速收敛,单次问答的差距不断缩小,价格战却愈演愈烈。单纯卖Token的商业模式,天花板已经肉眼可见。行业逐渐意识到,AI产业的核心价值不在模型输出,而在场景落地——同样的Token消耗,闲聊问答价值微薄,而修复一个Bug、完成一次功能开发,可以创造数倍的商业价值。

DeepSeek 的新发布,意味着其从卖算力走向交结果。在旧模式下,客户按Token调用量付费,不管模型有没有真正解决问题;在新模式下,付费的锚点从消耗了多少算力转向完成了什么任务。

这次 DeepSeek Harness发布,广受开发者好评还在于另一重原因。 DeepSeek 不定义Agent本身,而是利用开源社区的扩散能力,去推崇一种Harness方案。这是一条更宏大,也更难的路径。

并且,Harness面对的是Claude Code和Codex已经验证过的成熟市场,后两者背靠Anthropic和OpenAI的模型迭代与商业资源,先发优势明显。开源之后,插件生态能否自发繁荣、开发者是否愿意把生产环境托付给一个0.1版本、国内市场的付费习惯能否支撑“按结果付费”的新逻辑,都是悬而未决的问题。 DeepSeek 自己也表示,核心插件和API仍将快速演进,这意味着早期接入者要承担接口变动的成本。

8月13日这一天,太平洋两岸各有一场发布。马斯克为Grok 4.6站台,称其“智能、快速且性价比极高” ; DeepSeek 什么都没说,只是一味丢更新。

喧嚣与沉默之间,是两种商业哲学的分野。

我们想要称其为 DeepSeek “史上最具潜力项目”,不是因为它今天已经最强。V4 Pro与sota模型仍有差距,Harness还只是v0.1的开发者预览版,涨价后的市场反应有待观察,生态建设更是长路漫漫。潜力之所以是潜力,恰恰因为它尚未兑现。

但如果把时间线拉长,从R1到V4,从模型到Harness, DeepSeek 走的每一步都踩在同一个方向上:把前沿能力压到可负担的成本,再把可负担的能力接入可落地的系统。

黑鲸已经出水。至于它能游多远,在属于开源社区的未来里。

本文来自微信公众号“凤凰网科技”,作者:Dale,编辑:董雨晴

Trending Cryptos

Related Questions

QDeepSeek Harness 是什么?它的核心设计理念是什么?

ADeepSeek Harness 是 DeepSeek 官方发布的智能体执行系统,是一个基于 Cordis 插件系统构建的工作台,可以让模型在真实研发任务中更好地工作。其核心设计理念是“一切皆插件”,即模型、工具、技能、会话、UI等所有Agent能力均由插件组合而成,开发者无需修改源码即可灵活配置和扩展。

Q根据文章中的实验,Harness 为什么重要?

A实验表明,同一个 DeepSeek V4-Flash 模型接入不同的 Harness(执行系统)后,完成复杂任务的成功率、效率和成本差异巨大。这说明,Harness 虽然不决定模型能力的上限,但决定了模型能力最终能在多大程度上、以多少成本被有效兑现。一个好的 Harness 能够通过优化提示词、上下文管理、工具调用和错误处理等执行策略,显著提升AI的实际工作效果。这表明了执行系统在Agent时代的重要性。

QDeepSeek Harness 相比于 Claude Code 或 Codex 有什么不同?

ADeepSeek Harness 并非仅仅是“DeepSeek版Claude Code”。其主要不同点在于:1. 更加彻底的插件化架构:不只是工具,而是包括模型、UI在内的所有组件都可替换为插件。2. 开放的模型支持:默认支持接入 Kimi、OpenAI、Anthropic、Google 等近四十家大模型,不局限于 DeepSeek 自身的模型。3. 强大的可追溯性:采用只追加的会话日志,确保执行过程的每一步都可追溯、恢复和回放。总的来说,它更像是一套供开发者自定义Agent的基础组件,而非一个封闭、固定的成品Agent。

Q文章提到DeepSeek Harness的负责人崔添翼的背景有何特点?这对Harness的设计有何影响?

AHarness 团队负责人崔添翼本科毕业于浙江大学计算机系,并曾在量化交易机构 Jane Street 任职九年。高频量化交易系统强调极端复杂环境下的稳定执行、异常兜底、全程日志追溯和风险可控。这种背景使 Harness 的设计带有浓厚的系统工程思维,追求长任务执行的稳定性(如“稳如老狗”、任务中断可续跑)和全流程的可追溯性,而这正是当前AI Agent从演示走向实际生产环境所急需的关键能力。

QDeepSeek 发布 Harness 意味着公司怎样的战略转变或野心?

A发布 Harness 意味着 DeepSeek 正在从一家单纯“卖Token”(按模型调用量收费)的模型公司,向“交付结果”(按完成的任务价值收费)的解决方案提供商转变。这表明其战略重心开始从模型能力的竞争,延伸到如何让模型能力更高效、稳定地在实际场景中落地。同时,通过将Harness以MIT协议开源并赋能社区,DeepSeek 旨在构建一个开放的生态系统,将定义Agent形态和能力的权限交给开发者和开源社区,这是一种更具野心和更长远的发展路径。

Related Reads

Jameson Lopp's BIP-110 Postmortem: Bitcoin Is Driven by Game Theory, Not Morality

Jameson Lopp's analysis concludes that BIP-110, a proposal to restrict arbitrary data (like inscriptions) on Bitcoin, failed due to economic and technical realities, not moral arguments. The proposal, championed by Luke Dashjr and others, aimed to "cleanse" the chain but never reached its 55% miner activation threshold. Upon its forced signaling deadline in August, only the OCEAN pool (with ~1% hash rate) supported it, creating a short-lived fork that quickly died as miners faced unredeemable block rewards. Lopp's earlier predictions about its economic infeasibility were proven correct. Technically, BIP-110 was flawed; workarounds to embed data compliant with its rules were demonstrated almost immediately, proving it couldn't achieve its stated goal. The debate often devolved into moralistic rhetoric, with supporters accusing opponents of supporting child exploitation material—a tactic Lopp criticizes as ineffective for building consensus. Lopp argues Bitcoin is driven by incentives and game theory, not morality. Past forks like Bitcoin Cash promised economic benefits, while BIP-110 offered only restrictions and reduced miner fees, gaining no substantial support from major economic players. He predicts the "puritans" behind the failed fork will continue complaining but their new chain will remain insignificant. The episode reaffirms that attempting to censor data on a permissionless, anti-censorship network like Bitcoin is a futile battle.

marsbit23m ago

Jameson Lopp's BIP-110 Postmortem: Bitcoin Is Driven by Game Theory, Not Morality

marsbit23m ago

Tokenization Scale Soars to $4.3 Billion, But Why Did Securitize Incur a $5.5 Million Loss?

Securitize's first quarterly report post-IPO reveals a paradox: while its tokenized assets under management hit a record $4.3 billion (up 16% YoY) and platform trading volume surged 147% to $5.3 billion, total revenue fell 5% to $14.4 million. Tokenization revenue specifically dropped ~12% to $7.8 million, leading to an adjusted EBITDA loss of $5.5 million. CFO Francisco Flores explained that most trading volume is not yet monetized, with the majority of tokenization revenue still coming from one-time projects like new protocol integrations. In contrast, asset servicing revenue, a more recurring stream, grew slightly to $6.6 million. The company has lowered its full-year revenue guidance to $70-$80 million from an initial projection of $110 million. Industry experts note this highlights a structural challenge for the tokenization sector. Scaling assets on-chain doesn't automatically scale a profitable business model. Current implementations often rely on costly, customized projects for each new asset or jurisdiction. The future, they argue, lies in building standardized infrastructure that generates recurring "infrastructure revenue" from post-issuance activities like compliance, distributions, and secondary trading—similar to enterprise software. Analysts caution against misinterpreting high trading volumes as indicative of a mature fee-based model, as Securitize's broad volume metric includes many non-monetized actions. The key test for the industry is whether adding billions in new assets can generate sustainable revenue without constant new custom projects.

marsbit25m ago

Tokenization Scale Soars to $4.3 Billion, But Why Did Securitize Incur a $5.5 Million Loss?

marsbit25m ago

The Five Paradoxes of Artificial Intelligence

**Five Paradoxes of Artificial Intelligence** Artificial intelligence (AI) is an era filled with paradoxes, which we navigate as we advance. **1. The Prediction Paradox** AI experts, from pioneers like Marvin Minsky to contemporary figures like Geoffrey Hinton and Demis Hassabis, have a history of inaccurate forecasts regarding AI's capabilities and timelines, such as achieving human-level machine intelligence or surpassing radiologists. Predictions about Artificial General Intelligence (AGI) vary wildly between optimistic entrepreneurs and skeptical academics, highlighting the inherent unpredictability of technological futures. **2. The Employment Quantification Paradox** Despite numerous studies from institutions like the OECD, IMF, and McKinsey attempting to quantify AI's impact on jobs, estimates of affected employment range from 0.4% to 67%, revealing vast inconsistencies. This paradox arises because isolating AI's effect from other economic, social, and technological factors is virtually impossible, and forecasts depend on static assumptions about a dynamically evolving technology. **3. The Productivity Paradox** While AI is a transformative General Purpose Technology, significant productivity growth has not yet materialized in major economies like the EU and has only matched historical averages in the US. This disconnect between rapid innovation and slow productivity gains, reminiscent of the "Solow Paradox" from the computer age, is often explained by time lags. History shows it takes decades for such technologies to diffuse and trigger complementary innovations that boost productivity. **4. The Data Value Paradox** Data is hailed as the "new oil" and critical for AI, yet its economic value is paradoxical. Its worth is realized only in use, not in straightforward trade. Despite policy emphasis and initiatives for data asset recognition on corporate balance sheets in China, the monetized value remains negligible—accounting for only about 0.06% of major telecom operators' total assets—highlighting the gap between perceived utility and financial valuation. **5. The Industrial Revolution Paradox** For decades, nearly every major new technology, from the internet and nanotechnology to blockchain and now AI, has been proclaimed as the driver of a "Fourth Industrial Revolution." This constant reassignment suggests prior labels were premature. True industrial revolutions are typically identified in hindsight, not in real-time. Furthermore, the coexistence of such a proclaimed transformative revolution with ongoing economic crises would be historically anomalous. Whether AI truly defines a new industrial revolution remains a narrative for the future to decide.

marsbit28m ago

The Five Paradoxes of Artificial Intelligence

marsbit28m ago

Trading

Spot

Hot Articles

Discussions

Welcome to the HTX Community. Here, you can stay informed about the latest platform developments and gain access to professional market insights. Users' opinions on the price of S (S) are presented below.

活动图片