DeepSeek V4 Official Version Arrives, New Capabilities Emerge, Value-for-Money King Enters the Fray

marsbitОпубликовано 2026-07-31Обновлено 2026-07-31

Введение

On July 31st, DeepSeek officially launched the public API beta for its DeepSeek-V4-Flash model. A key highlight is its performance on multiple Agent benchmark tests, reportedly nearing or even surpassing the level of the V4-Pro preview version from three months ago. Notably, the Flash model achieves this with significantly smaller scale (130B active parameters vs. Pro's 490B), suggesting that post-training optimization and data quality may be as crucial as raw model size. DeepSeek emphasized that the V4-Flash-0731 uses the same model architecture and size as its preview version, with improvements attributed solely to "re-trained post-training." The update also marks the official debut of DeepSeek's self-developed Agent framework, "Harness." The move signals DeepSeek's strategic push to position its cost-effective Flash model as a competitive base for Agent applications—scenarios requiring autonomous planning, tool usage, and complex task execution—where inference speed and cost are critical. By natively supporting OpenAI's Responses API format and adapting for code-generation scenarios, DeepSeek aims not just to be a cheaper alternative but to establish its own ecosystem in the Agent era. This release follows DeepSeek's record-breaking ~$50 billion fundraising round roughly two months prior, underscoring market confidence in its technology and commercialization prospects. The company is reportedly preparing for another funding round at a valuation of approximately $71 bill...

On the afternoon of July 31st, Phoenix Net Technology discovered upon checking the DeepSeek official website that the DeepSeek-V4-Flash official version API has been launched for public beta testing. Unlike the previous hierarchical logic of “Pro strong, Flash weak”, this update signals a noteworthy shift—the performance of the Flash official version in multiple Agent benchmark tests has approached or even surpassed the level of the V4-Pro preview version from three months ago.

The official update log shows that the V4-Flash official version scored 82.7 points on Terminal Bench 2.1, 54.2 points on NL2Repo, 76.7 points on Cybergym, and 70.3 points on Toolathlon verified. In contrast, the V4-Pro preview version scored 67.9 points on Terminal Bench 2.0.

It is important to note that Terminal Bench 2.0 and 2.1 are not the same version of the test suite, so a direct comparison is not entirely fair. However, for a lightweight version with only 13 billion activation parameters to achieve such scores in Agent capabilities suggests that the optimization space in the post-training phase might offer greater leverage than simply scaling up parameters.

Additionally, DeepSeek also specifically noted that the current public beta is limited to the API, and the latest capabilities are not yet available on the App and web interface. The DeepSeek-V4-Pro official version will be released as soon as possible.

“Only Underwent Post-Training Again”

The DeepSeek official statement in the update log indicates, “The model architecture, size, and parameters of DeepSeek-V4-Flash-0731 remain consistent with DeepSeek-V4-Flash-preview; only post-training was conducted again.”

According to DeepSeek's official technical report, V4-Flash has 284 billion total parameters and 13 billion activation parameters; V4-Pro has 1.6 trillion total parameters and 49 billion activation parameters. The two differ by an order of magnitude in model scale. If Flash can bring its Agent capabilities close to Pro's level through post-training, it implies that for specific tasks, model scale is not the decisive factor—the weight of training methods and data quality is rising.

The official also specially noted that for the Code Agent tasks in the public benchmark tests, the DeepSeek Harness minimal mode was used as the framework for testing, with max setting, topp=0.95, temperature=1.0. This detail suggests that DeepSeek may have made targeted optimizations at the Agent framework level, not just improvements in the model itself.

This is also the first time DeepSeek's self-developed Harness has appeared under an official name. Previously, Liang Wenfeng compared the path to AGI to climbing stairs: language models are the first step, CoT (Chain-of-Thought) was addressed last year, this year's step is Agent, and the problem that must be solved after Agent is continuous learning—enabling models to accumulate experience over time like humans, rather than requiring the full context to be fed in every time to work. Beyond that lies the “singularity” of self-iteration and embodied intelligence. “AI currently lacks not taste or intuition, but the ability for continuous learning,” he said. “Investors look at Agent; we look at how to solve learning.”

Continuous learning sounds like a model-level proposition, but its engineering focal point lies precisely in the Harness. DeepSeek's Agent Harness team was formed in March this year. Leading it is Cui Tianyi, born in the 1990s, a Zhejiang University computer science graduate, holder of six ACM Asia Regional Competition gold medals, who previously worked for nine years at top quantitative firm Jane Street, joined DeepSeek in March this year, and subsequently aggressively recruited in May.

It is reported that DeepSeek plans to launch the Harness concurrently with the V4 official version release. Furthermore, as DeepSeek stated at the end of the log, “The DeepSeek-V4-Pro official version will be released as soon as possible.”

A Direct Confrontation at the Ecosystem Level

If large model competition in 2023 was about “competing on general capabilities” and 2024 was about “competing on long context,” then the keyword for 2026 is undoubtedly Agent.

Agent capability—the model's ability to autonomously plan, call tools, and execute complex tasks—is becoming the new yardstick for measuring large model strength. From Terminal Bench (terminal operations) and NL2Repo (code repository generation) to Cybergym (cybersecurity tasks) and SWE-bench (software engineering), a series of Agent benchmark tests are redefining what makes a good model.

Globally, the first tier of Agent capability is still dominated by closed-source giants. According to data from third-party evaluation platforms like benchlm.ai, GPT-5.6 Sol and the Claude Opus series rank at the top in most Agent benchmark tests. Among domestic players, GLM, Qwen, and others are also catching up quickly.

DeepSeek's significant boost to Flash's Agent capabilities this time has a clear strategic intent: to enter the vast market of Agent applications with a high-value, lightweight model.

After all, Agent scenarios are far more sensitive to inference speed and cost than pure dialogue scenarios. An Agent task requiring repeated tool calls and multi-step reasoning could cost several times or even tens of times more to run on a flagship model compared to Flash. If Flash's Agent capability reaches a level that is “sufficient or even good,” its cost-performance advantage will be highly disruptive.

The two internal test sets officially released also have clear targets. DSBench-FullStack (internal full-stack development test set) scored 68.7, and DSBench-Hard (internal Coding Agent difficult problem test set) scored 59.6. This indirectly confirms DeepSeek's positioning—to build Flash as the preferred foundation for developers and Agent applications.

A quiet ecosystem battle is also brewing. The official V4-Flash natively supports the Responses API format and is specifically adapted for Codex.

The Responses API is a new-generation API format strongly promoted by OpenAI. Compared to the traditional Chat Completions, it is more suitable for Agent scenarios—supporting more flexible tool calls, more complex multi-turn interactions, and finer-grained output control.

DeepSeek's native support for this format means developers can migrate Agent applications developed based on the OpenAI ecosystem to DeepSeek at a lower cost. This will be a direct confrontation between the two at the ecosystem level.

The adaptation for Codex targets the vertical scenario of code generation and software development. Codex is OpenAI's model for the code domain. DeepSeek's targeted adaptation is a direct challenge in OpenAI's traditional area of strength.

Viewed together, these moves indicate that DeepSeek's ambition is not just to be a “cheap alternative” but to establish its own niche in the Agent era.

After the 50 Billion RMB Financing: Time Window Under High Valuation

This update comes less than two months after DeepSeek completed its first round of external financing.

A little over a month ago, DeepSeek completed its first external financing round since its founding nearly three years ago, raising over 50 billion RMB, setting a single-round financing record in China's AI industry, with a post-money valuation of approximately $52 billion (about 350 billion RMB). Investors included industry giants like Tencent, CATL, JD.com, as well as several state-owned industrial funds.

In mid-July, DeepSeek intended to advance a new round of private financing with a pre-money valuation of about $71 billion (approximately 480 billion RMB), a roughly 37% increase from the $52 billion valuation after the first round. The interval from $52 billion to $71 billion was less than six weeks.

Behind the high valuation lies market recognition of DeepSeek's technical strength and bets on its commercialization prospects. However, high valuation also means high expectations and high pressure.

According to multiple media reports, DeepSeek founder Liang Wenfeng personally contributed approximately 20 billion RMB in the first financing round, maintaining firm control of the company through a special structure. This founder, who emerged from the quantitative firm Phantom, had insisted on self-funding for nearly three years previously. The shift from “no financing, no IPO” to actively embracing capital is itself a strong signal—DeepSeek is accelerating toward commercialization and an IPO.

The launch of the Flash official version can perhaps be seen as a technological realization by DeepSeek under the support of capital. But the real test lies ahead: Can the Pro official version arrive on schedule? Can the improvement in Agent capability translate into solid revenue? Under the pressure from giants like OpenAI and Anthropic, how will DeepSeek leverage its high-value route?

The curtain on the Agent war has just risen. DeepSeek, with a significant evolution of a lightweight model, poses a new question to the industry—when the dividends of post-training are fully exploited, when efficiency gains begin to offset the parameter gap, the competitive logic of large models may need to be rewritten.

This article is from the WeChat public account “Phoenix Net Technology,” author: Phoenix Net Technology

Трендовые криптовалюты

Связанные с этим вопросы

QWhat key feature of DeepSeek-V4-Flash's official release is highlighted by the updated benchmark scores, and how does it compare to the V4-Pro preview?

AThe key feature is the significant enhancement of Agent capabilities. Although direct comparison is limited because Terminal Bench 2.0 and 2.1 are different versions, the article notes that the smaller V4-Flash model (with 130B activated parameters) achieved benchmark scores that approach or even surpass those of the much larger V4-Pro preview model (with 490B activated parameters) from three months prior, particularly in Agent-related tests. This suggests that post-training optimizations can be highly effective, challenging the notion that model size alone determines performance.

QAccording to the article, what does the release of the DeepSeek-V4-Flash official version represent in the broader AI model competition landscape?

AThe release represents a strategic move in the Agent-centric competition era of 2026. By boosting the Agent performance of its more cost-effective, lighter Flash model, DeepSeek aims to capture the growing market for Agent applications where inference speed and cost are critical. This positions it as a high-value, 'good enough' alternative to more expensive flagship models like GPT-5.6 Sol and Claude Opus, directly challenging OpenAI and others on their home turf, especially in tool-use and code generation scenarios.

QWhat role does the newly named DeepSeek Harness play, as mentioned by CEO Liang Wenfeng, in the path toward AGI?

ACEO Liang Wenfeng likens the path to AGI to climbing stairs. While 2023 focused on language models and 2024 on long-context, the current step is Agent capability. Beyond that, he identifies 'sustained learning'—the ability for models to accumulate experience like humans without needing full context every time—as the critical next challenge. The DeepSeek Harness, their self-developed Agent framework, is identified as the key engineering tool for enabling this crucial 'sustained learning' capability, making it central to their long-term AGI strategy.

QHow does the DeepSeek-V4-Flash's native support for the Responses API format and Codex adaptation impact its competitive position?

AIt represents a direct ecosystem-level challenge to OpenAI. Native support for OpenAI's Responses API format lowers the barrier for developers to migrate Agent applications built on OpenAI's ecosystem to DeepSeek's platform. Additionally, targeted adaptation for Codex, OpenAI's code-specific model, indicates a direct assault on OpenAI's traditional stronghold of code generation. These moves are part of a strategy for DeepSeek to establish its own ecosystem in the Agent era, moving beyond just being a cost-effective alternative to becoming a primary platform.

QWhat recent financial event for DeepSeek does the article connect to this technical release, and what pressures does it imply?

AThe article connects the V4-Flash release to DeepSeek's recent record-breaking fundraising of over 50 billion RMB (~$5.2 billion post-money valuation) and its subsequent plans for a new round at a valuation of approximately $71 billion. The high valuation reflects market confidence in its technology and commercialization potential but also creates significant pressure to deliver results. This technical release is seen as an initial 'delivery on the promise' following the capital influx. The true test will be whether improved Agent capabilities can translate into real revenue growth and if the upcoming V4-Pro official version can meet high expectations amidst intense competition.

Похожее

«Голос в пустыне» с Уолл-Стрит вновь прицелился в Nvidia

Новая сделка Майкла Берри, известного как «пророк биржевого апокалипсиса» после успешного шорта во время ипотечного кризиса, вновь привлекла внимание. Через свой Substack он объявил о шорте акций Nvidia (по цене входа $198.09), Tesla, Applied Materials, Caterpillar и ETF SOXX, а позднее добавил Micron Technology ($1051.87). 25 июля он увеличил позиции против Nvidia, Micron и SOXX. Его аргументы против Nvidia сосредоточены на проблемах в цепочке поставок ИИ: завышенные сроки амортизации оборудования (6 лет вместо реалистичных 2-3 лет) у таких клиентов, как Microsoft и Meta, что искусственно завышает их прибыль, и риски «внебалансового циклического финансирования», когда спрос на чипы может быть поддержан самой Nvidia через гарантии для покупателей. Третий аргумент о выкупе акций был оспорен компанией как основанный на ошибочных данных. Цена акций Nvidia колебалась после заявлений Берри, в целом оставаясь вблизи его цены входа, что привело к небольшому убытку по его более поздним позициям. История Берри после 2008 года неоднозначна: были как провалы (шорт Tesla в 2021), так и успехи (предупреждение о мемных акциях). Его методология, основанная на анализе свободного денежного потока и первичных документов, хорошо выявляет структурные риски, но плохо определяет время кризиса. Другие известные инвесторы разделяют осторожность, но действуют иначе. Стив Эйсман («Большой шорт») пока не шортит Nvidia, отмечая сильные текущие показатели, но сократил позиции. Джим Чанос согласен с тезисом об «учетном несоответствии», но шортит не чипмейкеров, а частные фонды, сочетающие ставки на ИИ и коммерческую недвижимость. Общий вывод: опытные инвесторы признают признаки перегрева в секторе ИИ, но их конкретные ставки и сроки сильно различаются. Для рядового инвестора ценным является не копирование их сделок, а понимание их методов анализа — куда смотреть, чтобы выявить скрытые риски и завышенные оценки, особенно когда рынок убежден, что «на этот раз все по-другому».

marsbit12 мин. назад

«Голос в пустыне» с Уолл-Стрит вновь прицелился в Nvidia

marsbit12 мин. назад

Когда рынок начинает сомневаться в капитальных затратах на ИИ: полный анализ квартальных отчетов пяти технологических гигантов

В конце июля 2026 года ведущие технологические гиганты — Alphabet (Google), Intel, Microsoft, Meta и Apple — представили квартальные отчёты. Все компании показали рост выручки и прибыли, превысив ожидания рынка, однако реакция инвесторов сильно различалась из-за расхождений в подходах к капитальным затратам (CAPEX) на ИИ и ожиданиях относительно свободного денежного потока. **Alphabet** продемонстрировал рекордный 12-й квартал двузначного роста выручки, а облачный бизнес вырос на 82%. Однако акции упали более чем на 7% после новостей о рекордных квартальных CAPEX в $44,9 млрд и первом в истории отрицательном свободном денежном потоке. Компания также повысила годовой прогноз по CAPEX. **Intel** сообщил о самом сильном росте выручки за 15 лет (25%), во многом благодаря бизнесу центров обработки данных и ИИ. Несмотря на это, акции испытали волатильность после увеличения годового прогноза CAPEX с $18 млрд до более чем $20 млрд, что усилило опасения по поводу денежного потока. **Microsoft** стала фаворитом рынка: выручка Azure впервые превысила $100 млрд в годовом исчислении. Ключевым положительным сигналом стало *снижение* прогноза по CAPEX на 2026 год с $190 млрд до $175 млрд и обещание положительного свободного денежного потока в 2027 финансовом году. Акции выросли, показав лучший однодневный результат за 18 лет. **Meta**, несмотря на рост выручки на 28%, столкнулась с самым резким падением акций (почти 10%). Причина — рост расходов на 55%, резкое сокращение свободного денежного потока на 90% и повышение годового прогноза по CAPEX. Это уже второй квартал подряд, когда Meta наказывают за рост затрат. **Apple** отчиталась о рекордной выручке и прибыли за июньский квартал, но её акции упали более чем на 8%, а рыночная капитализация сократилась на $300 млрд. Это произошло из-за консервативного прогноза на следующий квартал, который оказался ниже ожиданий Уолл-стрит, и опасений по поводу ограничений в цепочке поставок. **Общий вывод:** Рынок смещает фокус с абсолютных показателей выручки и прибыли на **доходность капитальных вложений**. Компании, демонстрирующие контроль над расходами и путь к положительному свободному денежному потоку (как Microsoft), вознаграждаются. Те, кто агрессивно наращивает CAPEX без четких краткосрочных перспектив возврата денежных средств (как Google и Meta), сталкиваются с продажами. В следующем отчётном сезоне объяснение стратегии CAPEX и его окупаемости будет критически важным для котировок акций.

Odaily星球日报26 мин. назад

Когда рынок начинает сомневаться в капитальных затратах на ИИ: полный анализ квартальных отчетов пяти технологических гигантов

Odaily星球日报26 мин. назад

a16z: От компании к DAO, DUNA может стать следующей организационной формой

Автор: a16z crypto Компиляция: Deep Tide TechFlow Основная задача бизнеса на протяжении веков остаётся неизменной: как организовать сотрудничество незнакомых людей с разными ролями, асимметричной информацией и интересами для достижения общей цели? История бизнеса — это история поиска новых организационных форм. Компания стала великим прорывом для индустриальной эпохи, но цифровые технологии и интернет-протоколы снижают потребность в традиционных иерархических структурах. Существующее законодательство не предназначено для нового мира децентрализованных сетей. DAO (децентрализованные автономные организации), хотя и являются перспективной формой, сталкиваются с правовыми проблемами: отсутствием юридического признания, что подвергает участников неограниченной ответственности, и неопределённостью в регулировании, особенно в свете теста Хоуи в США. В качестве решения предлагается DUNA (Децентрализованная Некорпоративная Некоммерческая Ассоциация) — новая правовая форма, уже принятая в нескольких штатах США. DUNA предоставляет группе людей юридическое лицо и ограниченную ответственность, позволяя децентрализованным сообществам легально владеть активами, заключать контракты и вести деятельность, не полагаясь на централизованное управление. Она заполняет правовой вакуум для интернет-нативных организаций, не решая всех проблем управления, но давая им законный статус. Эволюция организационных форм — от семейных предприятий и компаний до LLC и DUNA — расширяет инструменты для координации людей. DUNA представляет собой следующий шаг, позволяющий анонимным, географически разбросанным участникам блокчейн-сетей сотрудничать как единое юридическое лицо, минимизируя личные риски и открывая новую главу в организации коллективных действий.

marsbit1 ч. назад

a16z: От компании к DAO, DUNA может стать следующей организационной формой

marsbit1 ч. назад

Отчет о токенизированных активах реального мира (RWA) за середину 2026 года: рыночная капитализация токенизированных акций удвоилась за год, но 90% прав – пустышка

**Отчет о RWA на блокчейне за середину 2026 года: капитализация токенизированных акций удвоилась за год, но 90% прав — пустышка** В июле 2026 года стоимость распределенных токенизированных акций достигла $1,89 млрд, почти удвоившись с марта. Однако рост сконцентрирован вокруг немногих инструментов (SECZ, FGRS, STRCx дали ~49% прироста) и платформ (Ondo, xStocks, Securitize контролируют 85% рынка). Рынок расколот: офшорные продукты (Ondo, xStocks) обеспечивают ликвидность и доступность в DeFi, но предлагают слабые юридические права на базовый актив. Регулируемая инфраструктура США (например, через Nasdaq) обеспечивает юридическую определенность, но страдает от низкой ликвидности и ограниченного распределения. Ни один продукт не сочетает полноценные права собственности, широкое распределение, ликвидность и независимое ценовое образование. Общие данные по RWA ($36,8 млрд распределенных активов) требуют осторожной интерпретации из-за изменений в методологии учета. Рынок токенизированных акций расширяется технически, но остается фрагментированным в правовом поле, где разные токены, отслеживающие одну акцию, представляют собой отдельные юридические обязательства.

marsbit2 ч. назад

Отчет о токенизированных активах реального мира (RWA) за середину 2026 года: рыночная капитализация токенизированных акций удвоилась за год, но 90% прав – пустышка

marsbit2 ч. назад

Внимание, пользователи биткоина! Сегодняшний взлом может быть масштабнее, чем вы думали. Вот что нужно делать

Компания Coinkite, производитель аппаратных биткоин-кошельков Coldcard, обнаружила критическую уязвимость в процессе генерации сид-фраз на устройствах Coldcard Mk3 с прошивкой версий 4.0.1–5.0.3. Все биткоин-адреса, созданные из этих сид-фраз, потенциально скомпрометированы. Компания настоятельно рекомендует пользователям немедленно перевести средства на новый, безопасный кошелёк, создав новую сид-фразу. Просто сменить устройство недостаточно — старую сид-фразу использовать нельзя. Это предупреждение следует на фоне масштабной атаки 30 июля, в ходе которой автоматизированный инструмент за 41 минуту опустошил 1196 кошельков, похитив около 1082 BTC (на тот момент ~$70,2 млн). Исследователи связывают эту атаку с обнаруженной уязвимостью, так как она произошла за 30 часов до публичного раскрытия проблемы Coinkite. Украденные средства пока остаются на четырёх адресах без движения.

cryptonews.ru4 ч. назад

Внимание, пользователи биткоина! Сегодняшний взлом может быть масштабнее, чем вы думали. Вот что нужно делать

cryptonews.ru4 ч. назад

Торговля

Спот

Популярные статьи

Как купить T

Добро пожаловать на HTX.com! Мы сделали приобретение Threshold Network Token (T) простым и удобным. Следуйте нашему пошаговому руководству и отправляйтесь в свое крипто-путешествие.Шаг 1: Создайте аккаунт на HTXИспользуйте свой адрес электронной почты или номер телефона, чтобы зарегистрироваться и бесплатно создать аккаунт на HTX. Пройдите удобную регистрацию и откройте для себя весь функционал.Создать аккаунтШаг 2: Перейдите в Купить криптовалюту и выберите свой способ оплатыКредитная/Дебетовая Карта: Используйте свою карту Visa или Mastercard для мгновенной покупки Threshold Network Token (T).Баланс: Используйте средства с баланса вашего аккаунта HTX для простой торговли.Третьи Лица: Мы добавили популярные способы оплаты, такие как Google Pay и Apple Pay, для повышения удобства.P2P: Торгуйте напрямую с другими пользователями на HTX.Внебиржевая Торговля (OTC): Мы предлагаем индивидуальные услуги и конкурентоспособные обменные курсы для трейдеров.Шаг 3: Хранение Threshold Network Token (T)После приобретения вами Threshold Network Token (T) храните их в своем аккаунте на HTX. В качестве альтернативы вы можете отправить их куда-либо с помощью перевода в блокчейне или использовать для торговли с другими криптовалютами.Шаг 4: Торговля Threshold Network Token (T)С легкостью торгуйте Threshold Network Token (T) на спотовом рынке HTX. Просто зайдите в свой аккаунт, выберите торговую пару, совершайте сделки и следите за ними в режиме реального времени. Мы предлагаем удобный интерфейс как для начинающих, так и для опытных трейдеров.

985 просмотров всегоОпубликовано 2024.03.29Обновлено 2026.06.02

Как купить T

Обсуждения

Добро пожаловать в Сообщество HTX. Здесь вы сможете быть в курсе последних новостей о развитии платформы и получить доступ к профессиональной аналитической информации о рынке. Мнения пользователей о цене на T (T) представлены ниже.

活动图片