DeepSeek V4 Official Version Arrives, New Capabilities Emerge, Value-for-Money King Enters the Fray

marsbitPublié le 2026-07-31Dernière mise à jour le 2026-07-31

Résumé

On July 31st, DeepSeek officially launched the public API beta for its DeepSeek-V4-Flash model. A key highlight is its performance on multiple Agent benchmark tests, reportedly nearing or even surpassing the level of the V4-Pro preview version from three months ago. Notably, the Flash model achieves this with significantly smaller scale (130B active parameters vs. Pro's 490B), suggesting that post-training optimization and data quality may be as crucial as raw model size. DeepSeek emphasized that the V4-Flash-0731 uses the same model architecture and size as its preview version, with improvements attributed solely to "re-trained post-training." The update also marks the official debut of DeepSeek's self-developed Agent framework, "Harness." The move signals DeepSeek's strategic push to position its cost-effective Flash model as a competitive base for Agent applications—scenarios requiring autonomous planning, tool usage, and complex task execution—where inference speed and cost are critical. By natively supporting OpenAI's Responses API format and adapting for code-generation scenarios, DeepSeek aims not just to be a cheaper alternative but to establish its own ecosystem in the Agent era. This release follows DeepSeek's record-breaking ~$50 billion fundraising round roughly two months prior, underscoring market confidence in its technology and commercialization prospects. The company is reportedly preparing for another funding round at a valuation of approximately $71 bill...

On the afternoon of July 31st, Phoenix Net Technology discovered upon checking the DeepSeek official website that the DeepSeek-V4-Flash official version API has been launched for public beta testing. Unlike the previous hierarchical logic of “Pro strong, Flash weak”, this update signals a noteworthy shift—the performance of the Flash official version in multiple Agent benchmark tests has approached or even surpassed the level of the V4-Pro preview version from three months ago.

The official update log shows that the V4-Flash official version scored 82.7 points on Terminal Bench 2.1, 54.2 points on NL2Repo, 76.7 points on Cybergym, and 70.3 points on Toolathlon verified. In contrast, the V4-Pro preview version scored 67.9 points on Terminal Bench 2.0.

It is important to note that Terminal Bench 2.0 and 2.1 are not the same version of the test suite, so a direct comparison is not entirely fair. However, for a lightweight version with only 13 billion activation parameters to achieve such scores in Agent capabilities suggests that the optimization space in the post-training phase might offer greater leverage than simply scaling up parameters.

Additionally, DeepSeek also specifically noted that the current public beta is limited to the API, and the latest capabilities are not yet available on the App and web interface. The DeepSeek-V4-Pro official version will be released as soon as possible.

“Only Underwent Post-Training Again”

The DeepSeek official statement in the update log indicates, “The model architecture, size, and parameters of DeepSeek-V4-Flash-0731 remain consistent with DeepSeek-V4-Flash-preview; only post-training was conducted again.”

According to DeepSeek's official technical report, V4-Flash has 284 billion total parameters and 13 billion activation parameters; V4-Pro has 1.6 trillion total parameters and 49 billion activation parameters. The two differ by an order of magnitude in model scale. If Flash can bring its Agent capabilities close to Pro's level through post-training, it implies that for specific tasks, model scale is not the decisive factor—the weight of training methods and data quality is rising.

The official also specially noted that for the Code Agent tasks in the public benchmark tests, the DeepSeek Harness minimal mode was used as the framework for testing, with max setting, topp=0.95, temperature=1.0. This detail suggests that DeepSeek may have made targeted optimizations at the Agent framework level, not just improvements in the model itself.

This is also the first time DeepSeek's self-developed Harness has appeared under an official name. Previously, Liang Wenfeng compared the path to AGI to climbing stairs: language models are the first step, CoT (Chain-of-Thought) was addressed last year, this year's step is Agent, and the problem that must be solved after Agent is continuous learning—enabling models to accumulate experience over time like humans, rather than requiring the full context to be fed in every time to work. Beyond that lies the “singularity” of self-iteration and embodied intelligence. “AI currently lacks not taste or intuition, but the ability for continuous learning,” he said. “Investors look at Agent; we look at how to solve learning.”

Continuous learning sounds like a model-level proposition, but its engineering focal point lies precisely in the Harness. DeepSeek's Agent Harness team was formed in March this year. Leading it is Cui Tianyi, born in the 1990s, a Zhejiang University computer science graduate, holder of six ACM Asia Regional Competition gold medals, who previously worked for nine years at top quantitative firm Jane Street, joined DeepSeek in March this year, and subsequently aggressively recruited in May.

It is reported that DeepSeek plans to launch the Harness concurrently with the V4 official version release. Furthermore, as DeepSeek stated at the end of the log, “The DeepSeek-V4-Pro official version will be released as soon as possible.”

A Direct Confrontation at the Ecosystem Level

If large model competition in 2023 was about “competing on general capabilities” and 2024 was about “competing on long context,” then the keyword for 2026 is undoubtedly Agent.

Agent capability—the model's ability to autonomously plan, call tools, and execute complex tasks—is becoming the new yardstick for measuring large model strength. From Terminal Bench (terminal operations) and NL2Repo (code repository generation) to Cybergym (cybersecurity tasks) and SWE-bench (software engineering), a series of Agent benchmark tests are redefining what makes a good model.

Globally, the first tier of Agent capability is still dominated by closed-source giants. According to data from third-party evaluation platforms like benchlm.ai, GPT-5.6 Sol and the Claude Opus series rank at the top in most Agent benchmark tests. Among domestic players, GLM, Qwen, and others are also catching up quickly.

DeepSeek's significant boost to Flash's Agent capabilities this time has a clear strategic intent: to enter the vast market of Agent applications with a high-value, lightweight model.

After all, Agent scenarios are far more sensitive to inference speed and cost than pure dialogue scenarios. An Agent task requiring repeated tool calls and multi-step reasoning could cost several times or even tens of times more to run on a flagship model compared to Flash. If Flash's Agent capability reaches a level that is “sufficient or even good,” its cost-performance advantage will be highly disruptive.

The two internal test sets officially released also have clear targets. DSBench-FullStack (internal full-stack development test set) scored 68.7, and DSBench-Hard (internal Coding Agent difficult problem test set) scored 59.6. This indirectly confirms DeepSeek's positioning—to build Flash as the preferred foundation for developers and Agent applications.

A quiet ecosystem battle is also brewing. The official V4-Flash natively supports the Responses API format and is specifically adapted for Codex.

The Responses API is a new-generation API format strongly promoted by OpenAI. Compared to the traditional Chat Completions, it is more suitable for Agent scenarios—supporting more flexible tool calls, more complex multi-turn interactions, and finer-grained output control.

DeepSeek's native support for this format means developers can migrate Agent applications developed based on the OpenAI ecosystem to DeepSeek at a lower cost. This will be a direct confrontation between the two at the ecosystem level.

The adaptation for Codex targets the vertical scenario of code generation and software development. Codex is OpenAI's model for the code domain. DeepSeek's targeted adaptation is a direct challenge in OpenAI's traditional area of strength.

Viewed together, these moves indicate that DeepSeek's ambition is not just to be a “cheap alternative” but to establish its own niche in the Agent era.

After the 50 Billion RMB Financing: Time Window Under High Valuation

This update comes less than two months after DeepSeek completed its first round of external financing.

A little over a month ago, DeepSeek completed its first external financing round since its founding nearly three years ago, raising over 50 billion RMB, setting a single-round financing record in China's AI industry, with a post-money valuation of approximately $52 billion (about 350 billion RMB). Investors included industry giants like Tencent, CATL, JD.com, as well as several state-owned industrial funds.

In mid-July, DeepSeek intended to advance a new round of private financing with a pre-money valuation of about $71 billion (approximately 480 billion RMB), a roughly 37% increase from the $52 billion valuation after the first round. The interval from $52 billion to $71 billion was less than six weeks.

Behind the high valuation lies market recognition of DeepSeek's technical strength and bets on its commercialization prospects. However, high valuation also means high expectations and high pressure.

According to multiple media reports, DeepSeek founder Liang Wenfeng personally contributed approximately 20 billion RMB in the first financing round, maintaining firm control of the company through a special structure. This founder, who emerged from the quantitative firm Phantom, had insisted on self-funding for nearly three years previously. The shift from “no financing, no IPO” to actively embracing capital is itself a strong signal—DeepSeek is accelerating toward commercialization and an IPO.

The launch of the Flash official version can perhaps be seen as a technological realization by DeepSeek under the support of capital. But the real test lies ahead: Can the Pro official version arrive on schedule? Can the improvement in Agent capability translate into solid revenue? Under the pressure from giants like OpenAI and Anthropic, how will DeepSeek leverage its high-value route?

The curtain on the Agent war has just risen. DeepSeek, with a significant evolution of a lightweight model, poses a new question to the industry—when the dividends of post-training are fully exploited, when efficiency gains begin to offset the parameter gap, the competitive logic of large models may need to be rewritten.

This article is from the WeChat public account “Phoenix Net Technology,” author: Phoenix Net Technology

Cryptos en tendance

Questions liées

QWhat key feature of DeepSeek-V4-Flash's official release is highlighted by the updated benchmark scores, and how does it compare to the V4-Pro preview?

AThe key feature is the significant enhancement of Agent capabilities. Although direct comparison is limited because Terminal Bench 2.0 and 2.1 are different versions, the article notes that the smaller V4-Flash model (with 130B activated parameters) achieved benchmark scores that approach or even surpass those of the much larger V4-Pro preview model (with 490B activated parameters) from three months prior, particularly in Agent-related tests. This suggests that post-training optimizations can be highly effective, challenging the notion that model size alone determines performance.

QAccording to the article, what does the release of the DeepSeek-V4-Flash official version represent in the broader AI model competition landscape?

AThe release represents a strategic move in the Agent-centric competition era of 2026. By boosting the Agent performance of its more cost-effective, lighter Flash model, DeepSeek aims to capture the growing market for Agent applications where inference speed and cost are critical. This positions it as a high-value, 'good enough' alternative to more expensive flagship models like GPT-5.6 Sol and Claude Opus, directly challenging OpenAI and others on their home turf, especially in tool-use and code generation scenarios.

QWhat role does the newly named DeepSeek Harness play, as mentioned by CEO Liang Wenfeng, in the path toward AGI?

ACEO Liang Wenfeng likens the path to AGI to climbing stairs. While 2023 focused on language models and 2024 on long-context, the current step is Agent capability. Beyond that, he identifies 'sustained learning'—the ability for models to accumulate experience like humans without needing full context every time—as the critical next challenge. The DeepSeek Harness, their self-developed Agent framework, is identified as the key engineering tool for enabling this crucial 'sustained learning' capability, making it central to their long-term AGI strategy.

QHow does the DeepSeek-V4-Flash's native support for the Responses API format and Codex adaptation impact its competitive position?

AIt represents a direct ecosystem-level challenge to OpenAI. Native support for OpenAI's Responses API format lowers the barrier for developers to migrate Agent applications built on OpenAI's ecosystem to DeepSeek's platform. Additionally, targeted adaptation for Codex, OpenAI's code-specific model, indicates a direct assault on OpenAI's traditional stronghold of code generation. These moves are part of a strategy for DeepSeek to establish its own ecosystem in the Agent era, moving beyond just being a cost-effective alternative to becoming a primary platform.

QWhat recent financial event for DeepSeek does the article connect to this technical release, and what pressures does it imply?

AThe article connects the V4-Flash release to DeepSeek's recent record-breaking fundraising of over 50 billion RMB (~$5.2 billion post-money valuation) and its subsequent plans for a new round at a valuation of approximately $71 billion. The high valuation reflects market confidence in its technology and commercialization potential but also creates significant pressure to deliver results. This technical release is seen as an initial 'delivery on the promise' following the capital influx. The true test will be whether improved Agent capabilities can translate into real revenue growth and if the upcoming V4-Pro official version can meet high expectations amidst intense competition.

Lectures associées

Ne pas investir n'est pas un laissez-passer pour Apple

Face aux géants technologiques comme Meta et Google qui font face à des critiques pour leurs dépenses d'investissement massives dans l'IA, Apple, bien qu'en retard dans ce domaine, se distingue par sa retenue budgétaire. Cette approche lui a même permis de retrouver brièvement la première place mondiale en termes de valorisation boursière. Le rapport trimestriel (T3 2026) d'Apple affiche des performances solides, avec un chiffre d'affaires en hausse de 16,4% et un bénéfice net en progression de 27,1%. L'iPhone et le Mac sont les principaux moteurs de cette croissance, compensant les résultats plus modestes de l'iPad (en baisse) et des services (ralentissement de la croissance). Cependant, le marché réagit négativement après la publication des résultats, en raison des perspectives prudentes pour le trimestre suivant. Apple anticipe des contraintes d'approvisionnement majeures, notamment pour les puces et la mémoire, entraînant des hausses de prix sur ses produits. Contrairement à ses concurrents qui investissent des milliards dans l'infrastructure IA, Apple maintient des dépenses d'investissement (capex) faibles, privilégiant les dépenses de R&D. Malgré cela, l'entreprise n'échappe pas aux répercussions de la frénésie de l'IA, qui exacerbe les tensions sur sa chaîne d'approvisionnement. Ce rapport marque la dernière conférence téléphonique de Tim Cook en tant que PDG, avant son départ prévu en septembre. Il exprime sa confiance dans l'avenir de l'entreprise.

marsbitIl y a 42 mins

Ne pas investir n'est pas un laissez-passer pour Apple

marsbitIl y a 42 mins

PA Figure | Une infographie pour comprendre les grands événements de l'écosystème Web3 en août 2026

Août 2026 s'annonce chargé pour l'écosystème Web3, marqué par plusieurs événements clés susceptibles d'influencer les marchés. L'agenda macroéconomique sera déterminant, avec la publication des données américaines sur l'emploi (non-farm payrolls) et l'inflation (CPI) de juillet, ainsi que les comptes-rendus de la Réserve Fédérale et le symposium annuel de Jackson Hole. Sur le front réglementaire, des développements majeurs sont attendus : le Sénat américain doit dévoiler un nouveau projet de loi (*CLARITY Act*), tandis que l'interdiction des transactions cryptos de l'UE envers la Biélorussie entre en vigueur. Les marchés devront également absorber des déblocages massifs de jetons (**ENA, AVAX, CONX, ZRO, KAITO**, etc.), susceptibles de créer de la volatilité. Par ailleurs, l'industrie continuera son consolidation, avec l'arrêt ou la refonte prévus de plusieurs services comme Exchange Art, Ctrl Wallet, Zapper, NFTfi et Summer.fi, incitant les utilisateurs à gérer leurs actifs en conséquence. Du côté des entreprises, les résultats du Q2 de **SpaceX, Circle et Nvidia** seront publiés, et des levées de fonds importantes sont au programme, notamment pour la société chinoise Moonshot AI (pré-IPO). Enfin, des événements sectoriels majeurs comme **Bitcoin Asia 2026** et le **China Digital Expo 2026** se tiendront. En résumé, le mois d'août sera structuré autour des anticipations macroéconomiques, de l'évolution réglementaire, des déblocages de tokens et de la consolidation continue du secteur.

marsbitIl y a 55 mins

PA Figure | Une infographie pour comprendre les grands événements de l'écosystème Web3 en août 2026

marsbitIl y a 55 mins

La voix la plus célèbre de « Cassandre » de Wall Street s'attaque cette fois à NVIDIA

Une récente divulgation de l’investisseur Michael Burry, connu pour avoir prédit la crise des subprimes et immortalisé dans « The Big Short », a relancé les débats sur le marché. Fin juin, il a annoncé avoir pris des positions à découvert sur plusieurs valeurs technologiques, dont Nvidia (à un prix d’entrée de 198,09 $), Tesla, Applied Materials, Caterpillar et l’ETF SOXX (semiconducteurs). Le 1er juillet, il a ajouté Micron à sa liste, avant d’augmenter fin juillet ses positions sur Nvidia, Micron et SOXX. Ses arguments portent principalement sur les pratiques comptables dans le secteur de l’IA : selon lui, la durée d’amortissement des puces (étirée à 6 ans par les géants du cloud comme Microsoft ou Google) ne reflète pas leur obsolescence rapide (2-3 ans), ce qui gonflerait artificiellement les profits. Il évoque également des risques de « financement circulaire hors bilan », où Nvidia pourrait garantir des prêts à des clients pour qu’ils achètent ses propres puces, créant ainsi une demande artificielle. Enfin, il critique les rachats d’actions de Nvidia, accusés de doper artificiellement le bénéfice par action – une affirmation que la société a contestée en soulignant des erreurs de calcul. Les réactions sont partagées. D’un côté, des voix comme Steve Eisman (autre figure de « The Big Short ») restent prudentes mais ne suivent pas la position découverte, notant la croissance soutenue des revenus et des investissements en IA. De l’autre, Jim Chanos, célèbre vendeur à découvert, partage l’inquiétude sur les écarts comptables mais préfère cibler d’autres acteurs financiers plutôt que les fabricants de puces. Historiquement, les appels de Burry ont connu des succès mitigés : corrects sur des crises structurelles (subprimes, COVID), mais souvent prématurés ou erronés sur des arguments de valorisation (Tesla, Nvidia en 2023). Aujourd’hui, les positions à découvert sur Nvidia restent marginales (environ 1,4 % des actions en circulation), même si les pertes cumulées des vendeurs à découvert sur la valeur dépassent 5 milliards de dollars. Pour les investisseurs, l’intérêt réside moins dans le suivi des positions de Burry que dans la méthodologie sous-jacente : scruter les flux de trésorerie, questionner les traitements comptables agressifs et identifier les risques structurels, surtout quand le marché semble euphorique. La question centrale n’est pas de savoir si Burry a raison cette fois, mais plutôt quels enseignements tirer de son analyse pour évaluer la solidité réelle de la bulle présumée de l’IA.

marsbitIl y a 1 h

La voix la plus célèbre de « Cassandre » de Wall Street s'attaque cette fois à NVIDIA

marsbitIl y a 1 h

Sélection de la semaine丨Émotions épiques sur le marché boursier, l'entrée en bourse de Changxin Tech redéfinit le paysage du stockage, Saylor vise à rétablir l'ancrage du STRC vers le 8 septembre

PANews présente un résumé hebdomadaire de contenus clés. Les marchés ont connu une volatilité significative, avec des indices boursiers sud-coréens subissant plusieurs interruptions et une chute des actions technologiques liées à l'IA. Dans ce contexte, le bitcoin apparaît comme un actif relativement stable. Le paysage technologique évolue rapidement. Le géant de la mémoire Longxin Technology a fait ses débuts en bourse avec une capitalisation importante, symbolisant une percée pour la DRAM nationale. Parallèlement, la convergence de l'IA, des agents autonomes et des besoins en calcul redéfinit les secteurs, des GPU aux infrastructures énergétiques, les entreprises de minage de bitcoin se repositionnant autour de la gestion de l'électricité. Dans le domaine crypto, l'innovation se poursuit. Les portefeuilles intelligents pour agents IA et les mécanismes de paiement programmables gagnent en importance, attirant l'attention de grandes plateformes. De nouveaux modèles économiques, comme le protocole à jeton FWA qui combine NFT et système de tirage, génèrent un fort engagement. Cependant, un décalage est observé entre la croissance des revenus des principaux protocoles et la performance de leurs jetons. Le secteur des Real World Assets (RWA) voit son volume augmenter mais peine à mobiliser ses actifs. Les perspectives macroéconomiques restent mitigées. La Réserve Fédérale américaine maintient des taux directeurs élevés, signalant une orientation durablement restrictive malgré des divisions internes. Des analystes comme Tom Lee considèrent les récentes corrections comme un assainissement nécessaire, affirmant que la logique de long terme du secteur reste intacte. Des personnalités anticipent un rôle accru du bitcoin dans le futur système monétaire. Les informations marquantes incluent : des mouvements réglementaires affectant les courtiers pour investisseurs chinois, le plan de développement technique d'Ethereum à horizon 2030, une importante migration de staking par Lido, et des performances boursières variées pour les entreprises liées à la crypto et à l'IA. Michael Saylor a également annoncé un objectif de réalignement pour le STRC autour du 8 septembre.

marsbitIl y a 1 h

Sélection de la semaine丨Émotions épiques sur le marché boursier, l'entrée en bourse de Changxin Tech redéfinit le paysage du stockage, Saylor vise à rétablir l'ancrage du STRC vers le 8 septembre

marsbitIl y a 1 h

Lorsque le marché commence à s'interroger sur les dépenses en capital liées à l'IA : Analyse complète des résultats du Q2 des cinq géants technologiques

À la fin juillet 2026, les résultats trimestriels d'Alphabet, Intel, Microsoft, Meta et Apple ont tous mis en évidence une croissance robuste des revenus et des bénéfices, largement portée par les investissements et la demande en IA. Cependant, les réactions des investisseurs ont divergé, soulignant un changement d'attention : le marché s'interroge désormais sur le calendrier du retour sur investissement de ces dépenses massives en capital. Alphabet a affiché une croissance record de son chiffre d'affaires et une performance cloud exceptionnelle, mais son augmentation des dépenses d'investissement, entraînant un flux de trésorerie libre négatif pour la première fois, a provoqué une chute de son cours. Intel, malgré ses meilleures ventes depuis quinze ans, a vu son rebond boursier anéanti après l'annonce d'une hausse significative de son budget d'investissement. À l'inverse, Microsoft, en abaissant ses prévisions de dépenses en capital et en promettant des flux de trésorerie positifs, a connu une forte hausse de son action. Meta, dont les dépenses ont explosé et le flux de trésorerie libre s'est effondré, a subi la plus forte vente. Enfin, Apple, malgré des résultats records, a vu son action chuter en raison de prévisions inférieures aux attentes, pointant des contraintes d'approvisionnement. En résumé, la demande en IA reste solide, mais le marché sanctionne désormais les entreprises dont les investissements massifs menacent à court terme la rentabilité et les flux de trésorerie, récompensant celles qui maîtrisent leur trajectoire financière.

Odaily星球日报Il y a 1 h

Lorsque le marché commence à s'interroger sur les dépenses en capital liées à l'IA : Analyse complète des résultats du Q2 des cinq géants technologiques

Odaily星球日报Il y a 1 h

Trading

Spot

Articles tendance

Comment acheter T

Bienvenue sur HTX.com ! Nous vous permettons d'acheter Threshold Network Token (T) de manière simple et pratique. Suivez notre guide étape par étape pour commencer votre parcours crypto.Étape 1 : Création de votre compte HTXUtilisez votre adresse e-mail ou votre numéro de téléphone pour ouvrir un compte sur HTX gratuitement. L'inscription se fait en toute simplicité et débloque toutes les fonctionnalités.Créer mon compteÉtape 2 : Choix du mode de paiement (rubrique Acheter des cryptosCarte de crédit/débit : utilisez votre carte Visa ou Mastercard pour acheter instantanément Threshold Network Token (T).Solde :utilisez les fonds du solde de votre compte HTX pour trader en toute simplicité.Prestataire tiers :pour accroître la commodité d'utilisation, nous avons ajouté des modes de paiement populaires tels que Google Pay et Apple Pay.P2P :tradez directement avec d'autres utilisateurs sur HTX.OTC (de gré à gré) : nous offrons des services personnalisés et des taux de change compétitifs aux traders.Étape 3 : stockage de vos Threshold Network Token (T)Après avoir acheté vos Threshold Network Token (T), stockez-les sur votre compte HTX. Vous pouvez également les envoyer ailleurs via un transfert sur la blockchain ou les utiliser pour trader d'autres cryptos.Étape 4 : tradez des Threshold Network Token (T)Tradez facilement Threshold Network Token (T) sur le marché Spot de HTX. Il vous suffit d'accéder à votre compte, de sélectionner la paire de trading, d'exécuter vos trades et de les suivre en temps réel. Nous offrons une expérience conviviale aux débutants comme aux traders chevronnés.

678 vues totalesPublié le 2024.12.10Mis à jour le 2026.06.02

Comment acheter T

Discussions

Bienvenue dans la Communauté HTX. Ici, vous pouvez vous tenir informé(e) des derniers développements de la plateforme et accéder à des analyses de marché professionnelles. Les opinions des utilisateurs sur le prix de T (T) sont présentées ci-dessous.

活动图片