Liang Wenfeng Launches Surprise Attack on Musk: DeepSeek V4 Pro vs. Grok 4.6, First Tests Are Spectacular

marsbitОпубліковано о 2026-08-12Востаннє оновлено о 2026-08-12

Анотація

**DeepSeek V4 Pro vs. Grok 4.6: A Benchmark and Real-World Showdown** In a dramatic AI clash, DeepSeek V4 Pro and Elon Musk's Grok 4.6 launched head-to-head, targeting long-task capabilities like tool calling and code validation. **Benchmark Performance:** DeepSeek V4 Pro excelled in Agent tests, topping CyberGym (83.3) and AutomationBench (31.8), surpassing models like Fable 5 and Claude Opus 4.8. It closed the gap on Terminal-Bench 2.1 (87.9) and saw massive gains in DeepSWE (software engineering). Grok 4.6 matched GPT-5.6 in overall ability (61 index), outperformed it in coding (CursorBench 69.9%), and led in three knowledge-work evaluations, including legal tasks (Harvey LAB, 15.8%). **Pricing Disruption:** DeepSeek's standout feature is its radically low cost: $0.87 per million output tokens—roughly 1/7th of Grok 4.6 ($6) and a tiny fraction of rivals like Fable 5 ($50). **Real-World Tests:** In practical challenges, both models demonstrated impressive skill. DeepSeek V4 Pro built a 3D interactive Earth, websites, 60 design styles, and a polished Flappy Bird clone (costing only $0.019 vs. Grok's $0.03 for a simpler version). Head-to-head comparisons were tight: Grok 4.6 won on some design aesthetics, while DeepSeek often delivered higher visual detail and completeness. However, in complex tasks like generating a Three.js cherry blossom tree, V4 Pro lagged behind GPT-5.6 Sol and Claude Opus 5. **Conclusion:** The simultaneous release signals a shift: high-end AI capa...

Today, this showdown is simply surreal!

On one side, Liang Wenfeng finally released the official version of DeepSeek V4 Pro in the early hours.

On the other, Musk dropped the next-generation flagship Grok 4.6, touting low cost and high performance.

Two giants, releasing at the same time, instantly filling the air with gunpowder.

They are targeting almost the same thing—

enabling models to continuously call tools, modify code, verify results, and finally deliver a truly usable product within long-running tasks.

DeepSeek V4 Pro Arrives

Musk's Grok 4.6 Takes the Stage

The most dazzling achievement of the official DeepSeek V4 Pro is securing absolute first place in two evaluations, directly countering Fable 5.

It scored 83.3 points in Cybersecurity Agent Test CyberGym, surpassing Fable 5's 83.1 and Opus 4.8's 78.3.

In Automation Task AutomationBench, it scored 31.8 points, pushing down Fable 5's 29.1 and Opus 4.8's 27.2.

The most "in-your-face" moment occurred in Terminal-Bench 2.1.

DeepSeek scored 87.9 points, exceeding Opus 4.8's 85.0 and trailing Fable 5's 88.0 by only 0.1.

In several other high-difficulty Agent tests, DeepSeek achieved a complete overtaking of Opus 4.8.

In the tool-assisted "Final Human Exam", it scored 60.0 points, surpassing Opus 4.8's 57.9 and continuing to close in on Fable 5's 63.0.

Even in the previously weakest area of software engineering agents, DeepSWE surged from 12.8 points in the preview version to 62.7 points, nearly 4.9 times the original.

This score not only exceeded Opus 4.8's 58.0 but also left only a 7.3-point gap to Fable 5's 70.0.

Simultaneously, Musk also put Fable 5 and GPT-5.6 Sol on the hot seat.

Grok 4.6 first caught up with GPT-5.6 in comprehensive capabilities, then achieved consecutive overtakes in coding tests.

It scored 61 points in General Intelligence Index, just 1 point behind Fable 5's 62.

On CursorBench, it scored 69.9%, surpassing GPT-5.6's 67.2% and only 0.6 percentage points behind Fable 5's 70.5%.

On FrontierCode, it scored 61.3%, also edging out GPT-5.6's 60.6%, continuing to closely trail Fable 5's 63.6%.

In knowledge work closer to real workplace deliverables, Grok 4.6 directly turned the tables, securing first place in all three evaluations.

In GDPval-AA v2, it scored 1753 Elo, surpassing Fable 5's 1741 and GPT-5.6's 1728.

On AA-Briefcase, it scored another 1577 Elo, beating Fable 5's 1574 and GPT-5.6's 1502.

In the specialized legal task Harvey LAB, it scored 15.8%, while Fable 5 only managed 11.3%, and GPT-5.6 a mere 2.5%.

What's even more impressive is that DeepSeek V4 Pro and Grok 4.6 not only match the performance of OpenAI and Anthropic's top-tier models but also jointly bring down the price of cutting-edge intelligence.

Per million output tokens, Grok 4.6 costs $6, GPT-5.6 Sol costs $30, Claude Opus 5 costs $25, Fable 5 costs $50.

And DeepSeek costs only $0.87—

approximately 1/7th of Grok 4.6, 1/35th of GPT-5.6 Sol, 1/29th of Claude Opus 5, and 1/57th of Fable 5!

World's First Hands-on Test

DeepSeek vs. Grok

Now, the first batch of real-world tests is out. DeepSeek V4 Pro and Grok 4.6 have finally moved from benchmark charts into the field of real tasks.

Round one, we directly put pressure on DeepSeek.

With just one prompt, it built a complete 3D interactive Earth from scratch in a browser.

Basic interactions like dragging, rotating, and zooming were all functional; global data flows, dynamic flight paths, and geographic markers were fully laid out on the Earth's surface.

Looking deeper, atmospheric scattering, cloud rendering, day-night lighting, and even the entire UI interface were all included.

Next, we pitted the two models directly against each other in the same arena.

AI blogger "Xiangyang Qiaomu" first ran three small tasks with DeepSeek V4 Pro, then gave two of the same prompts to Grok 4.6, directly comparing the final products.

The first task was to call 3 Skills to develop and deploy a website.

DeepSeek smoothly ran through the entire process; judging by page design and completeness, its overall performance was very stable.

The second task was to replicate 60 design styles and generate a complete set of Bento cards for centralized display.

In this round, DeepSeek made clear distinctions in fonts, color schemes, and layouts, with an overall visual effect that was quite impressive.

The final task was to generate a 3D brick-breaker game from zero.

The game was not only playable but also featured a 3D scene, background music, and dynamic sound effects, with operational feedback and playability fully present.

Subsequently, the same prompts were given to Grok 4.6.

In the 3D brick-breaker round, the two models were almost evenly matched.

The game generated by Grok also had good quality; whether in visual effects, scene completeness, or playability, it was on par with DeepSeek V4 Pro.

In the 60 Bento designs round, Grok 4.6 pulled back a point.

Some of the pages it generated were more mature in layout, color scheme, and visual hierarchy, resulting in a more aesthetically pleasing final product.

An even fiercer duel occurred with Flappy Bird.

Developer Jun Song gave the exact same prompt, asking DeepSeek V4 Pro and Grok 4.6 to respectively "handcraft" a game from scratch.

As a result, the two models took completely different paths.

DeepSeek V4 Pro consumed over 20,000 tokens at a cost of only $0.019; Grok 4.6 used about 5,000 tokens but cost $0.03.

However, judging by the final product, DeepSeek clearly won this round.

The game featured distant mountain ranges and layered clouds; the pipes had gradients and a sense of volume. When the character passed through a pipe, a floating "+1" animation even popped up.

From scene layering to operational feedback, the details were almost all maxed out; the completeness of the end-to-end build was noticeably higher than Grok 4.6's.

In another front-end test conducted by developer Hamza, DeepSeek V4 Pro once again bested Grok 4.6.

Whether in page completeness or final visual effects, the product delivered by V4 Pro was superior.

However, V4 Pro did not reign supreme in every test.

In a pelican comparison test, compared to the Flash version, V4 Pro's overall image completeness was higher, and the elephant's shape was more aesthetically pleasing.

The only problem was—the pelican's direction of movement was completely drawn backward.

Increasing the difficulty further, asking DeepSeek V4 Pro, GPT-5.6 Sol, and Claude Opus 5 to generate the same cherry blossom tree using Three.js, the gap became more apparent.

Whether in the details of the trunk and branches or in lighting, depth of field, and overall atmosphere, V4 Pro fell short of the other two top-tier models.

DeepSeek V4 Pro; GPT-5.6 Sol; Claude Opus 5

After several rounds of hands-on tests, it's clear that DeepSeek V4 Pro and Grok 4.6 are indeed on the same level, neck and neck.

Two Major AIs, Same Day

Caught Up with OpenAI and Anthropic

Regarding the simultaneous breakthroughs of these two models, expert Rick De Oliveira gave the evaluation: "Powerful, and affordable."

With the empowerment of DeepSeek V4 Pro and Grok 4.6, these two words that were almost impossible to appear together have, for the first time, truly come together today.

From cybersecurity and knowledge-based tasks to Agents autonomously solving problems, they have already shattered the ceiling of cutting-edge capabilities in multiple battlefields, all at a lower cost.

Looking back at today's surreal AI "clash," DeepSeek and Grok have jointly rewritten a rule of the game:

Top-tier intelligence is no longer out of reach.

In the past, developers often had to make difficult choices between "can't afford it" and "not powerful enough."

Today, DeepSeek overturned the table with a price tag mere fractions of SOTA, while Grok followed closely with extremely high cost-effectiveness.

This frontal assault of "high performance at low price" has already torn a rift in the old order.

And the war is not over.

Musk has already previewed that Grok 4.7 is coming soon, aiming to surpass all top AIs; counterattacks from OpenAI and Anthropic are also bound to follow.

In this era evolving by the "hour," intelligence is no longer a privilege of a few giants. The explosion of applications for everyone has only just begun.

References:

https://api-docs.deepseek.com/zh-cn/quick_start/pricing/

https://x.ai/news/grok-4-6

https://x.com/vista8/status/2087577905081823559

https://x.com/vista8/status/2087566666477744620

https://x.com/jun_song/status/2087602149979254914

This article is from the WeChat public account "Xin Zhi Yuan," author: ASI Apocalypse, editors: Moses Peach

Трендові криптовалюти

Пов'язані питання

QWhat are the key performance benchmarks where DeepSeek V4 Pro outperformed other major AI models?

ADeepSeek V4 Pro achieved top scores in CyberGym (83.3), AutomationBench (31.8), and closely trailed Fable 5 in Terminal-Bench 2.1 (87.9 vs 88.0). It also showed significant improvement in DeepSWE, scoring 62.7, and outperformed Opus 4.8 in the 'Human Final Exam' with tools (60.0).

QHow does the pricing of DeepSeek V4 Pro compare to other leading AI models?

ADeepSeek V4 Pro is priced at $0.87 per million output tokens, which is approximately 1/7th of Grok 4.6's cost, 1/35th of GPT-5.6 Sol, 1/29th of Claude Opus 5, and 1/57th of Fable 5.

QIn which specific practical tests did DeepSeek V4 Pro demonstrate superior performance over Grok 4.6?

AIn head-to-head tests, DeepSeek V4 Pro generated a more detailed and visually complete Flappy Bird game with better scene hierarchy and feedback. It also produced a higher completion rate and visual effect in a front-end development test compared to Grok 4.6.

QWhat were Grok 4.6's notable achievements in benchmarks against GPT-5.6 Sol and Fable 5?

AGrok 4.6 matched GPT-5.6 Sol in comprehensive capability (61 points), outperformed it in CursorBench (69.9% vs 67.2%) and FrontierCode (61.3% vs 60.6%). It also led in three knowledge work evaluations: GDPval-AA v2 (1753 Elo), AA-Briefcase (1577 Elo), and the professional legal task Harvey LAB (15.8%).

QAccording to the article, what is the broader significance of the simultaneous release of DeepSeek V4 Pro and Grok 4.6?

AThe simultaneous release signifies a shift in the AI landscape where 'powerful and affordable' are no longer mutually exclusive. Both models challenge the high-cost barrier of top-tier AI by offering near-state-of-the-art performance at a fraction of the price, potentially democratizing access to advanced AI capabilities and intensifying competition among major players.

Пов'язані матеріали

The True Cost of Strategy Being Forced to 'Hoard' U.S. Dollars

The article analyzes MicroStrategy's (now Strategy) shift towards holding substantial US dollar reserves (currently $4.65 billion) while selling Bitcoin, contrasting this with its core identity as a Bitcoin-focused company. It argues this move is a strategic necessity specific to Strategy, driven by its unique business model of issuing "digital credit" (preferred securities) with fixed USD dividend obligations. To secure favorable credit ratings (e.g., S&P's B- rating, which penalizes Bitcoin-heavy balance sheets) and sustain this credit issuance, Strategy must demonstrate significant dollar liquidity. However, this cash hoarding imposes a heavy "invisible tax" or opportunity cost: capital allocated to low-yielding cash reserves cannot be deployed into Bitcoin, artificially raising the required return threshold on its actual Bitcoin investments to cover fixed dividend costs. The author concludes that Strategy is an extreme case, and most other Bitcoin-related companies should not blindly emulate this strategy. For them, dollar reserves should be dictated strictly by operational needs and prudent buffers, not arbitrary targets. Holding excess cash typically means forfeiting potential Bitcoin returns for uncertain benefits, making it economically unsound unless a company shares Strategy's specific credit-dependent capital structure. The decision ultimately hinges on a firm's unique business model and genuine cash flow requirements.

marsbit1 хв тому

The True Cost of Strategy Being Forced to 'Hoard' U.S. Dollars

marsbit1 хв тому

New Fed Correspondent: Inflation Data 'Neither Hot Nor Cold' Offers Fed Temporary Respite but Future Path Remains Uncertain

Inflation data for July came in as expected, providing the Federal Reserve with some breathing room ahead of its September meeting, though the longer-term policy path remains uncertain. The core CPI rose 0.2% month-over-month and 2.5% year-over-year, in line with forecasts, easing immediate pressure for a rate hike. Following the report, market expectations for a September rate increase fell below 50%. Despite the temporary respite, deep divisions persist within the Fed. At least six of the twelve voting members have recently signaled openness to further tightening, with three having voted for a hike in July. The debate centers on whether current rates are sufficiently restrictive to bring inflation back to the 2% target or if persistent factors like tariffs, energy prices, and surging demand from AI infrastructure necessitate more action. San Francisco Fed President Mary Daly highlighted the growing complexity, outlining two potential scenarios: a baseline where inflationary pressures fade, allowing for rates to hold steady, and an alternative where shocks persist and inflation gains self-reinforcing momentum. She suggested that if the latter materializes, the policy response might need to be more aggressive than the typical 25-basis-point increments. With Fed Chair Wash adopting a less forward-leaning public stance, markets are closely monitoring data and other officials' comments for clues. The upcoming August CPI report, due just days before the September FOMC meeting, is now seen as a critical determinant for the final decision.

marsbit1 хв тому

New Fed Correspondent: Inflation Data 'Neither Hot Nor Cold' Offers Fed Temporary Respite but Future Path Remains Uncertain

marsbit1 хв тому

Market Trend (Aug 13): Nasdaq Leads Gain as CPI Meets Expectations; Coherent Beats Estimates Then Falls in After-Hours

U.S. stock markets were mixed on August 13th. While the S&P 500 and Nasdaq Composite rose, the Dow Jones Industrial Average was essentially flat. The catalyst was the July CPI report, which came in line with expectations across the board, keeping market pricing for a September Federal Reserve rate hike steady near 45%. The trading day's standout was a surge in AI-related sectors following the CPI release. AI cloud service providers Nebius and CoreWeave soared, while companies in the optical communications and memory chip spaces, including Coherent and SK Hynix, also posted significant gains. However, a stark reality check emerged after the closing bell. Despite reporting quarterly results that exceeded expectations, shares of optical communications leader Coherent and AI chipmaker Cerebras plunged in after-hours trading. Cerebras was hit by an unexpected drop in hardware revenue. Similarly, Cisco Systems shares initially rose on strong earnings and news of $4 billion in AI-related orders, only to reverse sharply lower later, highlighting increased investor scrutiny on valuations and growth sustainability in the crowded AI space. In other markets, gold extended its rally, oil prices rose, and Bitcoin rebounded from pre-CPI declines. Market focus now shifts to digesting the latest corporate earnings and upcoming U.S. retail sales data.

marsbit21 хв тому

Market Trend (Aug 13): Nasdaq Leads Gain as CPI Meets Expectations; Coherent Beats Estimates Then Falls in After-Hours

marsbit21 хв тому

Bitwise Closes 8 ETFs and Cuts Staff by 14%, but Continues to Launch New Products

Bloomberg reported on August 12 that cryptocurrency asset management firm Bitwise has laid off approximately 14% of its staff, reducing headcount from around 180 to about 155 employees. CEO Hunter Horsley stated the adjusted team remains the largest in the company's eight-year history and expressed confidence in long-term growth, though current personnel and product configurations are tightening. Prior to the layoffs, Bitwise disclosed a significant drop in client assets from over $15 billion in early February to $11 billion by April 1st, a decrease of at least $4 billion. The company did not specify the contributions of price changes, fund flows, or changes in reporting scope to this decline. In a related product restructuring, Bitwise decided to liquidate eight ETFs between April and June. This included the Web3 ETF, a BTC/ETH/Treasuries rotation strategy ETF, and six options income ETFs tied to assets like Coinbase and Ethereum. These moves reduce the number of products requiring ongoing operational support. Despite these cuts, Bitwise continues to launch new products. Recent offerings include an Avalanche ETP with staking in Europe, the Hyperliquid ETF, and the takeover of Superstate's $267+ million Crypto Carry Fund, entering the tokenized fund management space. The company now manages 70 investment products for over 5,500 advisory teams and works with more than 20 banks and broker-dealers. The simultaneous staff reduction and product portfolio shift—away from thematic Web3 and single-asset options strategies towards direct crypto exposure, staking, and tokenized funds—indicates a strategic realignment. The operational burden on the smaller remaining team will depend heavily on the complexity of the revised product lineup. Bitwise has not disclosed the specific reasons for the layoffs, the affected departments, severance details, or any one-time costs associated with the job cuts.

marsbit1 год тому

Bitwise Closes 8 ETFs and Cuts Staff by 14%, but Continues to Launch New Products

marsbit1 год тому

CoreWeave: The Inflection Point Has Arrived. Has the 'Hard-Working Underdog' Finally Turned Profitable?

CoreWeave, an AI cloud unicorn, released its Q2 2026 earnings on August 12, with shares rising about 15% post-announcement. While overall performance did not significantly exceed expectations, key positive trends emerged. Revenue reached approximately $2.58 billion, up 112% year-over-year, though slightly below the high end of prior guidance. A major highlight was the acceleration in computing power deployment, with Active Power increasing by a record 500 Mw to 1,500 Mw, far surpassing market forecasts and signaling faster future revenue growth. Capital expenditures also hit a new high of $9.4 billion. Importantly, profitability showed signs of inflection. "True" gross margin (after deducting cost of revenue and Tech & Infrastructure expenses) rose to 7.3%, up 3 percentage points from the previous quarter. As revenue scales, depreciation and operating expenses are being diluted. Adjusted operating profit margin improved significantly to 5% from 1% last quarter. The company's guidance points to continued acceleration, with Q3 revenue growth expected at 158% and margins continuing to climb. Management raised full-year 2026 guidance, projecting Q4 revenue growth of around 194% and an adjusted operating margin of approximately 14.6%, suggesting a rapid path toward its 25%-30% long-term target. Recent developments, including a 25% price increase for its services and the launch of higher-margin Managed Inference offerings, support the improving profitability narrative. While long-term competitive challenges remain for new cloud providers, CoreWeave's near-term trajectory of accelerating growth and expanding margins presents a high-risk, high-reward opportunity, especially amid renewed market optimism for cloud stocks.

marsbit1 год тому

CoreWeave: The Inflection Point Has Arrived. Has the 'Hard-Working Underdog' Finally Turned Profitable?

marsbit1 год тому

Торгівля

Спот

Популярні статті

Як купити ONE

Ласкаво просимо до HTX.com! Ми зробили покупку Harmony (ONE) простою та зручною. Дотримуйтесь нашої покрокової інструкції, щоб розпочати свою криптовалютну подорож.Крок 1: Створіть обліковий запис на HTXВикористовуйте свою електронну пошту або номер телефону, щоб зареєструвати обліковий запис на HTX безплатно. Пройдіть безпроблемну реєстрацію й отримайте доступ до всіх функцій.ЗареєструватисьКрок 2: Перейдіть до розділу Купити крипту і виберіть спосіб оплатиКредитна/дебетова картка: використовуйте вашу картку Visa або Mastercard, щоб миттєво купити Harmony (ONE).Баланс: використовуйте кошти з балансу вашого рахунку HTX для безперешкодної торгівлі.Треті особи: ми додали популярні способи оплати, такі як Google Pay та Apple Pay, щоб підвищити зручність.P2P: Торгуйте безпосередньо з іншими користувачами на HTX.Позабіржова торгівля (OTC): ми пропонуємо індивідуальні послуги та конкурентні обмінні курси для трейдерів.Крок 3: Зберігайте свої Harmony (ONE)Після придбання Harmony (ONE) збережіть його у своєму обліковому записі на HTX. Крім того, ви можете відправити його в інше місце за допомогою блокчейн-переказу або використовувати його для торгівлі іншими криптовалютами.Крок 4: Торгівля Harmony (ONE)Легко торгуйте Harmony (ONE) на спотовому ринку HTX. Просто увійдіть до свого облікового запису, виберіть торгову пару, укладайте угоди та спостерігайте за ними в режимі реального часу. Ми пропонуємо зручний досвід як для початківців, так і для досвідчених трейдерів.

535 переглядів усьогоОпубліковано 2024.12.12Оновлено 2026.06.02

Як купити ONE

Обговорення

Ласкаво просимо до спільноти HTX. Тут ви можете бути в курсі останніх подій розвитку платформи та отримати доступ до професійної ринкової інформації. Нижче представлені думки користувачів щодо ціни ONE (ONE).

活动图片