Today, this showdown is simply surreal!
On one side, Liang Wenfeng finally released the official version of DeepSeek V4 Pro in the early hours.
On the other, Musk dropped the next-generation flagship Grok 4.6, touting low cost and high performance.
Two giants, releasing at the same time, instantly filling the air with gunpowder.

They are targeting almost the same thing—
enabling models to continuously call tools, modify code, verify results, and finally deliver a truly usable product within long-running tasks.
DeepSeek V4 Pro Arrives
Musk's Grok 4.6 Takes the Stage
The most dazzling achievement of the official DeepSeek V4 Pro is securing absolute first place in two evaluations, directly countering Fable 5.
It scored 83.3 points in Cybersecurity Agent Test CyberGym, surpassing Fable 5's 83.1 and Opus 4.8's 78.3.
In Automation Task AutomationBench, it scored 31.8 points, pushing down Fable 5's 29.1 and Opus 4.8's 27.2.
The most "in-your-face" moment occurred in Terminal-Bench 2.1.
DeepSeek scored 87.9 points, exceeding Opus 4.8's 85.0 and trailing Fable 5's 88.0 by only 0.1.
In several other high-difficulty Agent tests, DeepSeek achieved a complete overtaking of Opus 4.8.
In the tool-assisted "Final Human Exam", it scored 60.0 points, surpassing Opus 4.8's 57.9 and continuing to close in on Fable 5's 63.0.
Even in the previously weakest area of software engineering agents, DeepSWE surged from 12.8 points in the preview version to 62.7 points, nearly 4.9 times the original.
This score not only exceeded Opus 4.8's 58.0 but also left only a 7.3-point gap to Fable 5's 70.0.

Simultaneously, Musk also put Fable 5 and GPT-5.6 Sol on the hot seat.
Grok 4.6 first caught up with GPT-5.6 in comprehensive capabilities, then achieved consecutive overtakes in coding tests.
It scored 61 points in General Intelligence Index, just 1 point behind Fable 5's 62.
On CursorBench, it scored 69.9%, surpassing GPT-5.6's 67.2% and only 0.6 percentage points behind Fable 5's 70.5%.
On FrontierCode, it scored 61.3%, also edging out GPT-5.6's 60.6%, continuing to closely trail Fable 5's 63.6%.
In knowledge work closer to real workplace deliverables, Grok 4.6 directly turned the tables, securing first place in all three evaluations.
In GDPval-AA v2, it scored 1753 Elo, surpassing Fable 5's 1741 and GPT-5.6's 1728.
On AA-Briefcase, it scored another 1577 Elo, beating Fable 5's 1574 and GPT-5.6's 1502.
In the specialized legal task Harvey LAB, it scored 15.8%, while Fable 5 only managed 11.3%, and GPT-5.6 a mere 2.5%.

What's even more impressive is that DeepSeek V4 Pro and Grok 4.6 not only match the performance of OpenAI and Anthropic's top-tier models but also jointly bring down the price of cutting-edge intelligence.
Per million output tokens, Grok 4.6 costs $6, GPT-5.6 Sol costs $30, Claude Opus 5 costs $25, Fable 5 costs $50.
And DeepSeek costs only $0.87—
approximately 1/7th of Grok 4.6, 1/35th of GPT-5.6 Sol, 1/29th of Claude Opus 5, and 1/57th of Fable 5!


World's First Hands-on Test
DeepSeek vs. Grok
Now, the first batch of real-world tests is out. DeepSeek V4 Pro and Grok 4.6 have finally moved from benchmark charts into the field of real tasks.
Round one, we directly put pressure on DeepSeek.
With just one prompt, it built a complete 3D interactive Earth from scratch in a browser.
Basic interactions like dragging, rotating, and zooming were all functional; global data flows, dynamic flight paths, and geographic markers were fully laid out on the Earth's surface.
Looking deeper, atmospheric scattering, cloud rendering, day-night lighting, and even the entire UI interface were all included.

Next, we pitted the two models directly against each other in the same arena.
AI blogger "Xiangyang Qiaomu" first ran three small tasks with DeepSeek V4 Pro, then gave two of the same prompts to Grok 4.6, directly comparing the final products.
The first task was to call 3 Skills to develop and deploy a website.
DeepSeek smoothly ran through the entire process; judging by page design and completeness, its overall performance was very stable.
The second task was to replicate 60 design styles and generate a complete set of Bento cards for centralized display.
In this round, DeepSeek made clear distinctions in fonts, color schemes, and layouts, with an overall visual effect that was quite impressive.
The final task was to generate a 3D brick-breaker game from zero.
The game was not only playable but also featured a 3D scene, background music, and dynamic sound effects, with operational feedback and playability fully present.



Subsequently, the same prompts were given to Grok 4.6.
In the 3D brick-breaker round, the two models were almost evenly matched.
The game generated by Grok also had good quality; whether in visual effects, scene completeness, or playability, it was on par with DeepSeek V4 Pro.
In the 60 Bento designs round, Grok 4.6 pulled back a point.
Some of the pages it generated were more mature in layout, color scheme, and visual hierarchy, resulting in a more aesthetically pleasing final product.


An even fiercer duel occurred with Flappy Bird.
Developer Jun Song gave the exact same prompt, asking DeepSeek V4 Pro and Grok 4.6 to respectively "handcraft" a game from scratch.
As a result, the two models took completely different paths.
DeepSeek V4 Pro consumed over 20,000 tokens at a cost of only $0.019; Grok 4.6 used about 5,000 tokens but cost $0.03.

However, judging by the final product, DeepSeek clearly won this round.
The game featured distant mountain ranges and layered clouds; the pipes had gradients and a sense of volume. When the character passed through a pipe, a floating "+1" animation even popped up.
From scene layering to operational feedback, the details were almost all maxed out; the completeness of the end-to-end build was noticeably higher than Grok 4.6's.
In another front-end test conducted by developer Hamza, DeepSeek V4 Pro once again bested Grok 4.6.
Whether in page completeness or final visual effects, the product delivered by V4 Pro was superior.

However, V4 Pro did not reign supreme in every test.
In a pelican comparison test, compared to the Flash version, V4 Pro's overall image completeness was higher, and the elephant's shape was more aesthetically pleasing.
The only problem was—the pelican's direction of movement was completely drawn backward.

Increasing the difficulty further, asking DeepSeek V4 Pro, GPT-5.6 Sol, and Claude Opus 5 to generate the same cherry blossom tree using Three.js, the gap became more apparent.
Whether in the details of the trunk and branches or in lighting, depth of field, and overall atmosphere, V4 Pro fell short of the other two top-tier models.



DeepSeek V4 Pro; GPT-5.6 Sol; Claude Opus 5
After several rounds of hands-on tests, it's clear that DeepSeek V4 Pro and Grok 4.6 are indeed on the same level, neck and neck.
Two Major AIs, Same Day
Caught Up with OpenAI and Anthropic
Regarding the simultaneous breakthroughs of these two models, expert Rick De Oliveira gave the evaluation: "Powerful, and affordable."
With the empowerment of DeepSeek V4 Pro and Grok 4.6, these two words that were almost impossible to appear together have, for the first time, truly come together today.
From cybersecurity and knowledge-based tasks to Agents autonomously solving problems, they have already shattered the ceiling of cutting-edge capabilities in multiple battlefields, all at a lower cost.

Looking back at today's surreal AI "clash," DeepSeek and Grok have jointly rewritten a rule of the game:
Top-tier intelligence is no longer out of reach.
In the past, developers often had to make difficult choices between "can't afford it" and "not powerful enough."
Today, DeepSeek overturned the table with a price tag mere fractions of SOTA, while Grok followed closely with extremely high cost-effectiveness.
This frontal assault of "high performance at low price" has already torn a rift in the old order.

And the war is not over.
Musk has already previewed that Grok 4.7 is coming soon, aiming to surpass all top AIs; counterattacks from OpenAI and Anthropic are also bound to follow.


In this era evolving by the "hour," intelligence is no longer a privilege of a few giants. The explosion of applications for everyone has only just begun.
References:
https://api-docs.deepseek.com/zh-cn/quick_start/pricing/
https://x.ai/news/grok-4-6
https://x.com/vista8/status/2087577905081823559
https://x.com/vista8/status/2087566666477744620
https://x.com/jun_song/status/2087602149979254914
This article is from the WeChat public account "Xin Zhi Yuan," author: ASI Apocalypse, editors: Moses Peach








