Breaking News: Musk Delivers the Most Powerful Grok 4.5, Slashes Price of Top-tier Opus Intelligence Drastically

marsbit2026-07-09 tarihinde yayınlandı2026-07-09 tarihinde güncellendi

Özet

**Elon Musk Launches Grok 4.5: A Cost-Effective, High-Performance AI Rival** SpaceXAI, in collaboration with Cursor, has released Grok 4.5, its new flagship AI model designed specifically for coding and agentic tasks. Trained on tens of thousands of NVIDIA GB300 GPUs using massive, high-quality data filtered from trillions of Cursor developer interactions, the model emphasizes "per-token intelligence." In benchmark performance, Grok 4.5 is highly competitive. It scores 64.7% on SWE Bench Pro (surpassing GPT-5.5's 58.6% and Opus 4.7's 64.3%), 83.3% on Terminal Bench 2.1 (nearly matching GPT-5.5), and 62.0% on DeepSWE 1.0 (beating Opus 4.8). Overall, it ranks fourth in AAAI official tests and first in the Harvey legal agent benchmark. The model's key advantage is its combination of speed, efficiency, and low cost. It generates responses at 80 tokens per second and, crucially, uses far fewer tokens to complete tasks—4.2 times fewer than Opus 4.8 on SWE Bench Pro. It is priced at $2 per million input tokens and $6 per million output tokens, significantly undercutting competitors. Musk stated it is "roughly equivalent to Opus 4.7, but much faster." Early user tests show Grok 4.5 can generate functional code for applications like 3D solar system simulators and basic games from simple prompts, though some note it still lags behind top models in certain creative tasks. Musk has hinted at a major update next month, leveraging real-world engineering data from his companies, with an...

Grok 4.5, finally delivered!

Just now, Musk's SpaceXAI impressively released its own strongest flagship model to date, Grok 4.5.

This time, SpaceXAI and Cursor have joined forces powerfully.

On tens of thousands of GB300 beasts, they have fiercely "forged" this performance monster, born for coding and agents.

The report card is quite dazzling—

  • SWE Bench Pro: Sweeps a whopping 64.7%, directly challenging and surpassing Opus 4.7's 64.3%;
  • Terminal Bench 2.1: Races all the way up to 83.3%;
  • DeepSWE 1.0: Firmly stands at 62.0%, decisively crushing Opus 4.8.

Musk stated plainly, "Grok 4.5 is roughly equivalent to Opus 4.7, but much faster."

It's the combination of capability, speed, and cost that forms its competitiveness. In other words, the Tokens it uses for work are only a fraction of others'.

Grok 4.5 is priced at $2 per million tokens for input, $6 per million tokens for output.

Compared to Opus 4.8, Token consumption is dramatically reduced by 4.2 times.

Tens of thousands of GB300s forged an "Opus-level" model

Grok 4.5 is the first ace card after SpaceXAI went public and the first report card from their collaboration with Cursor.

So, how was it trained?

The answer: tens of thousands of NVIDIA GB300 GPUs, in one super large-scale training session. But stacking compute power was just the entry ticket.

Musk previews: Grok 4.5 context will be upgraded to 1 million next week

Where SpaceXAI truly put in exhaustive effort is in the data.

They subjected massive corpora to devilish filtering, deduplication, and quality scoring, ensuring every bit fed into the model was professional content with high information density.

Then, they focused RL efforts on an indicator rarely mentioned—"per-token intelligence".

Cursor's involvement is the most crucial link in this flywheel.

Grok 4.5's base is V9 (1.5T), and officially, it was trained on trillions of Cursor data points.

This data records how real developers interact with codebases, tools, and agents.

This means the model learns not just "what code looks like," but "how humans and AI code together."

And its training stack is designed for high asynchronicity—

Agents can run for hours continuously, the model keeps learning while working, and training on tens of thousands of GPUs never stops.

The result is that it not only solves problems but can also withstand complex, multi-step engineering tasks over hours.

Matches GPT-5.5, approaches Opus 4.8

The hardcore performance of Grok 4.5, trained with this approach, can absolutely withstand scrutiny.

Its performance on several core engineering benchmarks can be described as "steady"—not the strongest, but solidly within the top tier.

On DeepSWE 1.0, it scores 62.0%, surpassing Opus 4.8 (55.75%) and closely trailing GPT-5.5 (64.31%);

On Terminal Bench 2.1, it surges to 83.3%, only 0.1 behind GPT-5.5 (83.4%)!;

On the more hardcore SWE Bench Pro, it overtakes GPT-5.5 (58.6%) with a 64.7% solve rate, approaching Opus 4.8 (69.2%).

In the official AAAI test, Grok 4.5 ranks fourth, behind only Fable 5, GPT-5.5, and Opus 4.8.

In the Harvey legal agent benchmark test, it ranks first.

It must be said, this report card is quite impressive.

But it must be acknowledged that the current undisputed king is still Claude Fable.

In summary, Grok 4.5 is roughly on par with GPT-5.5, closely follows Opus 4.8, but there is still a distance to the true ceiling.

Musk himself said quite realistically: "Our internal evaluation is that Grok 4.5 is roughly equivalent to Opus 4.7, but much faster."

The real killer feature: Fast and cheap

Grok 4.5's real knockout punch lies in three words: speed, efficiency, and price.

Its inference speed reaches up to 80 TPS (Tokens per second). The official statement is "faster than flash-type models."

With just one sentence, it wrote a 3D solar system simulator using Three.js:

Supports time acceleration, realistic orbital motion of the eight planets, even the HUD panel is quite well-crafted.

Regarding efficiency, on SWE Bench Pro tasks, Grok 4.5 on average outputs only 15,954 Tokens to solve the problem.

Opus 4.8, doing the same job, spits out an average of 67,020 Tokens.

4.2 times—Grok 4.5 solves the same engineering problem using less than a quarter of its competitor's Tokens.

On price, it's $2 per million tokens for input, $6 per million tokens for output.

There's also an even faster premium version priced at $4 input / $18 output.

Compared to a host of top-tier models starting at over ten dollars, this price almost beats the cost down to the bone.

Put these three numbers together, and the conclusion is one sentence:

When it comes to "how much intelligence you can buy per unit of time and cost," Grok 4.5 is currently the "king of cost-effectiveness."

Global hands-on tests: Grok 4.5's real performance

Furthermore, from the hands-on feedback of netizens, we can glimpse a corner of Grok 4.5's true capabilities.

One sentence, directly outputs "Minecraft."

In a single HTML file, Grok 4.5 can easily handle a complete high-end SaaS page.

Grok 4.5 showcases its skills, completing a full set of 2D+3D design schemes within the app in less than a minute.

AI game master Danny Limanseta used Grok 4.5 to generate a game with various complete functions.

However, some developers stated that Grok 4.5 is completely not on the same level as Opus 4.7, and the generated lava lamp test performed poorly.

Not the strongest, saving energy to overturn the table next month

Today, Musk planted another bombshell:

Grok understands engineering deeply.

Next month's version will be another step-change improvement, because we are closing the loop on solving real engineering problems inside Tesla, SpaceX, Neuralink, and Boring Company.

Next month, another leap. And reportedly, a larger 2 trillion parameter version is already on the way.

Benchmarks are fireworks for outsiders; efficiency and cost are the real weapons to drag opponents into a war of attrition.

When model intelligence starts being billed like electricity per kilowatt-hour, the winning move is who can make intelligence fast, cheap, and omnipresent.

This time, Musk didn't play the strongest card, but he flipped the table.

References:

https://x.ai/news/grok-4-5

https://cursor.com/blog/grok-4-5

This article is from WeChat official account "New Zhiyuan," author: ASI Apocalypse, editor: Taozi

İlgili Sorular

QWhat is the main advantage of Grok 4.5 according to the article compared to models like Opus 4.8?

AIts main advantages are significantly higher speed, better efficiency (using far fewer tokens for the same tasks), and a much lower cost. For example, it uses about 4.2 times fewer tokens than Opus 4.8 on SWE Bench Pro tasks and is priced at $2/million tokens for input and $6/million for output.

QWhich company did SpaceX AI collaborate with to develop Grok 4.5, and what unique data did this partnership provide?

ASpaceX AI collaborated with Cursor. The training incorporated trillions of tokens of Cursor data, which records how real developers interact with codebases, tools, and AI agents. This taught the model not just 'what code looks like' but 'how humans and AI code together.'

QWhat are some of the key benchmark scores achieved by Grok 4.5 mentioned in the article?

AKey benchmark scores include: SWE Bench Pro: 64.7% (surpassing Opus 4.7's 64.3%), Terminal Bench 2.1: 83.3%, and DeepSWE 1.0: 62.0% (surpassing Opus 4.8's 55.75%). It also ranked 4th in AAAI official tests and 1st in the Harvey legal agent benchmark.

QWhat did Elon Musk hint at regarding the next version of Grok?

AElon Musk hinted that next month's version will be another 'step-change' improvement. This will be achieved by closing the loop on solving real engineering problems within his companies like Tesla, SpaceX, Neuralink, and Boring Company. A larger 2-trillion parameter version is also rumored to be in development.

QWhat is the core training focus for Grok 4.5's reinforcement learning, as stated in the article?

AThe core focus of its reinforcement learning (RL) was on optimizing 'per-token intelligence,' an indicator rarely emphasized by others. This aims to maximize the useful intelligence delivered by each token the model processes.

İlgili Okumalar

The 'Saving U.S. Treasuries' Baton Pass: Bessent Fumbled Last Week, This Week It's Wash's Turn

"Rescuing US Treasuries" Relay: After Bessent's Miss, All Eyes Are on Walsh Last week, US Treasury Secretary Bessent's announcement to at least double long-term Treasury buybacks failed to sustainably lower yields, which quickly rebounded. The market response saw a drop in the dollar alongside surges in gold and Bitcoin, interpreted as a "pressure release valve" for anxiety. The focus now shifts to Fed Chairman Walsh's upcoming Jackson Hole speech. Markets are highly sensitive to his message, seeking clarity on the Fed's policy response to stubborn inflation and worsening fiscal conditions. Analysts warn that a lack of new guidance could disappoint markets and worsen the sell-off in long-dated bonds. Analysts question the scale of Bessent's operations, noting they are too small relative to the overall debt market and do not constitute quantitative easing. A key issue is the Fed's massive holdings of long-term bonds, which distorts the market. With the Fed holding low-yielding short-term bonds that are losing money relative to its policy rate, discussion is growing around a potential Fed-led "Operation Twist." This would involve selling short-term bonds to buy long-term ones, aiming to lower long-end yields without expanding the balance sheet. The upcoming PCE inflation data will set the stage for Walsh's speech. However, the window for action is narrowing amid political pressures. A critical threshold is the 30-year yield at 5%; holding above it could increase stress on the dollar and leveraged sectors. Overall, the article suggests that without coordinated Fed action to anchor inflation expectations, Treasury interventions may ultimately fail, with investors increasingly looking to assets like gold as hedges.

marsbit22 dk önce

The 'Saving U.S. Treasuries' Baton Pass: Bessent Fumbled Last Week, This Week It's Wash's Turn

marsbit22 dk önce

Hyperliquid's Compliance Journey: From Permissionless to Permissioned via HIP-3

Hyperliquid’s Compliance Path: From Permissionless to Permissioned HIP-3 Hyperliquid currently blocks U.S. access because its permissionless, on-chain infrastructure conflicts with U.S. market structure laws, which restrict futures trading to registered exchanges, clearinghouses, and brokers. Through its Hyperliquid Policy Center (HPC), the project is advocating for regulatory modernization, proposing that regulated entities be allowed to build products on HyperCore (its exchange and clearing layer) while fulfilling their compliance obligations. The platform’s modular stack separates roles like a traditional exchange (DCM), clearinghouse (DCO), and broker (FCM), but reconstructs them on-chain with code. This enables permissionless access, self-custody, and 24/7 global trading, but clashes with U.S. rules requiring KYC, specific margin models, and custodial arrangements. To resolve this, HPC is engaging with U.S. regulators (CFTC, SEC) to seek clarity that deploying on-chain software does not itself trigger licensing, and to establish exemptions allowing non-custodial wallets to route users to regulated derivatives. Recent political signals suggest openness to this approach. On the technical side, Hyperliquid Labs has introduced permissioned HIP-3 deployers on testnet. These allow regulated entities to launch markets, perform KYC, and whitelist compliant users. While these create separate order books, whitelisted market makers can bridge liquidity between them, ensuring deep, shared liquidity across the same L1. Features like payload-based “PA” permissions enable DEX-level account controls (e.g., reduce-only orders), mirroring traditional broker authorities. The strategy is not to open the native, permissionless front-end to U.S. users, but to position Hyperliquid as neutral infrastructure that U.S. regulated firms can use while meeting their legal duties. This paves a compliant path for U.S. investor access while preserving the protocol’s core, permissionless nature.

marsbit46 dk önce

Hyperliquid's Compliance Journey: From Permissionless to Permissioned via HIP-3

marsbit46 dk önce

Two Funding Rounds in Three Months: The Chinese Version of Palantir is on Fire

Investment Community AI has learned that Beijing Zhongshu Ruizhi Technology Co., Ltd., a domestic industrial-grade causal intelligence and high-reliability decision-making AI company, has recently completed a strategic financing round worth hundreds of millions of RMB. This round saw participation from China Internet Investment Fund, Suzhou Chuangtou National Social Security Fund, Financial Street Capital, ICBC Capital, Kunlun Capital, among others, with existing shareholders also increasing their investment. This follows a Series B funding round in the hundreds of millions completed just three months prior. The rapid succession of two major funding rounds signifies strong market recognition of the company's underlying original technology and scaled commercial implementation. Often referred to as the "Chinese version of Palantir," Zhongshu Ruizhi is entering a new phase of accelerated technological iteration, widespread scenario replication, and scaled performance release, mirroring the explosive growth of China's AI market. Founded in April 2020 by Dr. Han Han, a Tsinghua University Ph.D. and former core drafter of national AI policies, the company is mission-driven to "move AI from the digital world to the physical world." It focuses on the high-reliability, strong-decision industrial AI track and enterprise-grade AI Agent full-stack infrastructure. The team tackles the challenge of applying AI to China's vast and complex industrial and energy systems by developing a new intelligent operating system from scratch. Its core technological breakthrough lies in three proprietary底层 technologies: meta-causal cognitive theory, causal models, and a dynamic ontology engine. These address critical pain points of generative large models in industrial settings—such as AI hallucinations, insufficient reasoning, lack of temporal logic, unverifiable decisions, and multi-source rule conflicts—thereby providing trustworthy, explainable, and executable智能决策 capabilities. Commercially, Zhongshu Ruizhi has achieved scaled deployment, serving over 50 central state-owned enterprises and industrial groups in sectors like power, petroleum, and aerospace, with implementations in more than 800 highly complex production scenarios. The company reported doubled revenue in 2025, demonstrating strong self-sufficiency and a viable business model—a rarity among new-generation AI firms. The latest funds will be allocated towards advancing foundational theoretical research, replicating successful application models to expand market presence (including overseas), and attracting top-tier talent. Lead investor China Internet Investment Fund highlighted that in the current shift from general AI capability contests to deep industrial empowerment, industrial-grade causal intelligence is crucial for building China's modern digital foundation and fostering new quality productive forces. They expressed support for the company's efforts to define decision-making paradigms and trustworthy standards for industrial intelligence, aiming to secure a rule-making voice in the global physical AI arena.

marsbit57 dk önce

Two Funding Rounds in Three Months: The Chinese Version of Palantir is on Fire

marsbit57 dk önce

The Biggest Political Economy Question in the AI Era: As Robots Become More Capable, How Do Humans Share the Value?

In the AI era, the most pressing political economy question is: as machines become increasingly capable, how can humanity share in the value they create? An article originally critiquing China's tech focus has sparked a deeper debate on this global challenge. Historically, industrial progress improved efficiency but still relied on human labor for wealth creation and distribution. AI is fundamentally different—it is now replacing cognitive and knowledge work. As AI and robots take over more tasks, economic growth may continue while direct human participation in value creation shrinks, creating a core tension between productivity gains and widespread income generation. The issue is not unique to China. While leading tech companies amass enormous wealth, labor's share of income is declining globally. The core problem is a broken link: technological innovation and corporate profits are not translating into sufficient consumer income and demand. Three potential paths forward are outlined: a traditional capitalist model where profits primarily go to capital owners; a state-capitalist approach with public investment in AI; and more innovative models like digital sovereign wealth funds, universal shareholding, or AI-era basic income schemes to directly distribute AI-generated value. The future competitive advantage may lie not just in technological supremacy, but in which society can build a new, inclusive distribution system for the intelligent economy. The ultimate challenge is ensuring that as AI creates value, humans have a means to obtain income and share in the resulting widespread social benefits.

marsbit1 saat önce

The Biggest Political Economy Question in the AI Era: As Robots Become More Capable, How Do Humans Share the Value?

marsbit1 saat önce

Generating Profits for Seven Consecutive Quarters, Emerging Markets Carry Trade Outperforms Everything

For the seventh consecutive quarter, dollar-funded emerging market carry trades have delivered positive returns, marking the longest winning streak since 2008. According to Bloomberg's index, this strategy has gained approximately 22% since late 2024, outperforming U.S. Treasuries, emerging market sovereign, and corporate dollar debt. The core of the trade involves borrowing low-interest currencies like the U.S. dollar, euro, or yen to invest in high-yielding emerging market assets, such as Turkish lira bonds offering over 40% returns. Returns were amplified by favorable currency moves, with the dollar weakening against most emerging market currencies and other traditional funding currencies. For instance, the trade gained 48% on the Colombian peso in the past year. A key test came in August 2024 with a historic joint U.S.-Japan currency intervention, which caused only a modest 1% dip in the carry trade risk premium as investors shifted funding from the yen to the euro and Swiss franc. Looking ahead, the primary risk is the timing of Federal Reserve policy changes. While persistent inflation allows the Fed to hold rates, a rapid rise in long-term U.S. yields could threaten the trade. Another concern is crowding, as massive inflows increase vulnerability to a sudden reversal. High interest rates in regions like Latin America and Eastern Europe, supported by external factors like Middle East tensions and energy prices, continue to sustain the opportunity. Major investors remain engaged, favoring currencies like the Mexican peso, South African rand, and Turkish lira.

marsbit1 saat önce

Generating Profits for Seven Consecutive Quarters, Emerging Markets Carry Trade Outperforms Everything

marsbit1 saat önce

İşlemler

Spot
活动图片