Replicating the "DeepSeek Moment"? Wall Street Unanimously Says: Kimi K3 Instead Strengthens Computing Power Demand

链捕手Published on 2026-07-21Last updated on 2026-07-21

Abstract

Title: Wall Street Sees Kimi K3 as a Catalyst for Compute Demand, Not a "DeepSeek Moment 2.0" Summary: Following the release of Moonshot AI's powerful open-source model Kimi K3, initial market reaction mirrored the "DeepSeek moment" that sparked a sell-off in compute stocks earlier in 2025, fearing reduced demand for AI infrastructure. However, major Wall Street banks including UBS, Nomura, BofA, and Citi argue the opposite: K3 will accelerate, not weaken, demand for compute, memory, storage, and networking. Their analysis centers on K3's specifications—2.8 trillion parameters, 1M token context, and MoE architecture—which represent a "scale" story rather than a pure "efficiency" one like DeepSeek R1. These features increase pressure on inference, memory (especially KV cache), and storage. Analysts invoke Jevons Paradox: as high-quality models become more affordable (K3 is cheaper than top closed models but not the cheapest), usage and token volumes expand, ultimately increasing total compute consumption. The reports highlight that competition will force leading US AI labs (OpenAI, Anthropic, Google) to invest more in training and iteration to maintain their edge. Furthermore, the rise of capable open-source models like K3 is expanding the global AI developer ecosystem, with Chinese models now accounting for over 45% of developer traffic. Key beneficiaries identified across the AI infrastructure chain include memory/storage players (e.g., Micron, Samsung), compute leaders ...

Author: Long Yue

The market panicked by treating Kimi K3 as "DeepSeek Moment 2.0," but this time, Wall Street's judgment is completely different.

Late night on July 16th, Moonshot AI released Kimi K3 in Shanghai. This open-source model with 2.8 trillion parameters scored 57 on the Artificial Analysis Intelligence Index, ranking third to fourth globally, on par with Anthropic's Claude Opus 4.8 and OpenAI's GPT-5.5. More crucially, on the Frontend Code Arena programming benchmark created by UC Berkeley, K3 topped the chart with 1679 points, surpassing Claude Fable 5 and GPT-5.6 Sol, becoming the first open-source model to outperform all leading closed-source overseas models on an authoritative programming benchmark.

On July 17th, the US semiconductor sector saw a notable decline. The market's knee-jerk reaction was understandable – at the beginning of 2025, the release of DeepSeek R1 triggered a sharp drop in computing power-related stocks, based on the logic: If Chinese models are getting stronger and stronger, do US AI companies still need to invest so much money in computing power? If Chinese models can approach state-of-the-art capabilities at lower costs, will the demand for NVIDIA GPUs, HBM, servers, and networking equipment be reassessed?

However, according to trading desk reports, the latest research notes from banks including UBS, Nomura, Bank of America Merrill Lynch, and Citi argue: Kimi K3 is not a terminator of computing power demand, but an accelerator.

Kimi K3 and DeepSeek R1 are not the same kind of shock. R1 mainly made the market see "efficiency"; K3 highlights "scale." With 2.8 trillion parameters, a 1M token context window, always-on reasoning, native multimodal capabilities, and an MoE architecture, these features do not tell a capital-light story. They will collectively increase pressure on inference, memory, networking, and storage.

How Powerful is Kimi K3?

Kimi K3 was released by Moonshot AI on July 16, 2026. The full model weights are planned to be released on July 27th. It is an open-weight large model with 2.8 trillion parameters, referred to by multiple institutions as the largest open-weight LLM currently available.

Its core configuration includes three points:

First, a 1M token context window. The model can handle longer texts, larger codebases, and more complex enterprise documents and research tasks.

Second, always-on reasoning. Not just simple Q&A, but oriented towards long-chain reasoning and Agent tasks.

Third, native vision capabilities. K3 doesn't just process text; it also handles multimodal tasks like video, images, game development, frontend design, CAD, etc.

Architecturally, K3 uses Kimi Delta Attention, Attention Residuals, and Stable LatentMoE. The MoE part activates 16 experts per token out of 896 experts. Moonshot AI stated that compared to Kimi K2, the overall scaling efficiency has improved by about 2.5x.

This explains why K3 is not a "cheaper K2." Price comparisons compiled by Nomura show that K3's input price is $3 per million tokens, cached hit input is $0.30, and output is $15 per million tokens; according to Artificial Analysis, K3's cost per task is about $0.94. This price is lower than Claude Fable 5's ~$2.75 and Claude Opus 4.8's ~$1.80, close to GPT-5.6 Sol's $1.04, but significantly higher than GLM-5.2's $0.32—$0.47, and far higher than DeepSeek V4 Pro's $0.04.

Thus, K3's positioning is not the lowest price, but getting close to frontier model capabilities at a lower price.

Four Investment Banks Intensively Set the Tone: This is Not a Weakening of Demand

Regarding market concerns about a "DeepSeek Moment" impact, Duan Bing, an analyst at Nomura Securities' Asia-Pacific tech team, wrote in a report: "We believe competition and innovation in the global large model market will not stop. As we approach Artificial General Intelligence (AGI), the application of generative AI in consumer and enterprise sectors will continue to expand. Frontier AI labs and hyperscale cloud platform companies are likely to continue investing in this stage to maintain their competitive positions – as scaling laws continue, we interpret this competition as beneficial for the AI infrastructure value chain."

In his report on July 19th, Citi semiconductor analyst Peter Lee directly titled it "Another Jevons Paradox." What is the Jevons Paradox? Simply put: increased efficiency of coal steam engines led to greater total coal consumption, because more people could afford it and it was applied to more scenarios. It's the same with AI models – when high-quality models become cheaper, developers and enterprises deploy more applications, process more tokens, ultimately leading to increased computing power consumption.

Peter Lee believes that even if K3 is widely adopted, demand for general-purpose memory like server DDR5 and eSSD will still increase. The reason is that K3's inference efficiency is comparable to other frontier models, but its KV cache footprint grows with the context length, increasing pressure on memory rather than reducing it.

Bank of America Securities semiconductor analyst Vivek Arya was more direct in his report on July 17th. He believes that the response from US frontier AI labs is "not less computing power, but more." If Chinese open-source models continue to approach their capabilities, OpenAI, Anthropic, and Google must maintain differentiation through larger-scale training, heavier inference, and faster iteration. Arya also mentioned an easily overlooked background: media reports suggest Google's Gemini 3.5 Pro has been delayed by months from its plan, with programming performance not meeting internal targets, making it "increasingly difficult to defend" the frontier leadership position.

In a report dated July 20th, the team of UBS analyst Timo Arcuri pointed out that parallels between K3 and DeepSeek R1 indeed exist, but K3 is more about scale – it is the world's largest open-source model, with 2.8 trillion parameters and a 1 million token context window. The analysts emphasized that open-source models are generally more memory-intensive than closed-source frontier models because their context windows are longer, and KV cache demand, even after quantization, continues to grow in absolute terms, making deployment of open-source models more dependent on HBM and storage.

Storage: The most direct beneficiary sector. UBS calculations show that the storage and memory sectors are expected to generate cumulative free cash flow (FCF) reaching about 30% of market capitalization by 2028 – the highest among all subsectors, with Micron (MU) alone reaching 47%. Both Citi and Nomura maintain Buy ratings on Samsung Electronics, citing the global memory market being in a state of extremely tight supply. Citi analyst Peter Lee pointed out that Kimi K3's inference-side memory demand is no less than other frontier models, and the expansion of KV cache volume will directly drive demand for server DDR5 and enterprise SSDs (eSSD). He specifically noted: Large-scale deployment of Kimi K3 requires "super-node" cluster configurations with over 64 GPUs.

Computing Infrastructure: TSMC and NVIDIA are the primary beneficiaries. Whether it's scaling laws still effective on the training side or token demand growth on the inference side, both ultimately point to increased demand for advanced node chips. Nomura reiterated Buy ratings on TSMC, ASE, MediaTek, etc. NVIDIA has publicly stated that modern MoE model inference delivers up to 25x performance-per-watt improvement on the GB300 NVL72 compared to the previous Hopper architecture. Models like K3 naturally benefit from NVIDIA's latest hardware.

Networking: Super-node trend creates structural opportunities. Kimi K3 requires super-node clusters. Domestic Chinese computing power, constrained by high-end chip export controls, relies even more on super-node architectures to compensate for single-card performance gaps, driving demand for network layer suppliers like optical modules and optical chips. Nomura is bullish on Inphi and Suzhou Inphi.

Cloud Platforms: Benefiting from ecosystem agglomeration effects. Cloud platforms hosting multiple frontier open-source models have stronger bargaining power, not relying on a single closed-source model supplier. Nomura is bullish on Alibaba (BABA) as a core player in China's AI cloud ecosystem, and data center operators like GDS and 21Vianet (VNET).

This might be the most easily underestimated data point in the entire narrative.

According to statistics from the open API gateway OpenRouter, the proportion of token usage by Chinese AI models in global developer traffic has exceeded 45%, up from less than 2% a year ago. Data from Bank of America Merrill Lynch confirms the acceleration of overall AI penetration: about 55% of US companies have now subscribed to AI models, platforms, or tools, with enterprise adoption rates reaching 42% for Anthropic and 40% for OpenAI. Top AI consumers (the top 1% of enterprise users) spend $4,833 per employee per month on AI.

The market is diverging. On one hand are Chinese open-source models represented by DeepSeek and Kimi K3, covering the economy and mid-to-high-end value-for-money markets; on the other hand, top US frontier models are focusing on more complex workloads (like scientific computing) to maintain technology and pricing premiums. Nomura's judgment is: leading large model players on both the Chinese and US sides will benefit – provided they can continuously stay at the forefront of the technology curve.

After K3's release, the market's first reaction was still to compare it with DeepSeek R1. This comparison is useful but shouldn't stop at the level of "is AI hardware going to be sold off again?"

DeepSeek made the market re-evaluate training efficiency. K3 shows the market another thing: open-source models can also push scale, long context, Agent, and multimodality into frontier territory.

This will force US frontier labs to keep investing, and also allow Chinese models to continue expanding in the global developer ecosystem. Closed-source top models retain technology and pricing premiums, open-source models cover more price points and deployment scenarios, cloud vendors provide model distribution and enterprise adoption, and the hardware chain bears training and inference pressure.

In the short term, trading may fluctuate due to "DeepSeek memories." In the medium term, as long as token usage continues to grow, and long context and Agent continue to spread, computing power, HBM, storage, networking, and IDC remain unavoidable cost items.

This is also why multiple institutions reached similar conclusions after K3: stronger open-source models are not the endpoint for AI infrastructure demand; instead, they could be the entry point for the next wave of demand expansion.

However, Bank of America Merrill Lynch also clearly left a tail risk: "If the speed of efficiency gains outpaces workload growth, we might see a certain pullback in infrastructure construction." In other words, if models become much cheaper, but usage does not expand significantly in tandem, the growth logic for computing power demand would be discounted.

Related Questions

QWhat is the main argument presented by Wall Street analysts regarding the impact of Kimi K3 on AI infrastructure demand?

AWall Street analysts from firms like UBS, Nomura, Bank of America Merrill Lynch, and Citi argue that Kimi K3 is not a threat to AI infrastructure demand but rather an accelerator. They believe its advanced features (large parameters, long context, MoE architecture) will increase pressure on computing, memory, network, and storage, driving more demand for hardware like GPUs, HBM, and servers, similar to the 'Jevons Paradox' where efficiency gains lead to higher overall consumption.

QHow does Kimi K3 differ from DeepSeek R1 in terms of its market impact and technical characteristics?

ADeepSeek R1 primarily impacted the market by demonstrating 'efficiency'—showing that capable models could be built cost-effectively, leading to concerns about reduced spending on compute. Kimi K3, however, emphasizes 'scale.' With 2.8 trillion parameters, a 1M token context window, native multimodal capabilities, and an MoE architecture, K3 represents a large, complex model that is not a 'light asset' story. It increases demands on inference, memory, and storage, thus potentially boosting infrastructure demand rather than reducing it.

QAccording to the analysts, which specific hardware sectors are expected to benefit from the deployment of models like Kimi K3?

AAnalysts identify several hardware sectors as beneficiaries: 1) **Storage & Memory**: Due to increased KV cache needs from long contexts, boosting demand for server DDR5 and enterprise SSDs. Companies like Micron (MU) and Samsung are highlighted. 2) **Compute Infrastructure**: Companies like TSMC and NVIDIA benefit as larger models and growing token usage drive demand for advanced chips. 3) **Networking**: The need for 'super-node' clusters to run K3 boosts demand for optical modules and components (e.g., Zhongji Innolight). 4) **Cloud Platforms & Data Centers**: Platforms hosting multiple models (e.g., Alibaba Cloud) and data center operators (e.g., GDS, VNET) gain from ecosystem aggregation.

QWhat is the 'Jevons Paradox' mentioned in the context of AI models like Kimi K3?

AThe 'Jevons Paradox,' referenced by Citi analyst Peter Lee, describes a phenomenon where increased efficiency in using a resource (like coal in steam engines) leads to an increase in the total consumption of that resource, because lower costs make it accessible for more uses and scenarios. Applied to AI, as models like Kimi K3 become more capable and cost-effective, developers and enterprises will deploy more applications and process more tokens, ultimately leading to higher overall computing power consumption rather than a reduction.

QWhat potential tail risk do analysts see for the AI infrastructure growth narrative despite the demand boost from models like K3?

ABank of America Merrill Lynch notes a tail risk: If the rate of efficiency gains in model development outpaces the growth in actual AI workloads and usage, there could be a potential slowdown or 'pullback' in infrastructure investment. In other words, if models become significantly cheaper to run but the expansion in their application volume does not keep up, the fundamental logic for continuous, massive growth in computing demand could be undermined.

Related Reads

Global Stock Market's Storm Center: South Korea's Stock Market De-leveraging Is Largely Complete

Storm's Eye: South Korean Market De-leveraging Nears Completion The recent sharp correction in South Korean equities, with the KOSPI index dropping 32% from its June high, has been a key trigger for global tech stock volatility. The core driver was not a fundamental shift but a forced de-leveraging process within the market's unique structure, which is now largely complete. Two main leverage channels amplified the sell-off: 1. **Leveraged ETFs:** Their size, proportionally four times larger than in the U.S., peaked near $50 billion. Their mandatory daily rebalancing mechanism created a vicious cycle of "price drop → forced selling → further drop." Approximately 75% of this excess has been unwound, shrinking to $26 billion, with regulatory curbs now blocking new inflows. 2. **Hedge Fund Leverage:** Using swaps to magnify exposure, hedge funds saw their net long positioning fall by over 50% from peak levels. The most intense phase of this institutional de-leveraging is over. In contrast, **retail margin debt** poses minimal systemic risk. At 0.5% of market cap, it is far lower than in the U.S. or China, lacks automatic triggers, and is concentrated in smaller stocks. The conclusion: the high-leverage structures most prone to "chain-reaction selling" have been substantially cleared. The market is transitioning from a liquidity-driven crash to one priced more on fundamentals. The article argues that the AI trend—centered on Korean memory chips—remains intact. This episode represents a painful but necessary clearing of crowded trades, not the end of the AI revolution. For investors, the key question is conviction in the long-term AI direction; if the trend is real, current volatility is a cost of entry, not a terminal risk.

链捕手19m ago

Global Stock Market's Storm Center: South Korea's Stock Market De-leveraging Is Largely Complete

链捕手19m ago

The Eternal Fragments of Money: Third-Party Payment Has No First Principle

"The Enduring Fragments of Money: Third-Party Payments Lack a First Principle" Stripe is reportedly attempting to acquire PayPal, marking a significant shift reminiscent of PayPal's merger with the original X.com 30 years ago. The article analyzes Stripe's strategic challenges and the broader payments industry landscape. Despite its initial success with a developer-friendly API model, Stripe missed its optimal IPO window during the pandemic and has since seen its valuation decline. Its attempts to expand through acquisitions and new ventures, particularly in stablecoins (like its OUSD project) and Agent-focused payments (ACP/MPP protocols), have faced headwinds. The author argues that the payment industry remains highly fragmented and is ultimately an adjunct to the traditional banking system. This structure limits the potential for any single player, including Stripe, to achieve complete dominance. While stablecoins and the future rise of autonomous Agent economies present potential growth avenues, they are not yet mainstream and still require integration with the existing financial system. For now, Agent-based transactions are largely used for speculative "volume boosting" rather than substantive business applications. Stripe's current move to acquire PayPal is seen as an attempt to bolster its weak consumer-facing (C-side) business after its stablecoin-focused strategies faltered. Meanwhile, PayPal is described as structurally outdated, unable to revive itself through new products like Venmo or PYUSD. The future of payments may lie not in payments themselves but in value-added services like more efficient settlement networks. The author suggests that companies like Stripe and Circle, which are building their own blockchains (Tempo, Arc) and stablecoins, are positioning themselves to eventually profit from high-efficiency settlement systems. These new networks could potentially bypass some traditional banking layers. In conclusion, the article posits that third-party payment is a perpetually fragmented battlefield where scale alone cannot ensure victory. Players must find new models, focusing on efficiency to compete with the entrenched banking system. Stripe's acquisition of PayPal represents a bet on this uncertain future.

链捕手42m ago

The Eternal Fragments of Money: Third-Party Payment Has No First Principle

链捕手42m ago

Trading

Spot
活动图片