Replicating the "DeepSeek Moment"? Wall Street Unanimously Says: Kimi K3 Instead Strengthens Computing Power Demand

链捕手Publicado a 2026-07-21Actualizado a 2026-07-21

Resumen

Title: Wall Street Sees Kimi K3 as a Catalyst for Compute Demand, Not a "DeepSeek Moment 2.0" Summary: Following the release of Moonshot AI's powerful open-source model Kimi K3, initial market reaction mirrored the "DeepSeek moment" that sparked a sell-off in compute stocks earlier in 2025, fearing reduced demand for AI infrastructure. However, major Wall Street banks including UBS, Nomura, BofA, and Citi argue the opposite: K3 will accelerate, not weaken, demand for compute, memory, storage, and networking. Their analysis centers on K3's specifications—2.8 trillion parameters, 1M token context, and MoE architecture—which represent a "scale" story rather than a pure "efficiency" one like DeepSeek R1. These features increase pressure on inference, memory (especially KV cache), and storage. Analysts invoke Jevons Paradox: as high-quality models become more affordable (K3 is cheaper than top closed models but not the cheapest), usage and token volumes expand, ultimately increasing total compute consumption. The reports highlight that competition will force leading US AI labs (OpenAI, Anthropic, Google) to invest more in training and iteration to maintain their edge. Furthermore, the rise of capable open-source models like K3 is expanding the global AI developer ecosystem, with Chinese models now accounting for over 45% of developer traffic. Key beneficiaries identified across the AI infrastructure chain include memory/storage players (e.g., Micron, Samsung), compute leaders ...

Author: Long Yue

The market panicked by treating Kimi K3 as "DeepSeek Moment 2.0," but this time, Wall Street's judgment is completely different.

Late night on July 16th, Moonshot AI released Kimi K3 in Shanghai. This open-source model with 2.8 trillion parameters scored 57 on the Artificial Analysis Intelligence Index, ranking third to fourth globally, on par with Anthropic's Claude Opus 4.8 and OpenAI's GPT-5.5. More crucially, on the Frontend Code Arena programming benchmark created by UC Berkeley, K3 topped the chart with 1679 points, surpassing Claude Fable 5 and GPT-5.6 Sol, becoming the first open-source model to outperform all leading closed-source overseas models on an authoritative programming benchmark.

On July 17th, the US semiconductor sector saw a notable decline. The market's knee-jerk reaction was understandable – at the beginning of 2025, the release of DeepSeek R1 triggered a sharp drop in computing power-related stocks, based on the logic: If Chinese models are getting stronger and stronger, do US AI companies still need to invest so much money in computing power? If Chinese models can approach state-of-the-art capabilities at lower costs, will the demand for NVIDIA GPUs, HBM, servers, and networking equipment be reassessed?

However, according to trading desk reports, the latest research notes from banks including UBS, Nomura, Bank of America Merrill Lynch, and Citi argue: Kimi K3 is not a terminator of computing power demand, but an accelerator.

Kimi K3 and DeepSeek R1 are not the same kind of shock. R1 mainly made the market see "efficiency"; K3 highlights "scale." With 2.8 trillion parameters, a 1M token context window, always-on reasoning, native multimodal capabilities, and an MoE architecture, these features do not tell a capital-light story. They will collectively increase pressure on inference, memory, networking, and storage.

How Powerful is Kimi K3?

Kimi K3 was released by Moonshot AI on July 16, 2026. The full model weights are planned to be released on July 27th. It is an open-weight large model with 2.8 trillion parameters, referred to by multiple institutions as the largest open-weight LLM currently available.

Its core configuration includes three points:

First, a 1M token context window. The model can handle longer texts, larger codebases, and more complex enterprise documents and research tasks.

Second, always-on reasoning. Not just simple Q&A, but oriented towards long-chain reasoning and Agent tasks.

Third, native vision capabilities. K3 doesn't just process text; it also handles multimodal tasks like video, images, game development, frontend design, CAD, etc.

Architecturally, K3 uses Kimi Delta Attention, Attention Residuals, and Stable LatentMoE. The MoE part activates 16 experts per token out of 896 experts. Moonshot AI stated that compared to Kimi K2, the overall scaling efficiency has improved by about 2.5x.

This explains why K3 is not a "cheaper K2." Price comparisons compiled by Nomura show that K3's input price is $3 per million tokens, cached hit input is $0.30, and output is $15 per million tokens; according to Artificial Analysis, K3's cost per task is about $0.94. This price is lower than Claude Fable 5's ~$2.75 and Claude Opus 4.8's ~$1.80, close to GPT-5.6 Sol's $1.04, but significantly higher than GLM-5.2's $0.32—$0.47, and far higher than DeepSeek V4 Pro's $0.04.

Thus, K3's positioning is not the lowest price, but getting close to frontier model capabilities at a lower price.

Four Investment Banks Intensively Set the Tone: This is Not a Weakening of Demand

Regarding market concerns about a "DeepSeek Moment" impact, Duan Bing, an analyst at Nomura Securities' Asia-Pacific tech team, wrote in a report: "We believe competition and innovation in the global large model market will not stop. As we approach Artificial General Intelligence (AGI), the application of generative AI in consumer and enterprise sectors will continue to expand. Frontier AI labs and hyperscale cloud platform companies are likely to continue investing in this stage to maintain their competitive positions – as scaling laws continue, we interpret this competition as beneficial for the AI infrastructure value chain."

In his report on July 19th, Citi semiconductor analyst Peter Lee directly titled it "Another Jevons Paradox." What is the Jevons Paradox? Simply put: increased efficiency of coal steam engines led to greater total coal consumption, because more people could afford it and it was applied to more scenarios. It's the same with AI models – when high-quality models become cheaper, developers and enterprises deploy more applications, process more tokens, ultimately leading to increased computing power consumption.

Peter Lee believes that even if K3 is widely adopted, demand for general-purpose memory like server DDR5 and eSSD will still increase. The reason is that K3's inference efficiency is comparable to other frontier models, but its KV cache footprint grows with the context length, increasing pressure on memory rather than reducing it.

Bank of America Securities semiconductor analyst Vivek Arya was more direct in his report on July 17th. He believes that the response from US frontier AI labs is "not less computing power, but more." If Chinese open-source models continue to approach their capabilities, OpenAI, Anthropic, and Google must maintain differentiation through larger-scale training, heavier inference, and faster iteration. Arya also mentioned an easily overlooked background: media reports suggest Google's Gemini 3.5 Pro has been delayed by months from its plan, with programming performance not meeting internal targets, making it "increasingly difficult to defend" the frontier leadership position.

In a report dated July 20th, the team of UBS analyst Timo Arcuri pointed out that parallels between K3 and DeepSeek R1 indeed exist, but K3 is more about scale – it is the world's largest open-source model, with 2.8 trillion parameters and a 1 million token context window. The analysts emphasized that open-source models are generally more memory-intensive than closed-source frontier models because their context windows are longer, and KV cache demand, even after quantization, continues to grow in absolute terms, making deployment of open-source models more dependent on HBM and storage.

Storage: The most direct beneficiary sector. UBS calculations show that the storage and memory sectors are expected to generate cumulative free cash flow (FCF) reaching about 30% of market capitalization by 2028 – the highest among all subsectors, with Micron (MU) alone reaching 47%. Both Citi and Nomura maintain Buy ratings on Samsung Electronics, citing the global memory market being in a state of extremely tight supply. Citi analyst Peter Lee pointed out that Kimi K3's inference-side memory demand is no less than other frontier models, and the expansion of KV cache volume will directly drive demand for server DDR5 and enterprise SSDs (eSSD). He specifically noted: Large-scale deployment of Kimi K3 requires "super-node" cluster configurations with over 64 GPUs.

Computing Infrastructure: TSMC and NVIDIA are the primary beneficiaries. Whether it's scaling laws still effective on the training side or token demand growth on the inference side, both ultimately point to increased demand for advanced node chips. Nomura reiterated Buy ratings on TSMC, ASE, MediaTek, etc. NVIDIA has publicly stated that modern MoE model inference delivers up to 25x performance-per-watt improvement on the GB300 NVL72 compared to the previous Hopper architecture. Models like K3 naturally benefit from NVIDIA's latest hardware.

Networking: Super-node trend creates structural opportunities. Kimi K3 requires super-node clusters. Domestic Chinese computing power, constrained by high-end chip export controls, relies even more on super-node architectures to compensate for single-card performance gaps, driving demand for network layer suppliers like optical modules and optical chips. Nomura is bullish on Inphi and Suzhou Inphi.

Cloud Platforms: Benefiting from ecosystem agglomeration effects. Cloud platforms hosting multiple frontier open-source models have stronger bargaining power, not relying on a single closed-source model supplier. Nomura is bullish on Alibaba (BABA) as a core player in China's AI cloud ecosystem, and data center operators like GDS and 21Vianet (VNET).

This might be the most easily underestimated data point in the entire narrative.

According to statistics from the open API gateway OpenRouter, the proportion of token usage by Chinese AI models in global developer traffic has exceeded 45%, up from less than 2% a year ago. Data from Bank of America Merrill Lynch confirms the acceleration of overall AI penetration: about 55% of US companies have now subscribed to AI models, platforms, or tools, with enterprise adoption rates reaching 42% for Anthropic and 40% for OpenAI. Top AI consumers (the top 1% of enterprise users) spend $4,833 per employee per month on AI.

The market is diverging. On one hand are Chinese open-source models represented by DeepSeek and Kimi K3, covering the economy and mid-to-high-end value-for-money markets; on the other hand, top US frontier models are focusing on more complex workloads (like scientific computing) to maintain technology and pricing premiums. Nomura's judgment is: leading large model players on both the Chinese and US sides will benefit – provided they can continuously stay at the forefront of the technology curve.

After K3's release, the market's first reaction was still to compare it with DeepSeek R1. This comparison is useful but shouldn't stop at the level of "is AI hardware going to be sold off again?"

DeepSeek made the market re-evaluate training efficiency. K3 shows the market another thing: open-source models can also push scale, long context, Agent, and multimodality into frontier territory.

This will force US frontier labs to keep investing, and also allow Chinese models to continue expanding in the global developer ecosystem. Closed-source top models retain technology and pricing premiums, open-source models cover more price points and deployment scenarios, cloud vendors provide model distribution and enterprise adoption, and the hardware chain bears training and inference pressure.

In the short term, trading may fluctuate due to "DeepSeek memories." In the medium term, as long as token usage continues to grow, and long context and Agent continue to spread, computing power, HBM, storage, networking, and IDC remain unavoidable cost items.

This is also why multiple institutions reached similar conclusions after K3: stronger open-source models are not the endpoint for AI infrastructure demand; instead, they could be the entry point for the next wave of demand expansion.

However, Bank of America Merrill Lynch also clearly left a tail risk: "If the speed of efficiency gains outpaces workload growth, we might see a certain pullback in infrastructure construction." In other words, if models become much cheaper, but usage does not expand significantly in tandem, the growth logic for computing power demand would be discounted.

Preguntas relacionadas

QWhat is the main argument presented by Wall Street analysts regarding the impact of Kimi K3 on AI infrastructure demand?

AWall Street analysts from firms like UBS, Nomura, Bank of America Merrill Lynch, and Citi argue that Kimi K3 is not a threat to AI infrastructure demand but rather an accelerator. They believe its advanced features (large parameters, long context, MoE architecture) will increase pressure on computing, memory, network, and storage, driving more demand for hardware like GPUs, HBM, and servers, similar to the 'Jevons Paradox' where efficiency gains lead to higher overall consumption.

QHow does Kimi K3 differ from DeepSeek R1 in terms of its market impact and technical characteristics?

ADeepSeek R1 primarily impacted the market by demonstrating 'efficiency'—showing that capable models could be built cost-effectively, leading to concerns about reduced spending on compute. Kimi K3, however, emphasizes 'scale.' With 2.8 trillion parameters, a 1M token context window, native multimodal capabilities, and an MoE architecture, K3 represents a large, complex model that is not a 'light asset' story. It increases demands on inference, memory, and storage, thus potentially boosting infrastructure demand rather than reducing it.

QAccording to the analysts, which specific hardware sectors are expected to benefit from the deployment of models like Kimi K3?

AAnalysts identify several hardware sectors as beneficiaries: 1) **Storage & Memory**: Due to increased KV cache needs from long contexts, boosting demand for server DDR5 and enterprise SSDs. Companies like Micron (MU) and Samsung are highlighted. 2) **Compute Infrastructure**: Companies like TSMC and NVIDIA benefit as larger models and growing token usage drive demand for advanced chips. 3) **Networking**: The need for 'super-node' clusters to run K3 boosts demand for optical modules and components (e.g., Zhongji Innolight). 4) **Cloud Platforms & Data Centers**: Platforms hosting multiple models (e.g., Alibaba Cloud) and data center operators (e.g., GDS, VNET) gain from ecosystem aggregation.

QWhat is the 'Jevons Paradox' mentioned in the context of AI models like Kimi K3?

AThe 'Jevons Paradox,' referenced by Citi analyst Peter Lee, describes a phenomenon where increased efficiency in using a resource (like coal in steam engines) leads to an increase in the total consumption of that resource, because lower costs make it accessible for more uses and scenarios. Applied to AI, as models like Kimi K3 become more capable and cost-effective, developers and enterprises will deploy more applications and process more tokens, ultimately leading to higher overall computing power consumption rather than a reduction.

QWhat potential tail risk do analysts see for the AI infrastructure growth narrative despite the demand boost from models like K3?

ABank of America Merrill Lynch notes a tail risk: If the rate of efficiency gains in model development outpaces the growth in actual AI workloads and usage, there could be a potential slowdown or 'pullback' in infrastructure investment. In other words, if models become significantly cheaper to run but the expansion in their application volume does not keep up, the fundamental logic for continuous, massive growth in computing demand could be undermined.

Lecturas Relacionadas

The Encryption Bill Clarity's Challenge: A Thorny Path of Bipartisan Compromise in the U.S.

U.S. lawmakers are attempting to advance the Clarity Act, a significant crypto market structure bill, but its path is fraught with partisan hurdles. The process has been rocky since January, when a prior bipartisan deal in the Senate Banking Committee was upended. A key compromise in May on "yield" issues allowed the bill to move forward in committee, but only with the conditional support of two Democratic senators, Angela Alsobrooks and Ruben Gallego. They emphasized that their final vote depends on reaching an agreement on ethics provisions for elected officials. Ultimately, the Senate Agriculture Committee passed its version along party lines without Democratic support. As Republicans push for a full Senate vote in July, the demand for strong ethics language has expanded beyond Democrats. Additional controversies surround provisions related to yields (aligning some Republicans with large banks) and developer protections (opposed by enforcement agencies). Core concerns about illicit finance and consumer protection remain central to the debate. Despite consensus on the need for legislation, achieving the necessary bipartisan compromise is proving difficult. While momentum exists—including recent meetings between senators and White House officials—a reconciled bill text faces skepticism. Senator Gallego has stated that without ethics terms acceptable to Democrats, they will not provide the needed votes. The immediate goals for the crypto community in Congress are unclear: a symbolic Senate vote before the August recess, eventual passage into law by 2026, or forging a final compromise framework. The arduous, vote-by-vote effort to build bipartisan support continues, mirroring the traditional legislative grind the industry must now navigate.

Foresight NewsHace 48 min(s)

The Encryption Bill Clarity's Challenge: A Thorny Path of Bipartisan Compromise in the U.S.

Foresight NewsHace 48 min(s)

Kalshi and Polymarket Founders at Odds? This Business War Is Far More Brutal Than You Imagine

The New York Times details the fierce, personal rivalry between Kalshi CEO Tarek Mansour and Polymarket founder Shayne Coplan, which has escalated beyond typical business competition into a conflict marked by legal complaints, regulatory battles, and public hostilities. The feud intensified in late 2024 when FBI agents raided Coplan's New York apartment. While Coplan publicly blamed political motives, sources indicate his team privately suspected Mansour, noting that Kalshi's lawyers had previously reported Polymarket's operational model to federal prosecutors, highlighting that U.S. users could still access its offshore platform despite a ban. The animosity extends through their companies' operations. Kalshi positions itself as a compliance-focused, fully licensed U.S. operator, while Polymarket has historically operated its core platform offshore without a U.S. license, offering more anonymity and controversial betting markets. Mansour has publicly called Polymarket's model "illegal and immoral," while Coplan privately dismisses Kalshi as a copycat. Their competition has played out in Washington lobbying, attempts to sabotage each other's major deals (such as Kalshi's efforts to dissuade Intercontinental Exchange from investing in Polymarket), competing sponsorships, and poaching staff. The rivalry continues as both platforms experience massive growth, with Kalshi currently holding a valuation and trading volume edge, but facing ongoing regulatory scrutiny alongside Polymarket.

Foresight NewsHace 1 hora(s)

Kalshi and Polymarket Founders at Odds? This Business War Is Far More Brutal Than You Imagine

Foresight NewsHace 1 hora(s)

Trading

Spot
活动图片