Author: Long Yue
The market panicked by treating Kimi K3 as "DeepSeek Moment 2.0," but this time, Wall Street's judgment is completely different.
Late night on July 16th, Moonshot AI released Kimi K3 in Shanghai. This open-source model with 2.8 trillion parameters scored 57 on the Artificial Analysis Intelligence Index, ranking third to fourth globally, on par with Anthropic's Claude Opus 4.8 and OpenAI's GPT-5.5. More crucially, on the Frontend Code Arena programming benchmark created by UC Berkeley, K3 topped the chart with 1679 points, surpassing Claude Fable 5 and GPT-5.6 Sol, becoming the first open-source model to outperform all leading closed-source overseas models on an authoritative programming benchmark.
On July 17th, the US semiconductor sector saw a notable decline. The market's knee-jerk reaction was understandable – at the beginning of 2025, the release of DeepSeek R1 triggered a sharp drop in computing power-related stocks, based on the logic: If Chinese models are getting stronger and stronger, do US AI companies still need to invest so much money in computing power? If Chinese models can approach state-of-the-art capabilities at lower costs, will the demand for NVIDIA GPUs, HBM, servers, and networking equipment be reassessed?
However, according to trading desk reports, the latest research notes from banks including UBS, Nomura, Bank of America Merrill Lynch, and Citi argue: Kimi K3 is not a terminator of computing power demand, but an accelerator.
Kimi K3 and DeepSeek R1 are not the same kind of shock. R1 mainly made the market see "efficiency"; K3 highlights "scale." With 2.8 trillion parameters, a 1M token context window, always-on reasoning, native multimodal capabilities, and an MoE architecture, these features do not tell a capital-light story. They will collectively increase pressure on inference, memory, networking, and storage.

How Powerful is Kimi K3?
Kimi K3 was released by Moonshot AI on July 16, 2026. The full model weights are planned to be released on July 27th. It is an open-weight large model with 2.8 trillion parameters, referred to by multiple institutions as the largest open-weight LLM currently available.
Its core configuration includes three points:
First, a 1M token context window. The model can handle longer texts, larger codebases, and more complex enterprise documents and research tasks.
Second, always-on reasoning. Not just simple Q&A, but oriented towards long-chain reasoning and Agent tasks.
Third, native vision capabilities. K3 doesn't just process text; it also handles multimodal tasks like video, images, game development, frontend design, CAD, etc.
Architecturally, K3 uses Kimi Delta Attention, Attention Residuals, and Stable LatentMoE. The MoE part activates 16 experts per token out of 896 experts. Moonshot AI stated that compared to Kimi K2, the overall scaling efficiency has improved by about 2.5x.
This explains why K3 is not a "cheaper K2." Price comparisons compiled by Nomura show that K3's input price is $3 per million tokens, cached hit input is $0.30, and output is $15 per million tokens; according to Artificial Analysis, K3's cost per task is about $0.94. This price is lower than Claude Fable 5's ~$2.75 and Claude Opus 4.8's ~$1.80, close to GPT-5.6 Sol's $1.04, but significantly higher than GLM-5.2's $0.32—$0.47, and far higher than DeepSeek V4 Pro's $0.04.
Thus, K3's positioning is not the lowest price, but getting close to frontier model capabilities at a lower price.

Four Investment Banks Intensively Set the Tone: This is Not a Weakening of Demand
Regarding market concerns about a "DeepSeek Moment" impact, Duan Bing, an analyst at Nomura Securities' Asia-Pacific tech team, wrote in a report: "We believe competition and innovation in the global large model market will not stop. As we approach Artificial General Intelligence (AGI), the application of generative AI in consumer and enterprise sectors will continue to expand. Frontier AI labs and hyperscale cloud platform companies are likely to continue investing in this stage to maintain their competitive positions – as scaling laws continue, we interpret this competition as beneficial for the AI infrastructure value chain."
In his report on July 19th, Citi semiconductor analyst Peter Lee directly titled it "Another Jevons Paradox." What is the Jevons Paradox? Simply put: increased efficiency of coal steam engines led to greater total coal consumption, because more people could afford it and it was applied to more scenarios. It's the same with AI models – when high-quality models become cheaper, developers and enterprises deploy more applications, process more tokens, ultimately leading to increased computing power consumption.
Peter Lee believes that even if K3 is widely adopted, demand for general-purpose memory like server DDR5 and eSSD will still increase. The reason is that K3's inference efficiency is comparable to other frontier models, but its KV cache footprint grows with the context length, increasing pressure on memory rather than reducing it.
Bank of America Securities semiconductor analyst Vivek Arya was more direct in his report on July 17th. He believes that the response from US frontier AI labs is "not less computing power, but more." If Chinese open-source models continue to approach their capabilities, OpenAI, Anthropic, and Google must maintain differentiation through larger-scale training, heavier inference, and faster iteration. Arya also mentioned an easily overlooked background: media reports suggest Google's Gemini 3.5 Pro has been delayed by months from its plan, with programming performance not meeting internal targets, making it "increasingly difficult to defend" the frontier leadership position.
In a report dated July 20th, the team of UBS analyst Timo Arcuri pointed out that parallels between K3 and DeepSeek R1 indeed exist, but K3 is more about scale – it is the world's largest open-source model, with 2.8 trillion parameters and a 1 million token context window. The analysts emphasized that open-source models are generally more memory-intensive than closed-source frontier models because their context windows are longer, and KV cache demand, even after quantization, continues to grow in absolute terms, making deployment of open-source models more dependent on HBM and storage.

Storage: The most direct beneficiary sector. UBS calculations show that the storage and memory sectors are expected to generate cumulative free cash flow (FCF) reaching about 30% of market capitalization by 2028 – the highest among all subsectors, with Micron (MU) alone reaching 47%. Both Citi and Nomura maintain Buy ratings on Samsung Electronics, citing the global memory market being in a state of extremely tight supply. Citi analyst Peter Lee pointed out that Kimi K3's inference-side memory demand is no less than other frontier models, and the expansion of KV cache volume will directly drive demand for server DDR5 and enterprise SSDs (eSSD). He specifically noted: Large-scale deployment of Kimi K3 requires "super-node" cluster configurations with over 64 GPUs.

Computing Infrastructure: TSMC and NVIDIA are the primary beneficiaries. Whether it's scaling laws still effective on the training side or token demand growth on the inference side, both ultimately point to increased demand for advanced node chips. Nomura reiterated Buy ratings on TSMC, ASE, MediaTek, etc. NVIDIA has publicly stated that modern MoE model inference delivers up to 25x performance-per-watt improvement on the GB300 NVL72 compared to the previous Hopper architecture. Models like K3 naturally benefit from NVIDIA's latest hardware.
Networking: Super-node trend creates structural opportunities. Kimi K3 requires super-node clusters. Domestic Chinese computing power, constrained by high-end chip export controls, relies even more on super-node architectures to compensate for single-card performance gaps, driving demand for network layer suppliers like optical modules and optical chips. Nomura is bullish on Inphi and Suzhou Inphi.
Cloud Platforms: Benefiting from ecosystem agglomeration effects. Cloud platforms hosting multiple frontier open-source models have stronger bargaining power, not relying on a single closed-source model supplier. Nomura is bullish on Alibaba (BABA) as a core player in China's AI cloud ecosystem, and data center operators like GDS and 21Vianet (VNET).
This might be the most easily underestimated data point in the entire narrative.
According to statistics from the open API gateway OpenRouter, the proportion of token usage by Chinese AI models in global developer traffic has exceeded 45%, up from less than 2% a year ago. Data from Bank of America Merrill Lynch confirms the acceleration of overall AI penetration: about 55% of US companies have now subscribed to AI models, platforms, or tools, with enterprise adoption rates reaching 42% for Anthropic and 40% for OpenAI. Top AI consumers (the top 1% of enterprise users) spend $4,833 per employee per month on AI.
The market is diverging. On one hand are Chinese open-source models represented by DeepSeek and Kimi K3, covering the economy and mid-to-high-end value-for-money markets; on the other hand, top US frontier models are focusing on more complex workloads (like scientific computing) to maintain technology and pricing premiums. Nomura's judgment is: leading large model players on both the Chinese and US sides will benefit – provided they can continuously stay at the forefront of the technology curve.


After K3's release, the market's first reaction was still to compare it with DeepSeek R1. This comparison is useful but shouldn't stop at the level of "is AI hardware going to be sold off again?"
DeepSeek made the market re-evaluate training efficiency. K3 shows the market another thing: open-source models can also push scale, long context, Agent, and multimodality into frontier territory.
This will force US frontier labs to keep investing, and also allow Chinese models to continue expanding in the global developer ecosystem. Closed-source top models retain technology and pricing premiums, open-source models cover more price points and deployment scenarios, cloud vendors provide model distribution and enterprise adoption, and the hardware chain bears training and inference pressure.
In the short term, trading may fluctuate due to "DeepSeek memories." In the medium term, as long as token usage continues to grow, and long context and Agent continue to spread, computing power, HBM, storage, networking, and IDC remain unavoidable cost items.
This is also why multiple institutions reached similar conclusions after K3: stronger open-source models are not the endpoint for AI infrastructure demand; instead, they could be the entry point for the next wave of demand expansion.
However, Bank of America Merrill Lynch also clearly left a tail risk: "If the speed of efficiency gains outpaces workload growth, we might see a certain pullback in infrastructure construction." In other words, if models become much cheaper, but usage does not expand significantly in tandem, the growth logic for computing power demand would be discounted.





