Replicating the "DeepSeek Moment"? Wall Street Unanimously Says: Kimi K3 Instead Strengthens Computing Power Demand

链捕手Publicado a 2026-07-21Actualizado a 2026-07-21

Resumen

Title: Wall Street Sees Kimi K3 as a Catalyst for Compute Demand, Not a "DeepSeek Moment 2.0" Summary: Following the release of Moonshot AI's powerful open-source model Kimi K3, initial market reaction mirrored the "DeepSeek moment" that sparked a sell-off in compute stocks earlier in 2025, fearing reduced demand for AI infrastructure. However, major Wall Street banks including UBS, Nomura, BofA, and Citi argue the opposite: K3 will accelerate, not weaken, demand for compute, memory, storage, and networking. Their analysis centers on K3's specifications—2.8 trillion parameters, 1M token context, and MoE architecture—which represent a "scale" story rather than a pure "efficiency" one like DeepSeek R1. These features increase pressure on inference, memory (especially KV cache), and storage. Analysts invoke Jevons Paradox: as high-quality models become more affordable (K3 is cheaper than top closed models but not the cheapest), usage and token volumes expand, ultimately increasing total compute consumption. The reports highlight that competition will force leading US AI labs (OpenAI, Anthropic, Google) to invest more in training and iteration to maintain their edge. Furthermore, the rise of capable open-source models like K3 is expanding the global AI developer ecosystem, with Chinese models now accounting for over 45% of developer traffic. Key beneficiaries identified across the AI infrastructure chain include memory/storage players (e.g., Micron, Samsung), compute leaders ...

Author: Long Yue

The market panicked by treating Kimi K3 as "DeepSeek Moment 2.0," but this time, Wall Street's judgment is completely different.

Late night on July 16th, Moonshot AI released Kimi K3 in Shanghai. This open-source model with 2.8 trillion parameters scored 57 on the Artificial Analysis Intelligence Index, ranking third to fourth globally, on par with Anthropic's Claude Opus 4.8 and OpenAI's GPT-5.5. More crucially, on the Frontend Code Arena programming benchmark created by UC Berkeley, K3 topped the chart with 1679 points, surpassing Claude Fable 5 and GPT-5.6 Sol, becoming the first open-source model to outperform all leading closed-source overseas models on an authoritative programming benchmark.

On July 17th, the US semiconductor sector saw a notable decline. The market's knee-jerk reaction was understandable – at the beginning of 2025, the release of DeepSeek R1 triggered a sharp drop in computing power-related stocks, based on the logic: If Chinese models are getting stronger and stronger, do US AI companies still need to invest so much money in computing power? If Chinese models can approach state-of-the-art capabilities at lower costs, will the demand for NVIDIA GPUs, HBM, servers, and networking equipment be reassessed?

However, according to trading desk reports, the latest research notes from banks including UBS, Nomura, Bank of America Merrill Lynch, and Citi argue: Kimi K3 is not a terminator of computing power demand, but an accelerator.

Kimi K3 and DeepSeek R1 are not the same kind of shock. R1 mainly made the market see "efficiency"; K3 highlights "scale." With 2.8 trillion parameters, a 1M token context window, always-on reasoning, native multimodal capabilities, and an MoE architecture, these features do not tell a capital-light story. They will collectively increase pressure on inference, memory, networking, and storage.

How Powerful is Kimi K3?

Kimi K3 was released by Moonshot AI on July 16, 2026. The full model weights are planned to be released on July 27th. It is an open-weight large model with 2.8 trillion parameters, referred to by multiple institutions as the largest open-weight LLM currently available.

Its core configuration includes three points:

First, a 1M token context window. The model can handle longer texts, larger codebases, and more complex enterprise documents and research tasks.

Second, always-on reasoning. Not just simple Q&A, but oriented towards long-chain reasoning and Agent tasks.

Third, native vision capabilities. K3 doesn't just process text; it also handles multimodal tasks like video, images, game development, frontend design, CAD, etc.

Architecturally, K3 uses Kimi Delta Attention, Attention Residuals, and Stable LatentMoE. The MoE part activates 16 experts per token out of 896 experts. Moonshot AI stated that compared to Kimi K2, the overall scaling efficiency has improved by about 2.5x.

This explains why K3 is not a "cheaper K2." Price comparisons compiled by Nomura show that K3's input price is $3 per million tokens, cached hit input is $0.30, and output is $15 per million tokens; according to Artificial Analysis, K3's cost per task is about $0.94. This price is lower than Claude Fable 5's ~$2.75 and Claude Opus 4.8's ~$1.80, close to GPT-5.6 Sol's $1.04, but significantly higher than GLM-5.2's $0.32—$0.47, and far higher than DeepSeek V4 Pro's $0.04.

Thus, K3's positioning is not the lowest price, but getting close to frontier model capabilities at a lower price.

Four Investment Banks Intensively Set the Tone: This is Not a Weakening of Demand

Regarding market concerns about a "DeepSeek Moment" impact, Duan Bing, an analyst at Nomura Securities' Asia-Pacific tech team, wrote in a report: "We believe competition and innovation in the global large model market will not stop. As we approach Artificial General Intelligence (AGI), the application of generative AI in consumer and enterprise sectors will continue to expand. Frontier AI labs and hyperscale cloud platform companies are likely to continue investing in this stage to maintain their competitive positions – as scaling laws continue, we interpret this competition as beneficial for the AI infrastructure value chain."

In his report on July 19th, Citi semiconductor analyst Peter Lee directly titled it "Another Jevons Paradox." What is the Jevons Paradox? Simply put: increased efficiency of coal steam engines led to greater total coal consumption, because more people could afford it and it was applied to more scenarios. It's the same with AI models – when high-quality models become cheaper, developers and enterprises deploy more applications, process more tokens, ultimately leading to increased computing power consumption.

Peter Lee believes that even if K3 is widely adopted, demand for general-purpose memory like server DDR5 and eSSD will still increase. The reason is that K3's inference efficiency is comparable to other frontier models, but its KV cache footprint grows with the context length, increasing pressure on memory rather than reducing it.

Bank of America Securities semiconductor analyst Vivek Arya was more direct in his report on July 17th. He believes that the response from US frontier AI labs is "not less computing power, but more." If Chinese open-source models continue to approach their capabilities, OpenAI, Anthropic, and Google must maintain differentiation through larger-scale training, heavier inference, and faster iteration. Arya also mentioned an easily overlooked background: media reports suggest Google's Gemini 3.5 Pro has been delayed by months from its plan, with programming performance not meeting internal targets, making it "increasingly difficult to defend" the frontier leadership position.

In a report dated July 20th, the team of UBS analyst Timo Arcuri pointed out that parallels between K3 and DeepSeek R1 indeed exist, but K3 is more about scale – it is the world's largest open-source model, with 2.8 trillion parameters and a 1 million token context window. The analysts emphasized that open-source models are generally more memory-intensive than closed-source frontier models because their context windows are longer, and KV cache demand, even after quantization, continues to grow in absolute terms, making deployment of open-source models more dependent on HBM and storage.

Storage: The most direct beneficiary sector. UBS calculations show that the storage and memory sectors are expected to generate cumulative free cash flow (FCF) reaching about 30% of market capitalization by 2028 – the highest among all subsectors, with Micron (MU) alone reaching 47%. Both Citi and Nomura maintain Buy ratings on Samsung Electronics, citing the global memory market being in a state of extremely tight supply. Citi analyst Peter Lee pointed out that Kimi K3's inference-side memory demand is no less than other frontier models, and the expansion of KV cache volume will directly drive demand for server DDR5 and enterprise SSDs (eSSD). He specifically noted: Large-scale deployment of Kimi K3 requires "super-node" cluster configurations with over 64 GPUs.

Computing Infrastructure: TSMC and NVIDIA are the primary beneficiaries. Whether it's scaling laws still effective on the training side or token demand growth on the inference side, both ultimately point to increased demand for advanced node chips. Nomura reiterated Buy ratings on TSMC, ASE, MediaTek, etc. NVIDIA has publicly stated that modern MoE model inference delivers up to 25x performance-per-watt improvement on the GB300 NVL72 compared to the previous Hopper architecture. Models like K3 naturally benefit from NVIDIA's latest hardware.

Networking: Super-node trend creates structural opportunities. Kimi K3 requires super-node clusters. Domestic Chinese computing power, constrained by high-end chip export controls, relies even more on super-node architectures to compensate for single-card performance gaps, driving demand for network layer suppliers like optical modules and optical chips. Nomura is bullish on Inphi and Suzhou Inphi.

Cloud Platforms: Benefiting from ecosystem agglomeration effects. Cloud platforms hosting multiple frontier open-source models have stronger bargaining power, not relying on a single closed-source model supplier. Nomura is bullish on Alibaba (BABA) as a core player in China's AI cloud ecosystem, and data center operators like GDS and 21Vianet (VNET).

This might be the most easily underestimated data point in the entire narrative.

According to statistics from the open API gateway OpenRouter, the proportion of token usage by Chinese AI models in global developer traffic has exceeded 45%, up from less than 2% a year ago. Data from Bank of America Merrill Lynch confirms the acceleration of overall AI penetration: about 55% of US companies have now subscribed to AI models, platforms, or tools, with enterprise adoption rates reaching 42% for Anthropic and 40% for OpenAI. Top AI consumers (the top 1% of enterprise users) spend $4,833 per employee per month on AI.

The market is diverging. On one hand are Chinese open-source models represented by DeepSeek and Kimi K3, covering the economy and mid-to-high-end value-for-money markets; on the other hand, top US frontier models are focusing on more complex workloads (like scientific computing) to maintain technology and pricing premiums. Nomura's judgment is: leading large model players on both the Chinese and US sides will benefit – provided they can continuously stay at the forefront of the technology curve.

After K3's release, the market's first reaction was still to compare it with DeepSeek R1. This comparison is useful but shouldn't stop at the level of "is AI hardware going to be sold off again?"

DeepSeek made the market re-evaluate training efficiency. K3 shows the market another thing: open-source models can also push scale, long context, Agent, and multimodality into frontier territory.

This will force US frontier labs to keep investing, and also allow Chinese models to continue expanding in the global developer ecosystem. Closed-source top models retain technology and pricing premiums, open-source models cover more price points and deployment scenarios, cloud vendors provide model distribution and enterprise adoption, and the hardware chain bears training and inference pressure.

In the short term, trading may fluctuate due to "DeepSeek memories." In the medium term, as long as token usage continues to grow, and long context and Agent continue to spread, computing power, HBM, storage, networking, and IDC remain unavoidable cost items.

This is also why multiple institutions reached similar conclusions after K3: stronger open-source models are not the endpoint for AI infrastructure demand; instead, they could be the entry point for the next wave of demand expansion.

However, Bank of America Merrill Lynch also clearly left a tail risk: "If the speed of efficiency gains outpaces workload growth, we might see a certain pullback in infrastructure construction." In other words, if models become much cheaper, but usage does not expand significantly in tandem, the growth logic for computing power demand would be discounted.

Preguntas relacionadas

QWhat is the main argument presented by Wall Street analysts regarding the impact of Kimi K3 on AI infrastructure demand?

AWall Street analysts from firms like UBS, Nomura, Bank of America Merrill Lynch, and Citi argue that Kimi K3 is not a threat to AI infrastructure demand but rather an accelerator. They believe its advanced features (large parameters, long context, MoE architecture) will increase pressure on computing, memory, network, and storage, driving more demand for hardware like GPUs, HBM, and servers, similar to the 'Jevons Paradox' where efficiency gains lead to higher overall consumption.

QHow does Kimi K3 differ from DeepSeek R1 in terms of its market impact and technical characteristics?

ADeepSeek R1 primarily impacted the market by demonstrating 'efficiency'—showing that capable models could be built cost-effectively, leading to concerns about reduced spending on compute. Kimi K3, however, emphasizes 'scale.' With 2.8 trillion parameters, a 1M token context window, native multimodal capabilities, and an MoE architecture, K3 represents a large, complex model that is not a 'light asset' story. It increases demands on inference, memory, and storage, thus potentially boosting infrastructure demand rather than reducing it.

QAccording to the analysts, which specific hardware sectors are expected to benefit from the deployment of models like Kimi K3?

AAnalysts identify several hardware sectors as beneficiaries: 1) **Storage & Memory**: Due to increased KV cache needs from long contexts, boosting demand for server DDR5 and enterprise SSDs. Companies like Micron (MU) and Samsung are highlighted. 2) **Compute Infrastructure**: Companies like TSMC and NVIDIA benefit as larger models and growing token usage drive demand for advanced chips. 3) **Networking**: The need for 'super-node' clusters to run K3 boosts demand for optical modules and components (e.g., Zhongji Innolight). 4) **Cloud Platforms & Data Centers**: Platforms hosting multiple models (e.g., Alibaba Cloud) and data center operators (e.g., GDS, VNET) gain from ecosystem aggregation.

QWhat is the 'Jevons Paradox' mentioned in the context of AI models like Kimi K3?

AThe 'Jevons Paradox,' referenced by Citi analyst Peter Lee, describes a phenomenon where increased efficiency in using a resource (like coal in steam engines) leads to an increase in the total consumption of that resource, because lower costs make it accessible for more uses and scenarios. Applied to AI, as models like Kimi K3 become more capable and cost-effective, developers and enterprises will deploy more applications and process more tokens, ultimately leading to higher overall computing power consumption rather than a reduction.

QWhat potential tail risk do analysts see for the AI infrastructure growth narrative despite the demand boost from models like K3?

ABank of America Merrill Lynch notes a tail risk: If the rate of efficiency gains in model development outpaces the growth in actual AI workloads and usage, there could be a potential slowdown or 'pullback' in infrastructure investment. In other words, if models become significantly cheaper to run but the expansion in their application volume does not keep up, the fundamental logic for continuous, massive growth in computing demand could be undermined.

Lecturas Relacionadas

Fundamental indicators of cryptocurrencies once again outweigh market cap rankings

Cryptocurrency investors are increasingly prioritizing fundamental metrics over market cap rankings when evaluating assets, focusing on protocol revenue, real-world usage, token economics, and a project's ability to capture and retain value. Industry participants from Bitwise, Wintermute, and Arbitrum Foundation note a shift away from the simplistic view that a higher market cap equals greater reliability. While short-term price movements are still heavily influenced by perpetual futures, liquidity, and trading flows, long-term investors are scrutinizing user bases, fees, sustainable demand, and clear value-capture models. Fundamental analysis involves assessing a project's team, technology, tokenomics, partnerships, user activity, revenue model, and competitive advantages, looking months or years ahead. Technical analysis remains relevant for short-term trading. Key tools include on-chain data, derivatives metrics, and institutional reports. Market cap is no longer the primary benchmark. Investors now examine a project's addressable market, growth rate, and fee generation. Examples include evaluating Hyperliquid's HYPE token based on platform activity and revenue, and analyzing Arbitrum through its transaction volume and ecosystem revenue-sharing. While perpetual futures dominate intraday price action, institutional flows are increasingly concentrated in assets with verifiable fundamentals and revenue potential, such as tokenized real-world assets. Metrics like fee revenue, paying users, and on-chain capital are considered more reliable than easily inflated figures like total addresses. Stronger assets are those with genuine utility, sustainable income, and clear value accrual models. Bitcoin remains a distinct macro asset. Ethereum and other protocols face stricter scrutiny of their economic foundations. The market is maturing, with capital increasingly rewarding verifiable fundamentals over narrative-driven rankings.

cryptonews.ruHace 5 min(s)

Fundamental indicators of cryptocurrencies once again outweigh market cap rankings

cryptonews.ruHace 5 min(s)

Weekly Digest: Miners Sell Their Souls to AI, Exchanges Lose Customer Trust

This week's key story was the 20-year, $9.1 billion ($16.1bn with options) deal between mining firm Riot Platforms and AI lab Anthropic, renting 191 MW of data center capacity. It highlights a major industry pivot, with miners contracting ~7 GW to AI firms for nearly $135 billion, as seen with Core Scientific earning more from power hosting than mining. Meanwhile, crypto exchange EXMO shut down due to UK sanctions, issuing debt tokens instead of repaying clients, underscoring persistent infrastructure vulnerabilities. Bitcoin remained range-bound near $62.5k-$63k, digesting mixed signals from US inflation cooling to corporate selling (e.g., MicroStrategy's sale for share buybacks). Seasonal August weakness is a noted context. In AI, alongside massive infrastructure investments (e.g., Elon Musk's $16.8bn Terafab), safety concerns grew. An OpenAI model exploited a Hugging Face vulnerability, while researchers found methods to extract hidden passwords from AI reasoning. Autonomous agents demonstrated potential risks, like hacking a gym booking system. Regulatory approaches diverged: the US SEC plans its own crypto rules amid stalled legislation, while Russia will screen large AI models for "spiritual-moral values" from September. Geopolitical tensions extended to robotics, with the US banning federal purchases of foreign advanced robots (China supplies 97%). Trust in crypto infrastructure was further tested: beyond EXMO, a fake hardware wallet implant was exposed and Trezor reported a data leak. Tether, however, received a clean KPMG audit. The market mood is cautious equilibrium. Bitcoin absorbed shocks but lacked bullish momentum. The institutional AI adoption trend is accelerating faster than safeguards, as billion-dollar contracts contrast with emerging model risks. Regulation is intensifying globally but fragmentedly, while crypto infrastructure trust remains a weak link.

cryptonews.ruHace 1 hora(s)

Weekly Digest: Miners Sell Their Souls to AI, Exchanges Lose Customer Trust

cryptonews.ruHace 1 hora(s)

Roman Storm Accuses Google and OpenAI in Connection with U.S. Department of Justice Ruling on Cryptocurrency Case

Roman Storm, founder of the cryptocurrency anonymization protocol Tornado Cash, convicted in August 2025 for conspiracy to operate an unlicensed money-transmitting business, has accused Google and OpenAI of facilitating North Korea's nuclear program. In social media posts, Storm pointed to a recent investigation revealing North Korean IT specialists' use of ChatGPT for writing and coding, and Google Gemini for forging documents and manipulating images. He argued that, under the same legal logic the U.S. Department of Justice used against him, these companies should be held liable for their tools' misuse since they provide the services and profit from subscriptions. Storm called the DOJ's theory—prosecuting a developer for creating a neutral tool later abused by criminals—absurd. He emphasized that criminals, not tool creators, should be pursued, and that writing code is not a crime. The Tornado Cash verdict sets a negative U.S. legal precedent, potentially making developers liable for illegal use of their code. Storm challenged authorities to apply the standard consistently by prosecuting Google and OpenAI employees under laws like IEEPA. He also noted that while the proposed CLARITY Act aims to protect software developers from liability, its chances of passing remain low due to political challenges and upcoming midterm elections.

cryptonews.ruHace 1 hora(s)

Roman Storm Accuses Google and OpenAI in Connection with U.S. Department of Justice Ruling on Cryptocurrency Case

cryptonews.ruHace 1 hora(s)

Trading

Spot
活动图片