AMD has agreed to acquire Toronto-based startup Taalas, which addresses the main problem of neural network inference - the need to constantly transfer model weights from memory to the processor for generating each token. Taalas chips do away with this operation: the model weights are literally etched into the transistors. This data transfer is what limits the speed of modern inference and has turned high-bandwidth memory (HBM) into the most scarce commodity in the semiconductor industry.
The deal will be perceived as another round of competition between AMD and Nvidia in the inference sphere, but it changes little on that front. Much more interesting is what the acquisition says about the state of the memory market - the most overheated segment in semiconductors right now.
What Taalas created
Taalas was founded in 2023 by Ljubisa Bajic, who previously founded chip company Tenstorrent, and his wife Lejla Bajic, a veteran of ATI and AMD engineering divisions, who took the COO role. The company raised $219 million from Fidelity, Quiet Capital, and semiconductor investor Pierre Lamond - the funds went towards developing so-called model-specific integrated circuits.
The first test chip, manufactured using TSMC's 6nm process, served Meta's Llama 3.1 8B model at a speed of 16,960 tokens per second. According to the company, this is 48 times faster than Nvidia GPUs and 8.5 times faster than Cerebras accelerators at the time of comparison. A second-generation chip, designed for models with 20 billion parameters, is set to be released this year.
The architecture is divided into two zones: an area where model weights are hardwired as mask ROM, and regular SRAM for caches and fine-tuning adapters, which can still be changed. "It's this hard programming that partly gives us our speed," Bajic told The Next Platform in February.
AMD plans to integrate Taalas chips into Helios racks using a split scheme: Instinct accelerators will handle the prompt, while Taalas silicon will generate tokens, all managed by the ROCm software stack. The company emphasized that this is an acquisition, not just a team hire - the deal closure is expected in the fourth quarter. AMD's Senior Vice President of AI, Vamsi Boppana, described the purchase as a platform expansion.
The compromise Taalas makes
Etching weights into silicon has an obvious cost: each chip serves exactly one model - forever. Switching models requires partially recalculating the topology, and even with Taalas's shortened cycle - where only two metal layers on an almost-ready wafer are changed - it takes about two months using TSMC's capacity. Top models are updated faster. The bet only pays off where a model is stable, widely used, and valuable enough to be frozen.
This same compromise explains why the deal doesn't change the balance of power in the inference market. Nvidia paid $20 billion for Groq in December - its largest acquisition ever - to integrate specialized token-generation hardware precisely into the platform around which the entire industry is already built. The fight for inference is happening at the ecosystem and installed software level, and a chip tied to one model participates in neither. What Taalas's approach truly proves is a narrower but more interesting thesis: the memory bottleneck plaguing inference is an engineering problem, not a physical given, and it can be bypassed.
The memory question
The market is currently pricing in the assumption that this memory shortage is permanent. Prices for ordinary DRAM rose nearly 90% in the first quarter. High-bandwidth memory (HBM) is essentially sold out for all of 2026, with the HBM market volume this year expected to reach $54.6 billion. SK hynix, which controls more than half of HBM supply, surpassed a $1 trillion market cap and announced new factory construction worth $38.1 billion. For the average investor, the entire bet on memory hinges on one assumption: AI demand will keep memory scarce for years to come.
This assumption is already under attack from several sides. Taalas completely removes memory for weights from the model-serving process, and Nvidia engineers are compressing and quantizing models precisely to reduce the memory footprint of a deployed model.
Memory manufacturers themselves are working in the same direction. Samsung's zHBM technology, showcased at the FMS conference last week, stacks memory directly on the accelerator, multiplying effective bandwidth. SK hynix and Sandisk just published the first standard for high-speed flash memory aimed at replacing cheap NAND for the work HBM does today. Virtually every major industry player is funding its own way to reduce the need for the very resource the market believes will be permanently scarce.
AMD bought proof that inference can work without accessing the component whose scarcity defines the entire current AI demand cycle - and will sell this proof inside racks that still contain GPUs and HBM. Investors who consider today's memory prices a permanent feature of the entire AI infrastructure build cycle are betting against a large and growing engineering movement aimed at the opposite result.
Memory has always been a cyclical business. Those paying today's prices for it have just financed another reason why it will remain so.
AI Opinion
From the perspective of machine data analysis, the key precedent for Taalas chips can be found not in the world of AI, but in the crypto industry. Specialized integrated circuits for Bitcoin mining solve a similar problem - etching an algorithm into silicon for speed - and get the same side effect: complete inflexibility. Hash Telegraph has already described how hardware specialization, taken to the extreme, creates the risk of obsolescence when the underlying algorithm changes or a new computational paradigm emerges.
Technical aspects the article doesn't detail concern the economics of downtime: while a Taalas chip serves one model, competing labs release new versions every few months, and the topology re-programming cycle takes about two months on TSMC's capacity. ASIC history shows that specialization pays off only on a stable algorithm - the question is whether the architecture of large language models will be stable enough for this bet.
end-content





