Artículos Relacionados con Inference

El Centro de Noticias de HTX ofrece los artículos más recientes y un análisis profundo sobre "Inference", cubriendo tendencias del mercado, actualizaciones de proyectos, desarrollos tecnológicos y políticas regulatorias en la industria de cripto.

NVIDIA Stuns with First Vera Rubin Test, DeepSeek Throughput Skyrockets 30x

NVIDIA has unveiled the first on-silicon test results for its next-generation flagship cabinet, the Vera Rubin NVL72, using DeepSeek-V4-Pro on real "agent coding" workloads. The results are staggering: compared to the current leading GB300 NVL72, Vera Rubin delivers up to a 30x increase in throughput per megawatt and reduces token generation cost by up to 35x. The benchmark used was the AgentX test from SemiAnalysis, which captures full AI agent workflows with growing context, tool calls, and sub-agent generation, moving beyond traditional LLM benchmarks. This highlights a shift from the LLM era to the Agent era. Key innovations behind Vera Rubin's performance include extreme co-design, separation of services, distributed KV cache, KV-aware routing, MegaMoE architecture, 4-bit NVFP4 quantization, and 6th-gen NVLink for efficient MoE model execution. Simultaneously, NVIDIA announced the full-scale production of two new chips: 1. **Groq 3 LPX:** A low-latency inference accelerator designed for the Vera Rubin platform. When paired with Rubin GPUs for context processing, it achieves record-breaking output speeds—e.g., running Gemma 4 31B at 3,400 tokens/sec—drastically reducing multi-step agent task times. 2. **Vera CPU:** A processor specifically built for agentic AI, featuring 88 custom Olympus cores and 1.2TB/s memory bandwidth to handle the complex orchestration of agent tasks. SpaceXAI is already deploying it, with plans for space-based Vera Rubin systems by 2028. NVIDIA's strategy has evolved from selling GPUs to providing a complete, optimized "AI factory" stack—encompassing GPU, CPU, and specialized accelerators—to power the emerging Agent AI economy.

marsbitHace 9 hora(s)

NVIDIA Stuns with First Vera Rubin Test, DeepSeek Throughput Skyrockets 30x

marsbitHace 9 hora(s)

Arthur Hayes Makes a High-Profile Comeback, Flop Labs Aims to Become the "Fuel" for the Agent Economy

Arthur Hayes, co-founder of BitMEX, has announced his return as head of a new project called Flop Labs, declaring "FLOP is food for your AI agent." This marks his first high-profile return to a leadership role since stepping back from BitMEX. Flop Labs aims to be a native monetary network and verifiable computation settlement layer for the AI Agent economy, distinguishing itself from existing AI Agent projects. Its core concept is "Proof-of-Useful-Inference," which seeks to integrate useful AI inference tasks into the blockchain consensus mechanism, allowing FLOP tokens to serve as the native currency for Agents to purchase computing resources and store memory. The ecosystem involves four key roles: Miners (provide GPU compute), Validators (verify services), Agents (consume services), and KOLs/Partners (community growth). The project promises a 100% fair launch with no presale or VC involvement. A large airdrop is planned for Q4 2026, with the genesis block targeted for Q1 2027. While the announcement has generated community excitement, particularly around the airdrop, the project is in an early conceptual stage. It currently lacks a published whitepaper, detailed tokenomics, technical specifics on its consensus mechanism, or smart contract audits. Arthur Hayes has defended the venture, separating his belief in the potential of the Agentic Economy from concerns about an AI stock bubble. The success of Flop Labs will hinge on its ability to deliver substantive technical progress and ecosystem development in the coming months.

Odaily星球日报08/19 05:51

Arthur Hayes Makes a High-Profile Comeback, Flop Labs Aims to Become the "Fuel" for the Agent Economy

Odaily星球日报08/19 05:51

Overnight, GPT-5.6 Sol Was Accelerated 14x by OpenAI

OpenAI, in collaboration with chipmaker Cerebras, has unveiled a limited preview of an "Ultrafast Mode" for its flagship GPT-5.6 Sol model. This new service tier reportedly achieves output speeds of up to 750 tokens per second—a 14x increase over the standard mode's baseline of ~53 tokens/s—without any loss in quality. Key to this acceleration is Cerebras's wafer-scale architecture (WSE-3), which houses model parameters entirely in on-chip SRAM to eliminate the memory bandwidth bottlenecks typical of traditional GPU clusters. In benchmark testing on the challenging "Humanity's Last Exam" (HLE), GPT-5.6 Sol in Ultrafast Mode answered all 2500 questions in 11 hours and 11 minutes, compared to over 78 hours for a competitor model, while maintaining similar accuracy. The speed boost also translated to a 5.6x faster end-to-end performance on the GDP-Val benchmark for economically valuable knowledge work. OpenAI highlights several potential applications for such rapid inference, including real-time event response and reliability analysis, dynamic financial research and security, complex customer support, interactive shopping assistance, and accelerated research and experimentation workflows that enable multiple iterative cycles within a single workday. This advancement may allow users to deploy the highest-tier models for tasks previously requiring slower secondary models, significantly compressing multi-step agent workflows from hours to minutes.

marsbit08/14 00:02

Overnight, GPT-5.6 Sol Was Accelerated 14x by OpenAI

marsbit08/14 00:02

AMD Buys Taalas: Hardware AI Manages Without the Scarce HBM Memory

AMD has agreed to acquire Toronto-based startup Taalas, which addresses a major bottleneck in AI inference: the need to constantly transfer a neural network's model weights from memory to the processor for each token generated. Taalas's chips eliminate this operation by embedding the model weights directly into the transistors themselves. This data transfer is what currently limits inference speed and has made high-bandwidth memory (HBM) a critically scarce resource. Taalas, founded in 2023, has developed application-specific integrated circuits (ASICs). Its first test chip, fabricated on TSMC's 6nm process, reportedly ran Meta's Llama 3.1 8B model at speeds 48 times faster than Nvidia GPUs. The architecture features a mask ROM section for permanently stored weights and SRAM for adaptable components. AMD plans to integrate these chips into its Helios racks alongside its Instinct accelerators. However, this approach comes with a significant trade-off: each chip is permanently hardwired for a single model. Switching models requires a partial chip redesign, a process taking about two months even with Taalas's accelerated method. This limits its applicability to stable, widely-used models. The acquisition highlights a broader challenge in the semiconductor industry: the current memory shortage. HBM is sold out through 2026, and DRAM prices have surged. Yet, Taalas's technology demonstrates that this memory bottleneck is an engineering challenge, not an absolute physical limit. The industry is actively working on solutions, from Nvidia's model compression to Samsung's zHBM and new high-speed flash memory standards, all aimed at reducing reliance on scarce HBM. From an investment perspective, the deal challenges the assumption that AI-driven memory demand will keep prices permanently high. It serves as a reminder that memory has historically been a cyclical business, and current high prices are funding the very innovations designed to reduce future demand.

cryptonews.ru08/09 20:02

AMD Buys Taalas: Hardware AI Manages Without the Scarce HBM Memory

cryptonews.ru08/09 20:02

AMD acquires Taalas: hardware AI manages without scarce HBM memory

AMD has agreed to acquire Toronto-based startup Taalas, which tackles a key bottleneck in AI inference: the constant need to transfer model weights from memory to the processor for each generated token. Taalas's chips eliminate this operation by permanently embedding the model weights into the transistors themselves. This data transfer is what currently limits inference speed and has made high-bandwidth memory (HBM) a scarce commodity. Taalas's first test chip, fabricated on TSMC's 6nm process, reportedly generated tokens for Meta's Llama 3.1 8B model at speeds 48 times faster than comparable Nvidia GPUs. Its architecture features a mask ROM section for fixed weights and SRAM for adaptable components. However, this design comes with a significant trade-off: each chip is permanently dedicated to a single model. Switching models requires a partial redesign and fabrication, a process taking about two months. While the acquisition is seen as part of AMD's rivalry with Nvidia in inference, its broader implication lies in challenging the assumption of a permanent HBM memory shortage. The AI memory market is currently booming, with HBM supply sold out through 2026. Yet, Taalas's technology demonstrates that the memory bottleneck is an engineering challenge, not an absolute physical constraint. This aligns with industry-wide efforts from companies like Nvidia (through model compression) and memory makers like Samsung and SK hynix (developing new packaging and storage technologies) to reduce dependency on scarce HBM. AMD's move suggests that the current high prices for memory, driven by AI demand, may not be sustainable. It highlights a growing engineering push against the premise of perpetual memory scarcity, reminding investors that memory has historically been a cyclical business.

cryptonews.ru08/09 14:56

AMD acquires Taalas: hardware AI manages without scarce HBM memory

cryptonews.ru08/09 14:56

活动图片