Artículos Relacionados con CUDA

El Centro de Noticias de HTX ofrece los artículos más recientes y un análisis profundo sobre "CUDA", cubriendo tendencias del mercado, actualizaciones de proyectos, desarrollos tecnológicos y políticas regulatorias en la industria de cripto.

NVIDIA's 20-Year CUDA Moat Collapsed Over a Weekend, Claude Single-Handedly Got AMD's New GPU Running

In a single weekend, Claude, an AI agent from Anthropic, successfully ported and optimized its cutting-edge model to run on a brand-new AMD MI355X server rack without any manual code intervention. This feat demonstrates a potential breakthrough in overcoming NVIDIA's long-established CUDA software ecosystem dominance, built over two decades. Anthropic's team simply instructed Claude to get the AMD machine running. By Monday, it not only worked but was showing a continuously improving performance curve. The achievement impressed AMD CEO Lisa Su and accelerated a major deployment partnership: Anthropic plans to deploy up to 2GW of AMD Instinct GPUs starting in 2027. The key enabler is AMD's new ROCm.AI platform, a toolbox designed specifically for AI agents like Claude. It provides AI-readable documentation, including chip instruction sets (ISA), and tools like the Hyperloom service that allows agents to autonomously profile performance, identify bottlenecks, test configurations, and generate optimized kernels. In a demo, Hyperloom boosted the output speed of a model by 38%. This represents a fundamental shift. While CUDA's strength lies in its vast, human-expert-driven ecosystem of tools and tacit knowledge, AMD's strategy is to make its hardware and software stack directly accessible and optimizable by AI agents. An agent can parallelize tasks—debugging, profiling, coding—that would take human engineers years to master, compressing the traditional software adaptation timeline from years to tasks. The competition is no longer just about peak hardware specs but also about how well AI can read, utilize, and tune a platform.

marsbit07/28 00:09

NVIDIA's 20-Year CUDA Moat Collapsed Over a Weekend, Claude Single-Handedly Got AMD's New GPU Running

marsbit07/28 00:09

AMD Launches Compact AI Host, Directly Challenging NVIDIA DGX Spark

In June 2026, AMD announced the Ryzen AI Halo, a compact AI developer desktop to rival NVIDIA's DGX Spark. Both feature 128GB unified memory for running 200B+ parameter models locally. Priced from $2,949 to $3,999, AMD undercuts NVIDIA's $3,999+ DGX Spark. The core divergence lies in architecture and philosophy. Ryzen AI Halo uses an x86-based Ryzen AI Max+ 395 APU (CPU+GPU+NPU), runs standard Windows/Linux, and emphasizes general-purpose PC flexibility. DGX Spark uses an ARM-based Grace Blackwell Superchip, runs a custom DGX OS, and includes a high-speed ConnectX-7 NIC for cluster prototyping, anchoring it to NVIDIA's full-stack CUDA ecosystem. AMD's ROCm software has improved, with simpler installation and support for major frameworks, but still lags behind CUDA's 17-year maturity in community support and cutting-edge library availability. AMD's broader strategy focuses on becoming a viable second-source supplier. Key moves include acquiring design capabilities via ZT Systems (while outsourcing manufacturing) and securing two major 6GW GPU supply deals with OpenAI and Meta in late 2025/early 2026. These contracts validate AMD's role in diversifying the AI supply chain, rather than outright beating NVIDIA. NVIDIA counters with a tightly integrated stack from desktop (DGX Spark) to data center, emphasizing seamless scalability and enterprise software subscriptions (AI Enterprise). In summary, Ryzen AI Halo represents AMD's pragmatic path: offering a cost-effective, open-ecosystem alternative for developers wary of vendor lock-in, while its large data center contracts aim to capture share from customers seeking a second GPU supplier. The choice boils down to a familiar, flexible PC environment with potential software gaps (AMD) versus a premium, optimized, but locked-in ecosystem (NVIDIA).

marsbit06/16 09:14

AMD Launches Compact AI Host, Directly Challenging NVIDIA DGX Spark

marsbit06/16 09:14

It's Not Jensen Huang Who Wants to Change the PC, But the PC That's Revolting Against Itself

The 40-year-old PC industry is undergoing a fundamental transformation, driven by the rise of AI PCs. At the GTC Taipei 2026 event, NVIDIA, backed by Microsoft and major PC OEMs, announced the RTX Spark super chip for Windows PCs, marking its official entry into the PC core processor market. This move aims to redefine the AI PC by shifting its core from the CPU to an AI-focused SoC (System on Chip). NVIDIA envisions the PC evolving from a personal computer to a "personal AI"—a platform where local AI Agents can autonomously perform tasks. While Intel pioneered the AI PC concept earlier in 2026, NVIDIA's aggressive push, leveraging its vast CUDA developer ecosystem of 6 million, positions it to potentially reshape the industry's long-standing Wintel (Windows-Intel) power structure. NVIDIA's strategy extends beyond hardware; it's about embedding its CUDA, RTX, and AI software stack into the PC platform itself. The article identifies key shifts: 1) The move from a CPU-centric to an AI SoC-centric architecture, similar to Apple's approach with its M-series chips. 2) The PC's evolution from a human-operated tool to a platform for human-Agent collaboration. 3) The extension of NVIDIA's data center-centric CUDA ecosystem to personal devices via RTX Spark. Ultimately, the change is driven by the broader trend of AI moving to personal devices. Companies like Intel, AMD, Qualcomm, and Apple are all participating in this shift. NVIDIA's entry accelerates the competition, but the core driver is the technology itself finding its optimal expression in the PC. The industry is reinventing itself, with the outcome hinging on execution, ecosystem development, and the creation of compelling local AI applications.

marsbit06/12 11:14

It's Not Jensen Huang Who Wants to Change the PC, But the PC That's Revolting Against Itself

marsbit06/12 11:14

Autonomy or Compatibility: The Choice Facing China's AI Ecosystem Behind the Delay of DeepSeek V4

DeepSeek V4's repeated delay in early 2026 has sparked global discussions on "de-CUDA-ization" in AI. The highly anticipated trillion-parameter open-source model is undergoing deep adaptation to Huawei’s Ascend chips using the CANN framework, representing China’s first systematic attempt to run a core AI model outside the CUDA ecosystem. This shift, however, comes with significant engineering challenges. While the model uses a MoE architecture to reduce computational load, it places extreme demands on memory bandwidth, chip interconnects, and system scheduling—areas where NVIDIA’s mature CUDA ecosystem currently excels. Migrating to Ascend introduces complexities in hardware topology, communication latency, and software optimization due to CANN’s relative immaturity compared to CUDA. The move highlights a broader strategic dilemma: short-term compatibility with CUDA offers practical benefits and faster adoption, as seen in CANN’s efforts to emulate CUDA interfaces. Yet, long-term over-reliance on compatibility risks inheriting CUDA’s limitations and stifling native innovation. If global AI shifts away from transformer-based architectures, strict compatibility could lead to technological obsolescence. Despite these challenges, DeepSeek V4’s eventual release could demonstrate the viability of a full domestic AI stack and accelerate CANN’s ecosystem growth. However, true technological independence will require building an original software-hardware paradigm beyond compatibility—a critical task for China’s AI ambitions in the next 3-5 years.

marsbit04/21 10:16

Autonomy or Compatibility: The Choice Facing China's AI Ecosystem Behind the Delay of DeepSeek V4

marsbit04/21 10:16

NVIDIA's Market Share in China Drops Below 60%, Domestic AI Chips Seize Market with 1.65 Million Units Delivered Annually

Nvidia's market share in China's AI accelerator card market has declined significantly, dropping from approximately 95% to 55% in 2025, according to IDC data. During the same period, domestic Chinese manufacturers collectively captured 41% of the market, shipping 1.65 million units out of a total market of 4 million units. Huawei led the domestic suppliers with 812,000 units shipped, representing nearly half of the local market share. This shift is driven by both U.S. export controls and China’s aggressive domestic substitution policies. In November 2025, Beijing mandated that state-funded data centers must use domestic AI chips, accelerating the adoption of local alternatives. Huawei recently launched the Atlas 350 accelerator card, claiming 2.87 times the inference performance of Nvidia’s H20 in low-precision computing, though direct comparisons are complicated by architectural differences. While Chinese chips still lag behind in training large-scale AI models—estimated to be 5-10 years behind Nvidia—they have reached a "good enough" level for many commercial applications like inference tasks. The main challenge remains software ecosystem development, as Nvidia’s CUDA platform remains the industry standard. Chinese firms are responding with compatibility efforts and open-source initiatives. Several domestic AI chip companies are now pursuing IPOs, and Huawei continues heavy R&D spending to reduce foreign dependency. Even if U.S. export policies ease, the structural move toward domestic AI chips appears irreversible.

marsbit04/03 05:51

NVIDIA's Market Share in China Drops Below 60%, Domestic AI Chips Seize Market with 1.65 Million Units Delivered Annually

marsbit04/03 05:51

China's AI Computing Counterattack

Eight years after the ZTE crisis, China's AI industry is fighting back against U.S. chip restrictions. In 2018, ZTE nearly collapsed under U.S. sanctions but survived with heavy fines and oversight. Today, Chinese AI firms like DeepSeek are pivoting away from NVIDIA by developing domestic alternatives and optimizing algorithms to reduce reliance on foreign technology. DeepSeek’s V4 model will use entirely domestic chips, signaling a strategic shift toward computational independence. The real challenge isn’t just hardware—it’s NVIDIA’s CUDA ecosystem, which dominates global AI development with over 4.5 million developers. U.S. export controls have tightened since 2022, banning high-end chips like the A100, H100, and their downgraded versions. In response, Chinese companies are adopting technical workarounds like Mixture-of-Experts models, which activate only parts of the network during inference, slashing costs. DeepSeek’s API is up to 75x cheaper than competitors, driving rapid global adoption. By early 2026, Chinese models accounted for nearly 60% of API calls on OpenRouter. Domestic chips, such as Huawei’s Ascend series, are now capable of full-scale training, not just inference. Production lines in cities like Xinghua manufacture servers with homegrown processors, supporting major AI training projects. Meanwhile, the U.S. faces an electricity shortage as data centers consume growing power, while China benefits from greater energy capacity and lower costs. Chinese AI is also going global via “Token exports,” with services reaching users in India, Indonesia, and beyond. The situation echoes Japan’s semiconductor decline in the 1980s, but China is building an independent ecosystem rather than relying on global supply chains. Domestic chip firms report surging revenues but ongoing losses—reflecting the high cost of achieving true technological independence. The battle is difficult, but progress is underway.

marsbit03/04 05:09

China's AI Computing Counterattack

marsbit03/04 05:09

活动图片