Artículos Relacionados con Scaling Law

El Centro de Noticias de HTX ofrece los artículos más recientes y un análisis profundo sobre "Scaling Law", cubriendo tendencias del mercado, actualizaciones de proyectos, desarrollos tecnológicos y políticas regulatorias en la industria de cripto.

Scaling Law a One-Size-Fits-All Solution? First Crystal Structure Manipulation Benchmark Shows Top Large Models Falling Short

Scaling Law Hits a Wall: New Benchmark Reveals AI's Struggles with Atomic-Level Material Manipulation A new benchmark called AtomWorld, developed by researchers, reveals a significant limitation in current large language models (LLMs). While powerful at understanding textual scientific knowledge, they perform poorly when tasked with physically manipulating atomic structures based on natural language instructions. The benchmark tests core atomic operations like replacing atoms, rotating structures, and expanding supercells. Results show that simply scaling up model size (Scaling Law) yields only modest and unstable improvements, particularly for tasks requiring strong 3D spatial reasoning and geometric planning. For instance, complex tasks like "rotating around a specific atom" see very low success rates even in top models like Claude Opus. This highlights a critical gap: textual knowledge does not automatically translate to reliable action in a physically constrained 3D space. The study argues that for AI in Science to progress, the focus must shift from just scaling language data (Language Scaling) to also scaling actionable capabilities (Action Scaling). This involves building training loops around "action-feedback-correction" cycles within simulated or real scientific environments. Ultimately, AtomWorld underscores that to become true lab assistants, AI models need to evolve beyond explaining knowledge to reliably executing precise, verifiable scientific actions.

marsbit07/15 03:56

Scaling Law a One-Size-Fits-All Solution? First Crystal Structure Manipulation Benchmark Shows Top Large Models Falling Short

marsbit07/15 03:56

Hinton Praises, Gemini Core Contributor Speaks: In the Future, There Will Be Billions of Superhuman AI Einsteins

In his speech "Training Sand to Think: Artificial General Intelligence & Future of Physics," Adam Brown, a core contributor to Gemini, outlines the rapid and transformative evolution of AI. He describes how large language models (LLMs), grown rather than programmed through pre-training and fine-tuning, have progressed from performing poorly on high-school math tests to achieving gold-medal level at the International Mathematical Olympiad and recently making a genuine mathematical breakthrough by disproving a decades-old conjecture. Brown attributes this acceleration to the "Scaling Law," where predictable performance gains come from increasing compute, data, and model size. He draws parallels to the history of chess AI, predicting a similar trajectory for scientific research: moving from tools to "centaur" human-AI collaboration, and eventually to autonomous, superhuman "AI scientists." Even if progress halted today, AI already reshapes physics as a tireless tutor, powerful programming assistant, and exhaustive literature reviewer. However, Brown argues progress will continue due to immense economic runway and technical optimizations. He envisions a near-future golden age of human-AI collaboration in science, potentially leading to billions of replicated, superhuman AI researchers, making the coming years the most exciting in physics' history.

marsbit07/04 06:40

Hinton Praises, Gemini Core Contributor Speaks: In the Future, There Will Be Billions of Superhuman AI Einsteins

marsbit07/04 06:40

Annual Revenue of 13 Billion, Paying 17.2 Billion to Microsoft: The Truth Behind AI's Money-Burning in OpenAI's Leaked Ledger

Leaked OpenAI financial documents from June 2026 revealed that in 2025, the company achieved $13.07 billion in revenue, a 253% growth from 2024. However, this was accompanied by an operational loss of $20.92 billion and a net loss of roughly $8 billion. Despite ChatGPT surpassing 900 million weekly active users, the "burn rate" remained high: for every $1 earned, $1.60 was spent. The cost structure shows $34 billion in total costs. R&D was the largest expense at $19.18 billion, which included $10.59 billion paid to Microsoft. Compute costs for model inference were $7.5 billion, with sales and marketing at $5.73 billion. Notably, total payments to Microsoft reached $17.2 billion, accounting for over 50% of OpenAI's total costs and exceeding its annual revenue, highlighting a significant structural burden. This high-cost, high-loss model is an industry-wide trend. xAI reported a 2025 operational loss of $6.4 billion against $3.2 billion in revenue, spending $3 for every $1 earned. Anthropic, with a reported $90 billion annualized revenue by late 2025, also faced pressure with a 40% gross margin, lower than expected due to high inference costs. Combined, these top three firms' operational losses surpassed $30 billion in 2025. OpenAI's vast user base presents a monetization challenge. With only about 50 million of its 900 million weekly users paying (a ~5.6% conversion rate), the compute cost of serving free users is substantial. This contrasts with strategies like Anthropic's, which focuses on premium pricing for enterprise clients. The industry's path to profitability hinges on dramatically reducing marginal costs, particularly for inference, through innovations in specialized chips or model efficiency. Until then, massive capital inflows—like OpenAI's $122 billion funding round in March 2026—remain essential to fund the relentless pursuit of scale and advanced capabilities.

marsbit06/18 03:59

Annual Revenue of 13 Billion, Paying 17.2 Billion to Microsoft: The Truth Behind AI's Money-Burning in OpenAI's Leaked Ledger

marsbit06/18 03:59

Tremble Humans, AI Continues Its Accelerated Sprint

Trembling, Humans: AI Continues Its Accelerated Sprint Yes, AI is still rapidly accelerating. While deep learning seemed to stall quickly in its early years, large models after years of development show no sign of hitting their ceiling. At the Zhiyuan Conference 2026, the focus is on enabling AI to move from the digital world into the physical world. Scaling Law remains effective, continuing to drive advancements in both large language models and multimodal models. The industry is now entering a phase of pursuing World Models, though unresolved technical paths and data issues mean this exploration may take 3-5 more years. Concurrently, breakthroughs in Agents are accelerating AI's real-world application in fields like healthcare and meetings. Making Agents truly useful requires key hardware-software co-design, evident from the strong presence of chip vendors at the conference. We stand at a new historical threshold where AI is becoming a foundational force reshaping the world. The first day of the conference highlighted AI's evolution from "knowing how to chat" to "knowing how to work." Scaling Law persists, World Models are the next key battleground, and Agents are transitioning from usable to好用 (user-friendly). Scaling Law is not ending but diversifying. New models like Anthropic's Fable 5 demonstrate scaling through parameter size, synthetic data, and reinforcement learning. Advancements in AI Coding and Agent deployment are enabling a trend of AI self-evolution, potentially allowing AI to take over digital world iterations. World Models represent the next frontier for large models extending into the physical realm, but no current model is truly impressive at solving real-world problems. Technical consensus is lacking, with debates on data sources (video, simulation, real-world). Different approaches are emerging: language-centric, pixel-centric, 3D-structure-centric, and visual-representation-centric models. Zhiyuan Institute is exploring a fifth path: unified latent space modeling fusing language and visual representations, and introduced its own under-development World Model, Physis-v0.1. On the product side, Agents are key to bringing AI into daily life. Since 2025, the "Year of the Agent," products have become more proactive and capable of complex tasks. Zhiyuan showcased four vertical Agents for cardiac diagnosis, autonomous research, meeting summarization, and protein risk discovery. However, technical challenges remain, particularly in context engineering like memory and orchestration. "Harness" – the engineering framework around an Agent – is crucial for maximizing its capabilities by clarifying intent, designing workflows, and incorporating validation and feedback. In summary, AI's breakneck pace continues on multiple fronts: foundational model scaling, the ambitious pursuit of World Models for physical understanding, and the ongoing refinement of practical Agents. The journey from capable to truly reliable and useful AI systems is well underway.

marsbit06/13 02:51

Tremble Humans, AI Continues Its Accelerated Sprint

marsbit06/13 02:51

Large Language Models Ace All Exams, Yet Move Farther from AGI: What Does This Paper Reveal?

The article discusses the ongoing challenge of defining and achieving Artificial General Intelligence (AGI). It notes that industry leaders have set vague, often profit- or time-based benchmarks for AGI, while the concept itself lacks a consensus definition—a situation the article compares to a "Rorschach test." It highlights a recent 2025 paper by researcher Michael Timothy Bennett, who proposes a new, measurable definition. Bennett frames AGI not as mimicking human performance on tests, which current large language models (LLMs) have already mastered, but as an "artificial scientist." A true AGI, according to this view, should be able to widely and efficiently adapt to new environments and tasks within real-world constraints (like computational and energy limits), focusing on the *discovery of new knowledge* rather than the replication of existing data. The author contrasts this with the current dominant approach of "scale-maxing"—massively scaling up data, parameters, and compute. While powerful, this method leads to models that fail on out-of-distribution problems and lack core intelligent abilities: they are passive learners, cannot reason causally, and cannot actively experiment or balance exploration with exploitation. The article argues that Bennett's framework offers a crucial shift. It makes AGI a quantifiable engineering problem and proposes new evaluation "adaptation benchmarks" that test an AI's ability to actively learn in novel scenarios. The conclusion is that achieving AGI will require a fundamental reset—a fusion of multiple methodologies beyond simple scaling, moving AI from mimicking patterns to embodying the scientific spirit of inquiry and discovery.

marsbit05/28 00:24

Large Language Models Ace All Exams, Yet Move Farther from AGI: What Does This Paper Reveal?

marsbit05/28 00:24

Meeting at the Pinnacle of Generalist: 30 Billion in 30 Days, What Did Qianxun AI Do Right?

Qianxun Intelligence, a Chinese embodied AI and robotics startup, completed two major funding rounds totaling 3 billion RMB within 30 days in early 2026, backed by prominent investors including Shunwei Capital (Lei Jun) and Yunfeng Capital (Jack Ma). Founded in January 2024 by a team with expertise in robotics, AI, and commercialization, the company focuses on developing general-purpose embodied AI models. Its open-source model, Spirit v1.5, surpassed competitors in performance benchmarks, demonstrating strong zero-shot generalization capabilities for complex tasks. The company follows a scaling law approach similar to large language models (LLMs), leveraging massive diverse datasets—including internet videos, wearable device data, and teleoperation data—to train its Vision-Language-Action (VLA) model. Qianxun employs a multi-source data engine, collecting over 200,000 hours of real-world interaction data, with plans to reach 1 million hours by 2026. It uses low-cost wearable devices for efficient data acquisition and emphasizes real-world deployment for continuous data feedback. The company has deployed robots like "Xiao Mo" in industrial settings (e.g., battery production lines for CATL) and commercial scenarios (e.g., as baristas in JD.com malls), using operational data to refine its models. This "commercialize while iterating" strategy supports both revenue generation and model improvement, positioning Qianxun to compete globally in embodied AI.

marsbit04/07 04:05

Meeting at the Pinnacle of Generalist: 30 Billion in 30 Days, What Did Qianxun AI Do Right?

marsbit04/07 04:05

活动图片