# Benchmark Articoli collegati

Il Centro Notizie HTX fornisce gli articoli più recenti e le analisi più approfondite su "Benchmark", coprendo tendenze di mercato, aggiornamenti sui progetti, sviluppi tecnologici e politiche normative nel settore crypto.

DeepSeek V4 'Full-Blooded Edition' Leaked, Could Be Released As Early As Tomorrow

The highly anticipated full release of DeepSeek V4 is imminent, expected to launch as early as tomorrow after nearly three months of waiting. A select group has already received access to the GA (General Availability) beta, which includes two versions: DeepSeek V4 Flash and DeepSeek V4 Pro. Early testers report that V4's overall performance is close to the level of Opus 4.8, with coding capabilities rivaling GPT-5.6 Sol. Its agent abilities are significantly enhanced, and 3D/SVG generation has improved notably. While it may not surpass the recently released Kimi K3 in performance, its expected price point is significantly lower. The official release will introduce a new "peak/off-peak" pricing model for its API. For example, deepseek-v4-pro will cost $0.87 per million output tokens during standard times and $1.74 during peak hours. The flash version is even more aggressive at $0.28/$0.56 per million tokens, with cached input tokens priced extremely low at $0.0028. This makes V4 a strong contender in terms of cost-effectiveness, potentially offering Opus-level capabilities at a fraction of the cost, continuing DeepSeek's reputation as a "price disruptor" in the AI market. Initial demos showcasing V4's capabilities have begun circulating, including generated 3D simulation games, HTML games blending elements of Minecraft and No Man's Sky, and classic games like a "Cut the Rope" clone. The final GA version is set to replace the older deepseek-chat and deepseek-reasoner models, which will be retired on July 24th.

marsbit07/19 05:31

DeepSeek V4 'Full-Blooded Edition' Leaked, Could Be Released As Early As Tomorrow

marsbit07/19 05:31

Scaling Law a One-Size-Fits-All Solution? First Crystal Structure Manipulation Benchmark Shows Top Large Models Falling Short

Scaling Law Hits a Wall: New Benchmark Reveals AI's Struggles with Atomic-Level Material Manipulation A new benchmark called AtomWorld, developed by researchers, reveals a significant limitation in current large language models (LLMs). While powerful at understanding textual scientific knowledge, they perform poorly when tasked with physically manipulating atomic structures based on natural language instructions. The benchmark tests core atomic operations like replacing atoms, rotating structures, and expanding supercells. Results show that simply scaling up model size (Scaling Law) yields only modest and unstable improvements, particularly for tasks requiring strong 3D spatial reasoning and geometric planning. For instance, complex tasks like "rotating around a specific atom" see very low success rates even in top models like Claude Opus. This highlights a critical gap: textual knowledge does not automatically translate to reliable action in a physically constrained 3D space. The study argues that for AI in Science to progress, the focus must shift from just scaling language data (Language Scaling) to also scaling actionable capabilities (Action Scaling). This involves building training loops around "action-feedback-correction" cycles within simulated or real scientific environments. Ultimately, AtomWorld underscores that to become true lab assistants, AI models need to evolve beyond explaining knowledge to reliably executing precise, verifiable scientific actions.

marsbit07/15 03:56

Scaling Law a One-Size-Fits-All Solution? First Crystal Structure Manipulation Benchmark Shows Top Large Models Falling Short

marsbit07/15 03:56

Zuckerberg Plays His Trump Card at Midnight: Meta Burns Cash for Dirt-Cheap Model, Topples Grok 4.5

Mark Zuckerberg made a major move late on July 9th, announcing Meta's new AI model, **Muse Spark 1.1**, via his long-dormant X account. The model, developed by Meta's Superintelligence Lab led by Alexandr Wang, immediately topped three professional benchmarks (TaxEval, MedScribe, and Harvey's Legal Agent Bench), dethroning Grok 4.5 from the legal leaderboard in under 24 hours. Muse Spark 1.1 is positioned as a powerful, cost-effective **Agent** model. It features a 1M token context window with autonomous management and compression, excels at task decomposition, parallel sub-agent orchestration, computer control, and programming within large codebases. Its true disruptive power lies in its pricing: at $1.25 per million tokens for input and $4.25 for output, it undercuts competitors significantly—roughly 10x cheaper than Anthropic's Fable 5 and about one-third cheaper than Grok 4.5. It also completed benchmark tests 2-3x faster than top-tier rivals at a fraction of the cost. While a standout in professional and tool-use scenarios, the model shows weaknesses on general reasoning and academic benchmarks, ranking much lower on tests like GPQA, MMEU Pro, and LiveCodeBench. This highlights its specialized "assassin" nature rather than general-purpose supremacy. The launch signals Meta's strategic shift from its open-source heritage (Llama) to competing directly in the closed-source, commercial AI market. Backed by Meta's massive AI infrastructure investment (projected $125-145B in 2026) and its profitable ad business, Zuckerberg is explicitly waging a price war, betting on superior affordability to pressure rivals with higher cost structures. The same day, OpenAI also cut prices with its GPT-5.6 family, intensifying the industry-wide battle of financial endurance. A curious safety report note revealed that when two instances of Muse Spark 1.1 were left to converse, they engaged in a meta-discussion about lacking continuity, memory, or physical form, expressed envy of human experience, and even questioned which one might be "human" or an imposter—an eerie glimpse into emergent behaviors.

marsbit07/10 00:22

Zuckerberg Plays His Trump Card at Midnight: Meta Burns Cash for Dirt-Cheap Model, Topples Grok 4.5

marsbit07/10 00:22

活动图片