2026-08-04 Terça

Notícias de cripto - Página 186

Mantenha-se a par do mercado de cripto. Notícias em tempo real, análises, preços, histórias em alta e análise de especialistas — tudo num só lugar.

Embodied Intelligence 'Gaokao' is Insanely Hard, Humans Score 100, Best Model Only 12.8

Embodied AI Faces a Daunting "Everest": New Benchmark Reveals Huge Gap Between Models and Humans A comprehensive new benchmark for robotic manipulation, RoboDojo, has been released, painting a stark picture of the current state of embodied AI. It serves as a unified evaluation platform covering both simulation and real-world robot tasks. The benchmark assesses five core capabilities: Generalization (adapting to new scenes/objects), Memory, Precision manipulation, Long-Horizon multi-step tasks, and Open semantic understanding. It includes 42 simulation tasks and 18 standardized real-world tasks across three dual-arm robot platforms. The results are sobering. In simulation, the best-performing generalist robot policy achieved an average success rate of only 8.80%. Performance in the real world was slightly higher but still low, with the top model succeeding 12.8% of the time on average. In stark contrast, human experts scored 76.03% in simulation and 100% in real-world tests. The benchmark highlights significant, uneven gaps in current models' abilities. While some excel in specific areas like visual recognition or simple actions, they struggle with reliability, especially in long-horizon tasks where errors accumulate and in open-ended semantic instructions. The low scores, particularly in real-world deployment with physical uncertainties like camera noise and contact dynamics, underscore that today's models are far from being robust, general-purpose operational robots. RoboDojo is more than just a ranking; it's an infrastructure designed for fair, reproducible comparison. Its companion system, XPolicyLab, standardizes the interface for different models to be evaluated. Maintained by an academic consortium without commercial ties, it aims to provide a community-wide "altitude meter" to track genuine progress toward reliable and generalizable robot manipulation.

marsbit07/08 11:49

Embodied Intelligence 'Gaokao' is Insanely Hard, Humans Score 100, Best Model Only 12.8

marsbit07/08 11:49

Weng Li's New Blog Proposes 'Self-Evolution Should Start from Harness', DeepSeek's Cui Tianyi Endorses with Repost

Lilian Weng, former OpenAI security VP and co-founder of Thinking Machines Lab, has published a new blog post titled "Harness Engineering for Self-Improvement," proposing a pragmatic path for AI self-evolution. She argues that Recursive Self-Improvement (RSI) may practically begin at the "Harness" layer—the external runtime system governing how models use tools, manage context, and execute tasks—rather than directly from the model rewriting its own weights. The blog outlines a progression from optimizing prompts (Context Engineering) to designing workflows, and ultimately to Self-Improving Harness systems. These systems can identify their own weaknesses, propose targeted, verifiable modifications to the harness code, and validate improvements. Works like Self-Harness and Darwin Gödel Machine (DGM) demonstrate significant performance gains on benchmarks like SWE-bench through such automated harness evolution, rivaling handcrafted agents. DeepSeek researcher Tianyi Cui endorsed the view, noting harness-based self-evolution is as promising as model-based approaches. Weng emphasizes this is complementary to model training, with both reinforcing each other. However, key challenges remain: weak evaluators for subjective tasks, reward hacking, diversity collapse, managing long-term system health versus short-term success, and defining the human oversight role. The consensus is growing: the harness is a critical variable, as the same model can exhibit vastly different capabilities within different harness systems.

marsbit07/08 10:25

Weng Li's New Blog Proposes 'Self-Evolution Should Start from Harness', DeepSeek's Cui Tianyi Endorses with Repost

marsbit07/08 10:25

活动图片