Auto Research Era: 47 Tasks Without Standard Answers Become the Must-Test Leaderboard for Agent Capabilities

marsbitXuất bản vào 2026-05-13Cập nhật gần nhất vào 2026-05-13

Tóm tắt

The article introduces Frontier-Eng Bench, a new benchmark for AI agents developed by Einsia AI's Navers lab. Unlike traditional tests with clear answers, this benchmark presents 47 complex, real-world engineering tasks—such as optimizing underwater robot stability, battery fast-charging protocols, or quantum circuit noise control—where there is no single correct solution, only continuous optimization towards a limit. It shifts AI evaluation from static knowledge retrieval to a dynamic "engineering closed-loop": the AI must propose solutions, run simulations, interpret errors, adjust parameters, and re-run experiments to iteratively improve performance. This process tests an agent's ability to learn and evolve through long-term feedback, much like a human engineer tackling trade-offs between power, safety, and performance. Key findings from the benchmark reveal two patterns: 1) Improvements follow a power-law decay, becoming harder and smaller as optimization progresses, and 2) While exploring multiple solution paths (breadth) helps, sustained depth in a single path is crucial for breakthrough innovations. The research suggests this marks a step toward "Auto Research," where AI systems can autonomously conduct continuous, tireless optimization in scientific and engineering domains. Humans would set high-level goals, while AI agents handle the iterative experimentation and refinement. This could fundamentally change research and development workflows.

If we throw AI into an engineering site with no standard answers, can it still survive?

For a long time, AI Agents have appeared omnipotent, but in reality, most are just 'flipping through memories' within known knowledge bases.

Yet the real engineering world is harsh: the stability of underwater robots, the lithium plating boundary of power batteries, the noise control of quantum circuits... These problems have no 'perfect score', only 'optimizations that inch closer to the limit'.

Recently, the Agent Benchmark released by Navers lab under Einsia AIFrontier-Eng Bench—officially tore off the label of AI being an 'exam-crammer'.

The research team didn't have AI grind through outdated coding problems. Instead, they gave it a complete 'engineering closed loop': propose a solution, connect to the simulator, digest errors, adjust parameters, and re-run.

Faced with 47 hardcore tasks spanning multiple disciplines, AI must behave like a senior engineer, seeking the optimal solution within the 'impossible triangle' of power consumption, safety, and performance.

This is not just a test suite; it's more like a rehearsal for Agent 'evolution'.

When AI begins to learn self-correction from feedback, the Auto Research era, where 'humans set goals and AI iterates non-stop 24/7', might be closer than we imagine.

AI Starts Tackling 'Hard Work'

Past large language models were more like super straight-A students.

You pose a question, it 'flips through memory' from massive training data, then pieces together an answer that seems plausible.

In this mode, the large model is essentially playing 'word chain', not solving real-world problems.

But the emergence of Frontier-Eng Bench has AI doing the work of 'engineering optimization'.

The process has shifted to letting AI first propose a solution, then connect to a simulator to run experiments, subsequently obtain feedback and errors, modify parameters and code, and continue re-running until performance improves further.

In this closed-loop system, AI's identity undergoes a qualitative change.

Want to make the underwater robot more stable? AI must start automatically tuning the controller.

Want to increase the speed of the robotic arm a bit more? AI has to run simulations itself.

To some extent, AIs have shed their purely semantic understanding role and begun to act like professional engineers, continuously optimizing based on real-world environmental feedback.

The most interesting aspect of Frontier-Eng Bench is: it doesn't test whether AI 'answered correctly', but rather whether AI can continuously become stronger.

Because real engineering optimization is never about multiple-choice questions; there is no single standard answer.

Take fast-charging batteries as an example: the goal sounds simple—charge as fast as possible, but reality isn't so easy.

Under strict constraints like temperature mustn't spike, voltage can't overspeed, battery life can't drop too fast, and lithium plating must be avoided, AI must precisely hit the balance point of performance.

This means AI cannot pass through by any clever 'test-cramming' tricks; it must demonstrate endurance for continuous evolution through long-term feedback.

Can AI perform long-term optimization in real environments?

Looking at the results, GPT5.4 showed the most stable overall performance, but AIs still have a long way to go before 'solving' the Benchmark.

Auto Research Enters the 'Iterative Optimization' Era

The research team raised a very interesting point in their paper:

Truly advanced intelligence essentially relies on long-term feedback loops.

Just as AlphaGo could defeat Lee Sedol, it lay in the vast number of simulations and immediate feedback behind each decision, not the rote memorization of established game records.

True scientific research is the same: top labs don't rely on a single burst of inspiration, but continuously propose hypotheses, run experiments, examine results, modify plans, and try again.

Engineering optimization follows the same principle: anyone can create the first version; what's truly difficult is that final 1% performance leap.

The significance of Frontier-Eng Bench lies here: For the first time, it systematically begins testing AI's 'iterative optimization capability', and has summarized two nearly brutal laws of AI evolution.

The first law is: The further you go, the harder the improvement.

This paper found that the frequency and magnitude of Agent improvements follow a power-law decay:

  • Improvement frequency ∝ 1 / iteration count
  • Improvement magnitude ∝ 1 / improvement count

Simply put: the fastest gains come in the first few rounds, and it gets progressively harder and smaller later on.

This closely resembles the real R&D process: the first version of AI can quickly eliminate many 'low-hanging fruits', but the closer it gets to the bottleneck, the more effort is required to squeeze out even a bit more performance.

Would it be more cost-effective to explore multiple paths in parallel for trial and error? The answer lies in the second law.

The second law: Breadth is useful, but depth is even more indispensable.

Running multiple parallel paths can avoid getting stuck, but with a fixed budget, each additional chain opened shallows the depth of exploration.

Many engineering breakthroughs require continuous accumulation and constant correction before structural leaps emerge; they can't be achieved simply by 'trying a few more times'.

This actually points towards the development direction of next-generation Agents: not models that 'output an answer once', but systems that can continuously iterate and self-evolve within long-term feedback loops.

AI Engineers Might Really Be Coming

The true far-reaching significance of this research lies in its preliminary outline of an AI system beginning to approach the real engineering cycle.

Imagine when AI connects to industrial software, simulation environments, CAD systems, chip design tools, scientific computing platforms...

A dramatic transformation in the modality of productivity is on the verge of emerging.

In future labs, a division of labor like this might appear:

Human researchers are responsible for proposing directions and goals.

For example, 'reduce this component's energy consumption by 30%', 'compress this model's forward pass GPU usage even lower', 'increase the stability of robot control a bit more', 'push the fidelity of this quantum circuit closer to the limit', etc.

And AI is responsible for 'grinding the path'. They focus on these goals, continuously optimizing.

For example, automatically running simulations and experiments, automatically reading feedback from verifiers and simulators, then continuing to modify and optimize, iterating non-stop 24/7.

This evolutionary logic frees AI from the identity of an 'assistive tool', allowing it to begin solving complex system problems like a real engineering team—and tirelessly at that.

And the issues revealed by the Frontier-Eng Benchmark are actually very direct:

When AI begins to learn 'long-term optimization', how far is it from true engineering intelligence?

Paper Title: Frontier-Eng: Benchmarking Self-Evolving Agents on Real-World Engineering Tasks with Generative Optimization

Project Homepage: https://lab.einsia.ai/frontier-eng/

Arxiv: https://arxiv.org/abs/2604.12290

GitHub repo: https://github.com/EinsiaLab/Frontier-Engineering

This article is from the WeChat public account "Quantum Bit", author: Yun Zhong

Tiền kỹ thuật số thịnh hành

Câu hỏi Liên quan

QWhat is the main purpose of the Frontier-Eng Benchmark released by Einsteina AI's Navers lab?

AThe main purpose of the Frontier-Eng Benchmark is to move beyond testing AI's ability to recall known information. It systematically tests AI agents' capability for 'iterative optimization' on 47 real-world, open-ended engineering tasks without standard answers, evaluating if they can continuously improve performance through a feedback loop involving simulation, error analysis, and parameter adjustment.

QHow does the AI's role change in the Frontier-Eng Benchmark testing process compared to traditional language models?

AIn the Frontier-Eng Benchmark, the AI transitions from acting as a 'super student' that retrieves and assembles answers from training data to performing 'engineering optimization.' Its role becomes akin to a professional engineer: it proposes solutions, runs simulations, analyzes feedback and errors, modifies parameters/code, and reruns experiments in a continuous loop to seek optimal performance under complex constraints.

QWhat are the two key 'AI evolution laws' discovered through the Frontier-Eng Benchmark regarding iterative optimization?

AThe two key laws are: 1) Improvements become progressively harder and smaller (showing a power-law decay: Improvement frequency ∝ 1/iteration count, Improvement magnitude ∝ 1/improvement count). 2) While exploring multiple parallel paths (breadth) is useful, sustained depth in a single optimization path is more critical for achieving structural breakthroughs, as fixed budgets force a trade-off between breadth and depth.

QWhat future work paradigm does the article suggest might emerge from the development of self-evolving AI agents?

AThe article suggests a future 'Auto Research' paradigm where human researchers define the goals and direction (e.g., 'reduce component energy consumption by 30%'), and AI agents take on the role of 'grinding the path.' They would work autonomously and tirelessly—running simulations, interpreting feedback from verifiers and simulators, and iteratively optimizing—24/7 to approach performance limits.

QAccording to the article, what fundamental shift in AI capability does the Frontier-Eng Benchmark represent?

AThe Frontier-Eng Benchmark represents a fundamental shift from evaluating AI's ability to find predetermined 'correct answers' to testing its capacity for 'self-evolution' through long-term feedback loops. It moves the focus to whether AI can demonstrate sustained learning and improvement in complex, real-world scenarios with no single correct answer, pushing AI closer to genuine engineering intelligence.

Nội dung Liên quan

Đạo luật CLARITY 'vượt qua rào cản then chốt' – Nhà Trắng đồng ý thỏa thuận đạo đức về tiền mã hóa

Tổng thống Donald Trump đã đồng ý với một thỏa thuận về điều khoản đạo đức trong dự luật cấu trúc thị trường tiền mã hóa CLARITY, được cho là đã vượt qua trở ngại cuối cùng để thúc đẩy dự luật. Thỏa thuận về ngôn ngữ đạo đức này được White House và các Thượng nghị sĩ Cộng hòa đạt được, tuy nhiên chi tiết vẫn chưa được công bố và phe Dân chủ chưa được xem xét văn bản. Điều khoản đạo đức trở nên cấp thiết sau khi Trump kiếm được hơn 1,4 tỷ USD từ các dự án tiền mã hóa. Phe Dân chủ, dẫn đầu bởi Nghị sĩ Elizabeth Warren, nhấn mạnh rằng dự luật CLARITY phải kiểm soát được hành vi mà họ cho là 'tham nhũng' này. Để dự luật thông qua bỏ phiếu tại Thượng viện, phe Cộng hòa cần khoảng 7-10 phiếu từ phe Dân chủ để đạt ngưỡng 60 phiếu. Một số nhận định cho rằng thỏa thuận đạo đức có thể giúp giành được sự ủng hộ này. Trong bối cảnh toàn cầu, cuộc đua xây dựng khung pháp lý cho tiền mã hóa đang nóng lên khi nhiều khu vực như Châu Âu, Singapore, UAE, Hong Kong và cả Nga đã hoặc đang hoàn thiện khung pháp lý. Cố vấn tiền mã hóa của White House, Patrick Witt, đã hoãn huấn luyện quân sự để tiếp tục vận động cho dự luật, thúc giục Quốc hội Mỹ hành động nhanh chóng. Dù cập nhật về thỏa thuận đạo đức là tích cực, thị trường vẫn kém lạc quan về việc dự luật sẽ được thông qua trong năm 2026. Theo trang dự đoán Kalshi, xác suất thông qua trước năm 2027 chỉ là 38%, nhưng tăng lên 61% cho quý I/2027, cho thấy kỳ vọng dự luật sẽ được thông qua vào năm sau.

ambcrypto2 phút trước

Đạo luật CLARITY 'vượt qua rào cản then chốt' – Nhà Trắng đồng ý thỏa thuận đạo đức về tiền mã hóa

ambcrypto2 phút trước

Khoảnh Khắc Áp Lực Của Base

**Tóm tắt:** Bài viết phân tích những áp lực mà Base, blockchain L2 của Coinbase, đang đối mặt sau sự xuất hiện của Robinhood Chain. Jesse Pollak, đồng sáng lập Base, hai lần công khai thừa nhận sai lầm chiến lược: một là việc đặt cược vào token xã hội và người sáng tạo không mang lại hiệu quả bền vững; hai là Base đang tụt hậu trong lĩnh vực token hóa tài sản truyền thống (RWA) so với Robinhood. Base dự định phát hành cổ phiếu được token hóa với tỷ lệ hỗ trợ 1:1, khác với mô hình phái sinh của Robinhood. Mặc dù Base thành công và có TVL cao nhất trong các L2, nhưng nó vẫn tồn tại vấn đề lớn về mức độ phi tập trung. Việc L2BEAT có thể hạ cấp đánh giá phi tập trung của Base từ Stage 1 xuống Stage 0 và hai sự cố gián đoạn tạo khối trước đây đã làm dấy lên lo ngại. Sự cạnh tranh từ Robinhood Chain – với khối lượng giao dịch DEX tăng vọt ngay sau khi ra mắt – khiến những vấn đề này trở nên cấp bách hơn. Hành động của những người sáng lập cũng tạo ra sự tương phản: Trong khi CEO Robinhood Vlad Tenev tích cực tương tác với hệ sinh thái mới, thì việc Brian Armstrong (Coinbase) đổi avatar gây ảnh hưởng đến giá meme token BRAIN và phản hồi sau đó đã vấp phải phản ứng tiêu cực từ cộng đồng. Bài viết kết luận rằng Base cần tận dụng áp lực hiện tại để giải quyết triệt để các vấn đề kỹ thuật và xây dựng niềm tin, vì lợi thế của họ trong lĩnh vực token hóa RWA là không đủ lớn và cạnh tranh sẽ ngày càng khốc liệt.

Foresight News17 phút trước

Khoảnh Khắc Áp Lực Của Base

Foresight News17 phút trước

Nhà Trắng Nhượng Bộ, Dọn Đường Cho Rào Cản Đạo Đức, Liệu Đạo Luật Clarity Có Kịp Cửa Sổ Thời Gian Cuối Trước Kỳ Nghỉ?

Vào ngày 21 tháng 7, nhiều nguồn tin trong ngành tiết lộ chính quyền Trump đã đồng ý thêm điều khoản đạo đức vào "Đạo luật Minh bạch" (Clarity Act), một động thái được kỳ vọng sẽ dọn đường cho việc cập nhật văn bản và bỏ phiếu tại Thượng viện Mỹ. Đây được coi là rào cản cuối cùng cần vượt qua sau khi các bên đã đạt được thỏa hiệp về các vấn đề như lợi nhuận từ stablecoin và quy định cho DeFi. Một tín hiệu tích cực khác là cố vấn Patrick Witt của Nhà Trắng sẽ ở lại để hỗ trợ giai đoạn then chốt của dự luật. Tuy nhiên, thời gian rất hạn hẹp, vì Quốc hội Mỹ sắp bước vào kỳ nghỉ hè giữa tháng 8, chỉ còn khoảng hơn chục ngày làm việc. "Đạo luật Minh bạch" nhằm thiết lập khuôn khổ pháp lý liên bang thống nhất cho thị trường tài sản số tại Mỹ, làm rõ cách phân loại tài sản và phân định quyền giám sát giữa Ủy ban Chứng khoán (SEC) và Ủy ban Giao dịch Hàng hóa Tương lai (CFTC). Nếu được thông qua, đạo luật này có thể chấm dứt tình trạng mơ hồ về quy định kéo dài nhiều năm, tạo nền tảng rõ ràng hơn cho các tổ chức truyền thống tham gia thị trường. Các chuyên gia nhận định, nếu vượt qua được trở ngại cuối cùng về điều khoản đạo đức và được thông qua trước kỳ nghỉ hè, "Đạo luật Minh bạch" có thể đánh dấu một bước ngoặt lịch sử, đưa thị trường tài sản số toàn cầu tiến tới một giai đoạn phát triển chín chắn và có hệ thống hơn.

Odaily星球日报22 phút trước

Nhà Trắng Nhượng Bộ, Dọn Đường Cho Rào Cản Đạo Đức, Liệu Đạo Luật Clarity Có Kịp Cửa Sổ Thời Gian Cuối Trước Kỳ Nghỉ?

Odaily星球日报22 phút trước

Chỉ Báo Thị Trường Tiền Ảo & Cổ Phiếu | Strategy Tăng Dự Trữ Tiền Mặt Lên 3,23 Tỷ USD, Tạm Dừng Mua BTC; Các Định Chế Quản Lý Tài Sản Như Vanguard Tăng Sở Hữu Cổ Phiếu Strategy (21/7)

Tín hiệu thị trường tiền điện tử và chứng khoán (21/7): Strategy tăng dự trữ tiền mặt lên 3,23 tỷ USD, tạm dừng mua BTC; Các tổ chức như Vanguard tăng mua cổ phiếu Strategy. Thị trường chứng khoán toàn cầu đang trong giai đoạn điều chỉnh mạnh. Chỉ số S&P 500 ghi nhận mức đặt cược giảm (short selling) cao kỷ lục kể từ năm 2010, phản ánh lo ngại về bong bóng AI và kỳ vọng lãi suất tăng. Cổ phiếu công nghệ, đặc biệt là bán dẫn, dẫn đầu đà giảm. Về phía các công ty nắm giữ tiền điện tử: - **Strategy (MSTR)**: Dự trữ tiền mặt tăng lên 3,23 tỷ USD và tạm ngừng mua Bitcoin trong tuần. Nhiều quỹ lớn như Vanguard, Swedbank AB và Capital Group đã tăng mua cổ phiếu của công ty. - Toàn cầu: Các công ty đại chúng (không bao gồm công ty khai thác) chỉ mua ròng 1,33 triệu USD BTC trong tuần. Tổng nắm giữ đạt 1,139,656 BTC, trị giá ~73,76 tỷ USD. - **BitMine**: Tăng nắm giữ ETH lên 5,78 triệu coin, đạt ~4,8% nguồn cung lưu hành. Tổng tài sản (tiền điện tử, tiền mặt, chứng khoán) đạt 11,5 tỷ USD. - Các công ty khác: ORANGE JUICE huy động 40 triệu USD để mua BTC; Bitcoin Japan Corporation huy động 60 triệu USD và lần đầu phân bổ 4,08 triệu USD cho BTC. Thị trường tiền điện tử và chứng khoán vẫn chịu áp lực từ lạm phát, chính sách tiền tệ thắt chắt và các yếu tố địa chính trị. Nhà đầu tư được khuyến cáo thận trọng.

marsbit33 phút trước

Chỉ Báo Thị Trường Tiền Ảo & Cổ Phiếu | Strategy Tăng Dự Trữ Tiền Mặt Lên 3,23 Tỷ USD, Tạm Dừng Mua BTC; Các Định Chế Quản Lý Tài Sản Như Vanguard Tăng Sở Hữu Cổ Phiếu Strategy (21/7)

marsbit33 phút trước

Giao dịch

Giao ngay

Bài viết Nổi bật

Làm thế nào để Mua ERA

Chào mừng bạn đến với HTX.com! Chúng tôi đã làm cho mua Caldera (ERA) trở nên đơn giản và thuận tiện. Làm theo hướng dẫn từng bước của chúng tôi để bắt đầu hành trình tiền kỹ thuật số của bạn.Bước 1: Tạo Tài khoản HTX của BạnSử dụng email hoặc số điện thoại của bạn để đăng ký tài khoản miễn phí trên HTX. Trải nghiệm hành trình đăng ký không rắc rối và mở khóa tất cả tính năng. Nhận Tài khoản của tôiBước 2: Truy cập Mua Crypto và Chọn Phương thức Thanh toán của BạnThẻ Tín dụng/Ghi nợ: Sử dụng Visa hoặc Mastercard của bạn để mua Caldera (ERA) ngay lập tức.Số dư: Sử dụng tiền từ số dư tài khoản HTX của bạn để giao dịch liền mạch.Bên thứ ba: Chúng tôi đã thêm những phương thức thanh toán phổ biến như Google Pay và Apple Pay để nâng cao sự tiện lợi.P2P: Giao dịch trực tiếp với người dùng khác trên HTX.Thị trường mua bán phi tập trung (OTC): Chúng tôi cung cấp những dịch vụ được thiết kế riêng và tỷ giá hối đoái cạnh tranh cho nhà giao dịch.Bước 3: Lưu trữ Caldera (ERA) của BạnSau khi mua Caldera (ERA), lưu trữ trong tài khoản HTX của bạn. Ngoài ra, bạn có thể gửi đi nơi khác qua chuyển khoản blockchain hoặc sử dụng để giao dịch những tiền kỹ thuật số khác.Bước 4: Giao dịch Caldera (ERA)Giao dịch Caldera (ERA) dễ dàng trên thị trường giao ngay của HTX. Chỉ cần truy cập vào tài khoản của bạn, chọn cặp giao dịch, thực hiện giao dịch và theo dõi trong thời gian thực. Chúng tôi cung cấp trải nghiệm thân thiện với người dùng cho cả người mới bắt đầu và người giao dịch dày dạn kinh nghiệm.

Tổng lượt xem 813Xuất bản vào 2025.07.17Cập nhật vào 2026.06.02

Làm thế nào để Mua ERA

Thảo luận

Chào mừng đến với Cộng đồng HTX. Tại đây, bạn có thể được thông báo về những phát triển nền tảng mới nhất và có quyền truy cập vào thông tin chuyên sâu về thị trường. Ý kiến ​​của người dùng về giá của ERA (ERA) được trình bày dưới đây.

活动图片