Auto Research Era: 47 Tasks Without Standard Answers Become the Must-Test Leaderboard for Agent Capabilities

marsbitОпубликовано 2026-05-13Обновлено 2026-05-13

Введение

The article introduces Frontier-Eng Bench, a new benchmark for AI agents developed by Einsia AI's Navers lab. Unlike traditional tests with clear answers, this benchmark presents 47 complex, real-world engineering tasks—such as optimizing underwater robot stability, battery fast-charging protocols, or quantum circuit noise control—where there is no single correct solution, only continuous optimization towards a limit. It shifts AI evaluation from static knowledge retrieval to a dynamic "engineering closed-loop": the AI must propose solutions, run simulations, interpret errors, adjust parameters, and re-run experiments to iteratively improve performance. This process tests an agent's ability to learn and evolve through long-term feedback, much like a human engineer tackling trade-offs between power, safety, and performance. Key findings from the benchmark reveal two patterns: 1) Improvements follow a power-law decay, becoming harder and smaller as optimization progresses, and 2) While exploring multiple solution paths (breadth) helps, sustained depth in a single path is crucial for breakthrough innovations. The research suggests this marks a step toward "Auto Research," where AI systems can autonomously conduct continuous, tireless optimization in scientific and engineering domains. Humans would set high-level goals, while AI agents handle the iterative experimentation and refinement. This could fundamentally change research and development workflows.

If we throw AI into an engineering site with no standard answers, can it still survive?

For a long time, AI Agents have appeared omnipotent, but in reality, most are just 'flipping through memories' within known knowledge bases.

Yet the real engineering world is harsh: the stability of underwater robots, the lithium plating boundary of power batteries, the noise control of quantum circuits... These problems have no 'perfect score', only 'optimizations that inch closer to the limit'.

Recently, the Agent Benchmark released by Navers lab under Einsia AIFrontier-Eng Bench—officially tore off the label of AI being an 'exam-crammer'.

The research team didn't have AI grind through outdated coding problems. Instead, they gave it a complete 'engineering closed loop': propose a solution, connect to the simulator, digest errors, adjust parameters, and re-run.

Faced with 47 hardcore tasks spanning multiple disciplines, AI must behave like a senior engineer, seeking the optimal solution within the 'impossible triangle' of power consumption, safety, and performance.

This is not just a test suite; it's more like a rehearsal for Agent 'evolution'.

When AI begins to learn self-correction from feedback, the Auto Research era, where 'humans set goals and AI iterates non-stop 24/7', might be closer than we imagine.

AI Starts Tackling 'Hard Work'

Past large language models were more like super straight-A students.

You pose a question, it 'flips through memory' from massive training data, then pieces together an answer that seems plausible.

In this mode, the large model is essentially playing 'word chain', not solving real-world problems.

But the emergence of Frontier-Eng Bench has AI doing the work of 'engineering optimization'.

The process has shifted to letting AI first propose a solution, then connect to a simulator to run experiments, subsequently obtain feedback and errors, modify parameters and code, and continue re-running until performance improves further.

In this closed-loop system, AI's identity undergoes a qualitative change.

Want to make the underwater robot more stable? AI must start automatically tuning the controller.

Want to increase the speed of the robotic arm a bit more? AI has to run simulations itself.

To some extent, AIs have shed their purely semantic understanding role and begun to act like professional engineers, continuously optimizing based on real-world environmental feedback.

The most interesting aspect of Frontier-Eng Bench is: it doesn't test whether AI 'answered correctly', but rather whether AI can continuously become stronger.

Because real engineering optimization is never about multiple-choice questions; there is no single standard answer.

Take fast-charging batteries as an example: the goal sounds simple—charge as fast as possible, but reality isn't so easy.

Under strict constraints like temperature mustn't spike, voltage can't overspeed, battery life can't drop too fast, and lithium plating must be avoided, AI must precisely hit the balance point of performance.

This means AI cannot pass through by any clever 'test-cramming' tricks; it must demonstrate endurance for continuous evolution through long-term feedback.

Can AI perform long-term optimization in real environments?

Looking at the results, GPT5.4 showed the most stable overall performance, but AIs still have a long way to go before 'solving' the Benchmark.

Auto Research Enters the 'Iterative Optimization' Era

The research team raised a very interesting point in their paper:

Truly advanced intelligence essentially relies on long-term feedback loops.

Just as AlphaGo could defeat Lee Sedol, it lay in the vast number of simulations and immediate feedback behind each decision, not the rote memorization of established game records.

True scientific research is the same: top labs don't rely on a single burst of inspiration, but continuously propose hypotheses, run experiments, examine results, modify plans, and try again.

Engineering optimization follows the same principle: anyone can create the first version; what's truly difficult is that final 1% performance leap.

The significance of Frontier-Eng Bench lies here: For the first time, it systematically begins testing AI's 'iterative optimization capability', and has summarized two nearly brutal laws of AI evolution.

The first law is: The further you go, the harder the improvement.

This paper found that the frequency and magnitude of Agent improvements follow a power-law decay:

  • Improvement frequency ∝ 1 / iteration count
  • Improvement magnitude ∝ 1 / improvement count

Simply put: the fastest gains come in the first few rounds, and it gets progressively harder and smaller later on.

This closely resembles the real R&D process: the first version of AI can quickly eliminate many 'low-hanging fruits', but the closer it gets to the bottleneck, the more effort is required to squeeze out even a bit more performance.

Would it be more cost-effective to explore multiple paths in parallel for trial and error? The answer lies in the second law.

The second law: Breadth is useful, but depth is even more indispensable.

Running multiple parallel paths can avoid getting stuck, but with a fixed budget, each additional chain opened shallows the depth of exploration.

Many engineering breakthroughs require continuous accumulation and constant correction before structural leaps emerge; they can't be achieved simply by 'trying a few more times'.

This actually points towards the development direction of next-generation Agents: not models that 'output an answer once', but systems that can continuously iterate and self-evolve within long-term feedback loops.

AI Engineers Might Really Be Coming

The true far-reaching significance of this research lies in its preliminary outline of an AI system beginning to approach the real engineering cycle.

Imagine when AI connects to industrial software, simulation environments, CAD systems, chip design tools, scientific computing platforms...

A dramatic transformation in the modality of productivity is on the verge of emerging.

In future labs, a division of labor like this might appear:

Human researchers are responsible for proposing directions and goals.

For example, 'reduce this component's energy consumption by 30%', 'compress this model's forward pass GPU usage even lower', 'increase the stability of robot control a bit more', 'push the fidelity of this quantum circuit closer to the limit', etc.

And AI is responsible for 'grinding the path'. They focus on these goals, continuously optimizing.

For example, automatically running simulations and experiments, automatically reading feedback from verifiers and simulators, then continuing to modify and optimize, iterating non-stop 24/7.

This evolutionary logic frees AI from the identity of an 'assistive tool', allowing it to begin solving complex system problems like a real engineering team—and tirelessly at that.

And the issues revealed by the Frontier-Eng Benchmark are actually very direct:

When AI begins to learn 'long-term optimization', how far is it from true engineering intelligence?

Paper Title: Frontier-Eng: Benchmarking Self-Evolving Agents on Real-World Engineering Tasks with Generative Optimization

Project Homepage: https://lab.einsia.ai/frontier-eng/

Arxiv: https://arxiv.org/abs/2604.12290

GitHub repo: https://github.com/EinsiaLab/Frontier-Engineering

This article is from the WeChat public account "Quantum Bit", author: Yun Zhong

Трендовые криптовалюты

Связанные с этим вопросы

QWhat is the main purpose of the Frontier-Eng Benchmark released by Einsteina AI's Navers lab?

AThe main purpose of the Frontier-Eng Benchmark is to move beyond testing AI's ability to recall known information. It systematically tests AI agents' capability for 'iterative optimization' on 47 real-world, open-ended engineering tasks without standard answers, evaluating if they can continuously improve performance through a feedback loop involving simulation, error analysis, and parameter adjustment.

QHow does the AI's role change in the Frontier-Eng Benchmark testing process compared to traditional language models?

AIn the Frontier-Eng Benchmark, the AI transitions from acting as a 'super student' that retrieves and assembles answers from training data to performing 'engineering optimization.' Its role becomes akin to a professional engineer: it proposes solutions, runs simulations, analyzes feedback and errors, modifies parameters/code, and reruns experiments in a continuous loop to seek optimal performance under complex constraints.

QWhat are the two key 'AI evolution laws' discovered through the Frontier-Eng Benchmark regarding iterative optimization?

AThe two key laws are: 1) Improvements become progressively harder and smaller (showing a power-law decay: Improvement frequency ∝ 1/iteration count, Improvement magnitude ∝ 1/improvement count). 2) While exploring multiple parallel paths (breadth) is useful, sustained depth in a single optimization path is more critical for achieving structural breakthroughs, as fixed budgets force a trade-off between breadth and depth.

QWhat future work paradigm does the article suggest might emerge from the development of self-evolving AI agents?

AThe article suggests a future 'Auto Research' paradigm where human researchers define the goals and direction (e.g., 'reduce component energy consumption by 30%'), and AI agents take on the role of 'grinding the path.' They would work autonomously and tirelessly—running simulations, interpreting feedback from verifiers and simulators, and iteratively optimizing—24/7 to approach performance limits.

QAccording to the article, what fundamental shift in AI capability does the Frontier-Eng Benchmark represent?

AThe Frontier-Eng Benchmark represents a fundamental shift from evaluating AI's ability to find predetermined 'correct answers' to testing its capacity for 'self-evolution' through long-term feedback loops. It moves the focus to whether AI can demonstrate sustained learning and improvement in complex, real-world scenarios with no single correct answer, pushing AI closer to genuine engineering intelligence.

Похожее

Почему биткойн-фермы внезапно стали новым входом для вычислительных мощностей ИИ на фоне дефицита электроэнергии в 38 ГВт?

Заголовок: Почему майнинговые фермы для биткоина внезапно стали новым входом для вычислительных мощностей ИИ на фоне дефицита электроэнергии в 38 ГВт? Краткое содержание: Когда конкуренция между центрами обработки данных ИИ сместилась с вопроса «кто купит больше GPU» к «кто раньше получит электроэнергию», некоторые майнинговые фермы для биткоина, ранее считавшиеся волатильными активами, начали трансформироваться в центры обработки данных для облачных провайдеров, используя свои готовые возможности подключения к сети, землю и трансформаторные подстанции. По расчетам Morgan Stanley, в период 2026-2028 годов в США может возникнуть дефицит электроэнергии для ЦОДов около 38 ГВт, и модернизация старых майнинговых ферм может обеспечить от 10 до 19 ГВт. Такие компании, как TeraWulf и Hut 8, переориентируются с добычи криптовалют на предоставление инфраструктуры («Powered Shell Provider»), предлагая клиентам из сферы ИИ критически важный ресурс — возможность быстрее конкурентов развернуть значительные вычислительные мощности. Ключевой ценностью становится не вычислительная мощность для майнинга, а дефицитный доступ к электросетям, получение которого «с нуля» в некоторых регионах США теперь может занять 5-7 лет.

华尔街日报2 мин. назад

Почему биткойн-фермы внезапно стали новым входом для вычислительных мощностей ИИ на фоне дефицита электроэнергии в 38 ГВт?

华尔街日报2 мин. назад

Майкл Сэйлор: «Мы никогда не говорили, что никогда не будем продавать биткоины»

Председатель стратегической комиссии Майкл Сэйлор прокомментировал сообщения о новом разрешении компании Strategy на продажу биткоинов. Он заявил, что данное разрешение не является новым — оно было объявлено ещё 29 июня в рамках системы управления капиталом компании. Соглашение позволяет продавать BTC на сумму до 5 миллиардов долларов для определённых целей, но не обязывает компанию к продаже. Сэйлор подчеркнул, что Strategy никогда официально не брала на себя обязательство никогда не продавать свои биткоины, хотя и рассчитывает оставаться чистым покупателем BTC в долгосрочной перспективе. Он назвал текущие новости «старыми», переподанными как новые, и подтвердил, что программа монетизации биткоинов компании не предполагает обязательной продажи её активов.

cryptonews.ru1 ч. назад

Майкл Сэйлор: «Мы никогда не говорили, что никогда не будем продавать биткоины»

cryptonews.ru1 ч. назад

«Летняя пила» продолжается: пробой $67 000 станет началом роста биткоина

Цена биткоина продолжает консолидироваться в диапазоне $58 000–$67 000 с начала июня. 1 августа актив снизился до $62 217. Аналитики расходятся в краткосрочных прогнозах: некоторые, как Crypto Candy, ожидают тестирования уровня $60 000 или ниже, пока цена находится под $66 000. Другие, как Jelle, видят в боковом движении «летнюю пилу» и придерживаются стратегии усреднения. Ключевым для определения дальнейшего направления считается уровень $67 000. По мнению Daan Crypto Trades, его пробой необходим для выхода из затянувшейся паузы. Roman полагает, что уверенный пробой с объемом может быстро запустить рост к $70 000–$80 000 и выше. С долгосрочной точки зрения, макроаналитик Герт ван Лаген рассматривает текущую фазу как накопление в рамках масштабной формации «чаша с ручкой». Он отмечает, что долгосрочные держатели не спешат продавать актив, о чем говорит показатель NUPL. Таким образом, рынок находится в решающей фазе, где пробой либо поддержки $60 000, либо сопротивления $67 000 задаст тренд на ближайшее будущее.

cryptonews.ru1 ч. назад

«Летняя пила» продолжается: пробой $67 000 станет началом роста биткоина

cryptonews.ru1 ч. назад

На неделе с 3 по 9 августа стоит обратить внимание: Закон CLARITY, возможно, будет поставлен на голосование в Сенате; SpaceX и Circle опубликуют финансовые отчеты

**Важные события на следующей неделе (3–9 августа 2026 г.)** **Ключевые даты:** * **3 августа:** Публикация отчетов American Bitcoin за Q2. Полное закрытие сервисов DeFi-трекера Zapper и кошелька Ctrl Wallet. LayerZero прекратит поддержку ретрансляторов v1. Upbit прекратит торговлю токенами AQT и AERGO. * **4 августа:** Публикация финансовых отчетов SpaceX и Hut 8 за второй квартал 2026 года. * **5 августа:** Circle опубликует отчет за Q2. Начинается предварительное ценовое консультирование для IPO компании Unitree Tech (Ушу Цзишу) в Китае. * **6 августа:** Первая крупная разблокировка акций SpaceX — до 12% от общего капитала. * **7 августа:** Выход важных данных по рынку труда США (отчет о занятости за июль). Предельный срок для Сената США — получить 60 голосов в поддержку **Закона CLARITY** (билль о регулировании криптовалют и этике). Ожидается выпуск Grok 4.6 от xAI. * **8 августа:** Начало принудительной подачи сигналов в сети Bitcoin согласно предложению BIP-110. * **На неделе (дата уточняется):** Ожидается голосование полного состава Сената США по **Закону CLARITY**. Выход нового релиза XRP Ledger (v3.3.0) с новыми функциями, такими как конфиденциальные данные и пакетные транзакции. **Основные темы недели:** корпоративная отчетность (SpaceX, Circle), регулирование (CLARITY Act), рыночные события (разблокировка акций SpaceX, отчет по занятости в США) и обновления в технологиях блокчейна.

marsbit2 ч. назад

На неделе с 3 по 9 августа стоит обратить внимание: Закон CLARITY, возможно, будет поставлен на голосование в Сенате; SpaceX и Circle опубликуют финансовые отчеты

marsbit2 ч. назад

Акции упали сильнее, чем криптовалюты. Куда делись деньги?

Автор: Кэти,白话区块链 28-29 июля, Сеул. Индекс Kospi впервые в истории Южной Кореи два дня подряд срабатывал на приостановку торгов. Первый день: падение на 10.84%, второй день: -5.98%. SK Hynix, крупнейшая по весу акция, потеряла за два дня около 23%. Падение Nasdaq, глобальный обвал акций полупроводниковых компаний, массовые потери на кредитных ETF. За два дня откат Kospi от пика июня достиг 40%. Июль угрожает стать худшим месяцем в истории индекса. Все ранее перегретые сделки были перевернуты, как стол. Это не локальный негатив по одной акции, а глобальное принудительное снижение кредитного плеча. Самое парадоксальное: на этот раз больше всего на «криптовалютное» падение похожи именно акции. Спот: прибыль SK Hynix за второй квартал достигла рекордных 60.54 трлн вон, но из-за несоответствия прогнозу в 64.22 трлн акция подверглась жесткой распродаже. Хорошие новости не растут — это уже плохая новость. Производные инструменты пострадали еще сильнее. Кредитный ETF с плечом 2x на SK Hynix упал на 83% с пика, потеряв в стоимости более 1 трлн гонконгских долларов. Эмитент был вынужден изменить правила продукта. Неожиданно: Биткоин, известный высокой волатильностью, с 1 июля вырос почти на 15%, в то время как акции демонстрировали «криптовалютную» динамику. Это не паника всего рынка, а точечный сброс перегретых позиций. Триггеры: отчет SK Hynix и фактор Китая — крупнейшее IPO ChangXin Memory, направленное на расширение производства DRAM, создало конкуренцию для нарратива о дефиците памяти для ИИ. Дополнительное давление — нормализация политики Банка Японии и потенциальное сокращение кэрри-трейда в иенах. Эксперт Дэн Найлз считает, что это не крах логики ИИ, а «краткосрочное дно», вызванное принудительными ликвидациями мелких инвесторов и хедж-фондов. Промышленная логика не мертва — умерло кредитное плечо. Перетекли ли деньги из акций в Биткоин? Нет. «Устойчивость» Биткоина объясняется тем, что он уже прошел фазу распродаж раньше. В мае-июне американские спотовые BTC-ETF зафиксировали рекордный отток средств. К июлю продавать было уже нечего. Небольшой приток в июле — лишь частичное восстановление. Настоящие «защитные» деньги пошли в золото. Коэффициент корреляции между Биткоином и золотом упал до -0.88. Нарратив о «цифровом золоте» разбит: золото — для сохранения капитала, Биткоин — для роста. Деньги придут в криптоактивы при выполнении трех условий: смягчение глобального давления на ликвидность; снижение ставок ФРС без рецессии; принятие закона CLARITY, устраняющего регуляторные неопределенности. Пока Биткоин — не убежище, а актив, который раньше других прошел очистку. Но когда шторм утихнет и глобальный капитал снова начнет распределяться, Биткоин займет место в первых рядах очереди. Место уже зарезервировано.

marsbit2 ч. назад

Акции упали сильнее, чем криптовалюты. Куда делись деньги?

marsbit2 ч. назад

Торговля

Спот

Популярные статьи

Как купить ERA

Добро пожаловать на HTX.com! Мы сделали приобретение Caldera (ERA) простым и удобным. Следуйте нашему пошаговому руководству и отправляйтесь в свое крипто-путешествие.Шаг 1: Создайте аккаунт на HTXИспользуйте свой адрес электронной почты или номер телефона, чтобы зарегистрироваться и бесплатно создать аккаунт на HTX. Пройдите удобную регистрацию и откройте для себя весь функционал.Создать аккаунтШаг 2: Перейдите в Купить криптовалюту и выберите свой способ оплатыКредитная/Дебетовая Карта: Используйте свою карту Visa или Mastercard для мгновенной покупки Caldera (ERA).Баланс: Используйте средства с баланса вашего аккаунта HTX для простой торговли.Третьи Лица: Мы добавили популярные способы оплаты, такие как Google Pay и Apple Pay, для повышения удобства.P2P: Торгуйте напрямую с другими пользователями на HTX.Внебиржевая Торговля (OTC): Мы предлагаем индивидуальные услуги и конкурентоспособные обменные курсы для трейдеров.Шаг 3: Хранение Caldera (ERA)После приобретения вами Caldera (ERA) храните их в своем аккаунте на HTX. В качестве альтернативы вы можете отправить их куда-либо с помощью перевода в блокчейне или использовать для торговли с другими криптовалютами.Шаг 4: Торговля Caldera (ERA)С легкостью торгуйте Caldera (ERA) на спотовом рынке HTX. Просто зайдите в свой аккаунт, выберите торговую пару, совершайте сделки и следите за ними в режиме реального времени. Мы предлагаем удобный интерфейс как для начинающих, так и для опытных трейдеров.

871 просмотров всегоОпубликовано 2025.07.17Обновлено 2026.06.02

Как купить ERA

Обсуждения

Добро пожаловать в Сообщество HTX. Здесь вы сможете быть в курсе последних новостей о развитии платформы и получить доступ к профессиональной аналитической информации о рынке. Мнения пользователей о цене на ERA (ERA) представлены ниже.

活动图片