How can we determine if artificial intelligence is intelligent enough to conduct scientific research and achieve end-to-end autonomous scientific discovery?
The latest research from the University of Science and Technology of China (USTC) has allowed an AI "brain" to take over a machine scientist laboratory "body," advancing this question from knowledge Q&A and solution generation to experimental execution and feedback learning in the real physical world.

Paper link: https://arxiv.org/abs/2607.23045
The research team built a machine catalysis laboratory, which includes 45 modular automated workstations covering synthesis, characterization, and catalytic performance testing.
To enable the AI to understand and utilize the lab's capabilities, the research team encapsulated these capabilities into machine-readable skills. The AI agent can directly call on experimental equipment through these skills while being constrained by the actual equipment capabilities, operational protocols, and experimental conditions.

Figure 1. Architecture of an AI-readable machine scientist laboratory for catalysis research.

Figure 2. Path from scientific intent to machine laboratory execution.
Pressure Test
Based on this platform, the research team systematically evaluated 48 practical configurations composed of 6 agent frameworks and 9 large language models.
The test covered 32 research tasks defined by domain experts, totaling 4,608 evaluation runs. The assessment not only examined whether the agent could generate experimental plans but also tracked whether these plans could pass verification and be dispatched to the robotic system, and whether the generated workflows could be executed directly without human intervention.
Harsh Real-World Evaluation Results
Current leading large language model agents still have a significant gap from "taking over." Out of 4,608 tests, only 151 workflows could be executed without manual fixes, accounting for 3.3% of all tests. The best-performing combination, Claude Code and Claude Opus 4.7, achieved an executable rate of 28.1%; the Codex and GPT 5.5 combination achieved an executable rate of 19.8%.
Adjusting Parameters Does Not Equal Re-planning
The research further examined whether the agent could learn from real experimental results. The team placed the Codex/GPT 5.5 combination in an open-ended, five-round closed loop to continuously conduct experimental planning, robotic execution, evidence acquisition, and re-planning.
The agent was able to adjust material formulas and operational conditions based on the returned experimental results, but these adjustments mainly stayed at the level of local parameter optimization.
During the five rounds of experiments, the agent consistently retained the original workflow skeleton, failing to redesign analysis methods or correct some persistent key omissions, such as the lack of electrode binders and colorimetric reagents for specific analytes.
The results indicate that being able to read experimental feedback and adjust parameters does not equate to being able to identify problems with the research strategy itself, let alone perform scientific-level re-planning.

Figure 3. Performance of AI agents across executable planning, verification, and laboratory task dispatch.
Long-Range Planning Remains a Bottleneck
Although some agents generated workflows deemed executable by expert evaluation, with some containing up to 44 operational steps, only 3 workflows across all tests exceeded 30 steps. This suggests that as the task chain lengthens, the agent's ability to maintain experimental logical integrity and physical executability still faces significant challenges.
Redefining AI's Scientific Capabilities
The research thus distinguishes between three capabilities often conflated in current discussions about "AI scientists": first, fluently generating experimental plans; second, generating workflows that can actually be executed in a physical laboratory; and third, adjusting overall research strategies based on experimental results.
The research results show that linguistic-level planning capability does not automatically translate into reliable experimental execution capability, and local parameter adjustment cannot be directly regarded as scientific-level re-planning.
From "AI Proving Ground" to "AI Training Ground"
As intelligent research infrastructure for AI for Science, the machine laboratory can serve both as a "proving ground" to test AI's scientific capabilities and as a "training ground" to drive AI's continuous evolution. The AI agent, through machine-readable skills, directly converses with the machine lab, transforming scientific intent into executable tasks and receiving real feedback from machine execution, instrument status, and experimental results.
The resulting closed loop of "planning—execution—feedback—re-planning" can not only expose AI's deficiencies in knowledge, operation, and research strategy but also capture successful workflows, failure cases, experimental results, and expert evaluations as data for model training and agent alignment, continuously iterating AI models and their agent systems.
Thus, the machine scientist provides quantifiable, repeatable, and verifiable physical world evidence to answer "How can we determine if artificial intelligence is intelligent enough to conduct scientific research?" and drives AI from generating experimental protocols to gradually achieving end-to-end autonomous scientific discovery encompassing scientific question posing, experimental design, robotic execution, data analysis, and evidence-driven re-planning.

Figure 4. Five-round iteration of the AI-Machine Scientist on an open scientific question.
Reference: https://arxiv.org/abs/2607.23045
This article is from the WeChat public account "Xin Zhi Yuan," author: Xin Zhi Yuan; editor: LRST








