Solving a 20-Year Math Problem, Microsoft Open-Sources Argus, Using Evidence-Driven Automatic Research for 1,548 Hours
Microsoft, in collaboration with institutions like Shanghai Jiao Tong University, has open-sourced Argus, a general-purpose Agent reasoning runtime designed for long-term research tasks. The system addresses a key bottleneck in current AI agents: while they can execute actions (through "Harness"), they lack autonomous, long-term decision-making ("Driving") for projects spanning days. Argus introduces an "Evidence-Driven" approach, where the agent's next steps are determined by accumulated evidence rather than a rigid initial goal. This enables sustained, multi-day operation with minimal human intervention, averaging one human request per 40.7 hours over 1,548 hours of wall-clock time tested.
The runtime architecture organizes work into Campaigns and Missions, employing a multi-agent loop with Manager, Planner, Engineer, and Reviewer roles. This separation of planning, execution, and validation improves efficiency and prevents local optimization. Argus is designed with a decoupled core and vertical components, allowing domain experts to customize workflows for fields like mathematics, GPU optimization, and chip design.
In less than a month, Argus has delivered concrete research outcomes across AI4AI, GPU kernels, AI4Science, chip design, AI4Math, and AI4System tasks. These results demonstrate its ability to autonomously drive projects from execution to exploration of open-ended research questions, effectively transitioning the human role from a constant driver to a supervisory co-pilot. Argus represents a shift towards organizing research as a continuously evolving, evidence-guided system that progresses independently of human availability.
marsbit09/07 03:20