Who is the Truly Strongest Agent in OpenClaw? Leaderboard of 23 Real-World Task Evaluations Released
This report presents a comprehensive benchmark evaluating the performance of AI coding agents on 23 real-world OpenClaw tasks, focusing solely on the core metric of success rate. The transparent and reproducible testing methodology employs three scoring methods: automated checks, an LLM judge (Claude Opus), and a hybrid approach. The diverse task set covers areas like code/file operations, content creation, research, system tools, and memory persistence.
The top 10 models by success rate (Best % / Avg %) are:
1. anthropic/claude-opus-4.6 (93.3% / 82.0%)
2. arcee-ai/trinity-large-thinking (91.9% / 91.9%)
3. openai/gpt-5.4 (90.5% / 81.7%)
4. qwen/qwen3.5-27b (90.0% / 78.5%)
5. minimax/minimax-m2.7 (89.8% / 83.2%)
Claude Opus 4.6 leads in peak performance, while Arcee's Trinity demonstrates superior average success rate stability. The Qwen series shows strong cost-performance potential with multiple entries in the top ten. All task definitions and scoring logic are publicly available for independent verification.
marsbit04/08 14:45