DeepSeek V4 Pro Runs 5 Harness Sets: Pi Achieves Highest Success Rate, DSH Most Cost-Effective

08/20 11:20

According to the latest news from Dongcha Beating AI, the agent tool company Composio tested the same DeepSeek V4 Pro 0813 with Pi Agent, DeepSeek Harness 0.1, Claude Code, OpenCode, and Hermes Agent, all set to max reasoning, running over 30 agent tasks. The success rates were as follows: 1. Pi Agent: 21/30 2. DeepSeek Harness: 20/30 3. Claude Code: 19/30 4. OpenCode: 19/30 5. Hermes Agent: 18/30. Out of the 30 tasks, 15 were passed by all five agents, while 7 failed across the board. The remaining 8 tasks only changed the harness, resulting in a reversal of outcomes. Speed rankings were: 1. Claude Code: 181.8 seconds 2. DeepSeek Harness: 252.1 seconds 3. Hermes Agent: 273.6 seconds 4. OpenCode: 280.6 seconds 5. Pi Agent: 362.9 seconds. The cost per successful attempt was: 1. DeepSeek Harness: $0.028 2. Pi Agent: $0.031 3. OpenCode: $0.032 4. Hermes Agent: $0.037 5. Claude Code: $0.074. Pi had the highest success rate, DeepSeek's official Harness was the most cost-effective, and Claude Code was the fastest but the most expensive. Composio previously conducted similar tests using DeepSeek V4 Flash, where Pi also achieved the highest success rate.
看漲看跌按讚分享
免責聲明以上內容不代表 HTX 的任何立場HTX 不為任何交易提供相關決策建議

全部評論0最新熱門

avatar
最新熱門