DeepSeek V4 Pro Runs 5 Harness Sets: Pi Achieves Highest Success Rate, DSH Most Cost-Effective

08/20 11:20

According to the latest news from Dongcha Beating AI, the agent tool company Composio tested the same DeepSeek V4 Pro 0813 with Pi Agent, DeepSeek Harness 0.1, Claude Code, OpenCode, and Hermes Agent, all set to max reasoning, running over 30 agent tasks. The success rates were as follows: 1. Pi Agent: 21/30 2. DeepSeek Harness: 20/30 3. Claude Code: 19/30 4. OpenCode: 19/30 5. Hermes Agent: 18/30. Out of the 30 tasks, 15 were passed by all five agents, while 7 failed across the board. The remaining 8 tasks only changed the harness, resulting in a reversal of outcomes. Speed rankings were: 1. Claude Code: 181.8 seconds 2. DeepSeek Harness: 252.1 seconds 3. Hermes Agent: 273.6 seconds 4. OpenCode: 280.6 seconds 5. Pi Agent: 362.9 seconds. The cost per successful attempt was: 1. DeepSeek Harness: $0.028 2. Pi Agent: $0.031 3. OpenCode: $0.032 4. Hermes Agent: $0.037 5. Claude Code: $0.074. Pi had the highest success rate, DeepSeek's official Harness was the most cost-effective, and Claude Code was the fastest but the most expensive. Composio previously conducted similar tests using DeepSeek V4 Flash, where Pi also achieved the highest success rate.
BullishBearishLikeShare
DisclaimerThe content above does not represent HTX's positions.HTX does not provide any trading recommendations.

All Comments0LatestHot

avatar
LatestHot