DeepSeek V4 Pro Runs 5 Harness Sets: Pi Achieves Highest Success Rate, DSH Most Cost-Effective

08/20 11:20

According to the latest news from Dongcha Beating AI, the agent tool company Composio tested the same DeepSeek V4 Pro 0813 with Pi Agent, DeepSeek Harness 0.1, Claude Code, OpenCode, and Hermes Agent, all set to max reasoning, running over 30 agent tasks. The success rates were as follows: 1. Pi Agent: 21/30 2. DeepSeek Harness: 20/30 3. Claude Code: 19/30 4. OpenCode: 19/30 5. Hermes Agent: 18/30. Out of the 30 tasks, 15 were passed by all five agents, while 7 failed across the board. The remaining 8 tasks only changed the harness, resulting in a reversal of outcomes. Speed rankings were: 1. Claude Code: 181.8 seconds 2. DeepSeek Harness: 252.1 seconds 3. Hermes Agent: 273.6 seconds 4. OpenCode: 280.6 seconds 5. Pi Agent: 362.9 seconds. The cost per successful attempt was: 1. DeepSeek Harness: $0.028 2. Pi Agent: $0.031 3. OpenCode: $0.032 4. Hermes Agent: $0.037 5. Claude Code: $0.074. Pi had the highest success rate, DeepSeek's official Harness was the most cost-effective, and Claude Code was the fastest but the most expensive. Composio previously conducted similar tests using DeepSeek V4 Flash, where Pi also achieved the highest success rate.
БичачийВедмежийВподобайкаПоділитися
ЗастереженняНаведені вище матеріали не представляють позицію компанії HTX.HTX не дає жодних торгових рекомендацій.

Усі коментарі0НовіПопулярно

avatar
НовіПопулярно