DeepSeek V4 Pro Runs 5 Harness Sets: Pi Achieves Highest Success Rate, DSH Most Cost-Effective

08/20 11:20

According to the latest news from Dongcha Beating AI, the agent tool company Composio tested the same DeepSeek V4 Pro 0813 with Pi Agent, DeepSeek Harness 0.1, Claude Code, OpenCode, and Hermes Agent, all set to max reasoning, running over 30 agent tasks. The success rates were as follows: 1. Pi Agent: 21/30 2. DeepSeek Harness: 20/30 3. Claude Code: 19/30 4. OpenCode: 19/30 5. Hermes Agent: 18/30. Out of the 30 tasks, 15 were passed by all five agents, while 7 failed across the board. The remaining 8 tasks only changed the harness, resulting in a reversal of outcomes. Speed rankings were: 1. Claude Code: 181.8 seconds 2. DeepSeek Harness: 252.1 seconds 3. Hermes Agent: 273.6 seconds 4. OpenCode: 280.6 seconds 5. Pi Agent: 362.9 seconds. The cost per successful attempt was: 1. DeepSeek Harness: $0.028 2. Pi Agent: $0.031 3. OpenCode: $0.032 4. Hermes Agent: $0.037 5. Claude Code: $0.074. Pi had the highest success rate, DeepSeek's official Harness was the most cost-effective, and Claude Code was the fastest but the most expensive. Composio previously conducted similar tests using DeepSeek V4 Flash, where Pi also achieved the highest success rate.
Tăng giáGiảm giáThíchChia sẻ
Tuyên bố miễn trừ trách nhiệmNội dung trên không đại diện cho quan điểm của HTX.HTX không đưa ra bất kỳ lời khuyên giao dịch nào.

Tất cả bình luận0Mới nhấtPhổ biến

avatar
Mới nhấtPhổ biến