Kimi K3 White-Collar Tasks Approach Fable5, Costs Soar to 10x Previous Generation

07/22 09:10

According to Dongcha Beating's monitoring, Artificial Analysis has updated its AA-Briefcase ranking. This evaluation requires models to find information from nearly 2,000 emails, Slack records, and company documents, and then deliver tables, presentations, and interface prototypes. It consists of 4 long-term projects and 91 confidential tasks. Kimi K3 scored 1543 Elo, second only to Claude Fable 5's 1574. It surpasses GPT-5.6 Sol's 1501 and also leads Claude Sonnet 5 and Claude Opus 4.8. K3's objective requirement pass rate is 51%, while Fable 5's is 56%. However, K3's analysis quality score is slightly higher, 1754 vs. 1744. It mainly loses in final product presentation, scoring lower than GPT-5.6 Sol and Opus 4.8. The performance improvement also comes with higher token consumption. K3 costs an average of $10.57 per task, about 10 times that of Kimi K2.6. It executes an average of 83 rounds, outputs 120,000 tokens, and takes 56.4 minutes. The time required to complete similar tasks is about 2.5 times that of Fable 5.
看漲看跌按讚分享
免責聲明以上內容不代表 HTX 的任何立場HTX 不為任何交易提供相關決策建議

全部評論0最新熱門

avatar
最新熱門