Kimi K3 White-Collar Tasks Approach Fable5, Costs Soar to 10x Previous Generation
07/22 09:10
According to Dongcha Beating's monitoring, Artificial Analysis has updated its AA-Briefcase ranking. This evaluation requires models to find information from nearly 2,000 emails, Slack records, and company documents, and then deliver tables, presentations, and interface prototypes. It consists of 4 long-term projects and 91 confidential tasks. Kimi K3 scored 1543 Elo, second only to Claude Fable 5's 1574. It surpasses GPT-5.6 Sol's 1501 and also leads Claude Sonnet 5 and Claude Opus 4.8. K3's objective requirement pass rate is 51%, while Fable 5's is 56%. However, K3's analysis quality score is slightly higher, 1754 vs. 1744. It mainly loses in final product presentation, scoring lower than GPT-5.6 Sol and Opus 4.8. The performance improvement also comes with higher token consumption. K3 costs an average of $10.57 per task, about 10 times that of Kimi K2.6. It executes an average of 83 rounds, outputs 120,000 tokens, and takes 56.4 minutes. The time required to complete similar tasks is about 2.5 times that of Fable 5.
BoğaAyıBeğenPaylaş
Sorumluluk Reddi:Yukarıdaki içerik HTX'ın tutumunu temsil etmez.,HTX herhangi bir alım satım önerisinde bulunmaz.。
Tüm Yorumlar0En yeniPopüler