Kimi K3 achieved second place in the AI agent benchmark 'AA-Briefcase,' behind only Fable 5, but its execution cost was higher than Opus 4.8, taking nearly an hour per task on average.

Moonshot AI, a Chinese AI company, announced
Kimi K3: second only to Fable 5 on AA-Briefcase
https://artificialanalysis.ai/articles/kimi-k3-agentic-knowledge-benchmark
AA-Briefcase is a benchmark developed by Artificial Analysis to measure the practical knowledge work capabilities of AI agents. It involves having AI conduct research and analysis, then create deliverables such as spreadsheets, presentations, documents, and websites. Evaluation criteria include a pass rate indicating how well the AI met specified requirements, as well as a comparison of the quality of the analysis and the readability and overall finish of the deliverables between different models. AA-Briefcase Elo, which integrates these aspects, assesses the overall ability to perform multi-stage tasks using tools and complete usable deliverables, going beyond simple question-answering.
The Kimi K3 achieved a score of 1543 in the 'AA-Briefcase Elo' metric, which integrates the pass rate of evaluation criteria, analysis quality Elo, and presentation quality Elo. This places it second overall, behind the Claude Fable 5 with 1574, and surpasses the GPT-5.6 Sol (1501), the Claude Sonnet 5 (1388), and the Claude Opus 4.8 (1347). Furthermore, it represents a significant improvement in overall score compared to the 816 achieved by its predecessor, the Kimi K2.6.

Kimi K3 achieved a score of 1754 in 'Analysis Quality Elo,' which evaluates the content of the analysis of the deliverables, surpassing Claude Fable 5's 1744 and GPT-5.6 Sol's 1599, making it the best among the models compared. On the other hand, its 'Presentation Quality Elo,' which evaluates the structure, layout, and readability, scored 1471, falling below GPT-5.6 Sol's 1660 and Claude Opus 4.8's 1492.

The average cost per task for Kimi K3 in AA-Briefcase was $10.57 (approximately ¥1720). This was broken down as follows: answer generation was $5.86 (approximately ¥955), inference was $2.91 (approximately ¥475), and other input and cache-related costs were also included. Kimi K3 is cheaper than Claude Sonnet 5 at $14.43 (approximately ¥2350) and Claude Fable 5 at $22.30 (approximately ¥3640), but it is more expensive than Claude Opus 4.8 at $8.26 (approximately ¥1350), making it the fourth most expensive among the comparison.

Comparing the AA-Briefcase Elo with the cost per task, it ranks second in overall performance, behind only the Claude Fable 5. However, it falls outside the top-left region, which is considered low-cost and high-performance, indicating that it is a model that comes at a considerable cost in exchange for its high performance. The GPT-5.6 Sol shows similar performance at a lower cost than the Kimi K3, while the Claude Fable 5 is more powerful but also more expensive.

The Kimi K3 took an average of 56.4 minutes to complete one AA-Briefcase task, the longest among the models compared. This was significantly faster than Claude Sonnet 5's 38.3 minutes, Claude Opus 4.8's 24.3 minutes, and Claude Fable 5's 22.9 minutes. Compared to the previous generation Kimi K2.6's 14.5 minutes, the Kimi K3 represents an approximately 3.9-fold increase in processing time.

Furthermore, the Kimi K3 took an average of 83 turns to complete one AA-Briefcase task. This is the third highest, after the Claude Sonnet 5 (183 turns) and the MiniMax M3 (101 turns), and significantly higher than the previous generation Kimi K2.6's 54 turns. Since fewer turns are generally considered more efficient, this indicates that the Kimi K3 tends to progress by repeating complex tasks.

Kimi K3 output an average of approximately 120,000 tokens per AA-Briefcase task. This breaks down to approximately 56,000 tokens for the answer and 63,000 tokens for inference, a significant increase from the previous generation Kimi K2.6's approximately 42,000 tokens. This is the third largest output, following Claude Opus 4.8 (approximately 126,000 tokens), Claude Fable 5 (approximately 139,000 tokens), and Claude Sonnet 5 (approximately 209,000 tokens), and is likely one of the reasons for the increased cost and processing time of Kimi K3. Artificial Analysis concludes that 'while Kimi K3 possesses extremely high performance, it is a model that leaves significant challenges in execution cost and processing efficiency.'

Related Posts:
in AI, Posted by log1i_yk







