AUTO-UPDATED

Kimi K3: second only to Fable 5 on AA-Briefcase

Moonshot AI has released Kimi K3, a 2.8T parameter model that ranks second on the AA-Briefcase agentic benchmark, though it faces challenges regarding high operational costs and processing time.

Key Points

  • Kimi K3 achieved an AA-Briefcase Elo of 1543, trailing only Claude Fable 5 and outperforming GPT-5.6 Sol and Claude Opus 4.8.
  • The model recorded a 51% rubric pass rate and strong analytical quality, though its presentation quality remains lower than several competing frontier models.
  • Operational costs average $10.57 per task, driven by high token usage and an average of 83 turns per task.
  • Kimi K3 requires an average of 56.4 minutes to complete a task, which is significantly slower than other top-tier models on the benchmark.
  • The model represents a substantial performance leap over its predecessor, Kimi K2.6, which previously scored an Elo of 816.

Why it Matters

This release highlights the trade-offs between high-level analytical reasoning and operational efficiency in the current landscape of agentic AI models. While Kimi K3 demonstrates competitive intelligence, its high cost and slow processing speed may limit its immediate adoption for time-sensitive or budget-constrained enterprise workflows.
Artificialanalysis.ai Published by Artificial Analysis
Read original