Moonshot AI has released Kimi K3, a 2.8T parameter model that ranks second on the AA-Briefcase agentic benchmark, though it faces challenges regarding high operational costs and processing time.
Key Points
- Kimi K3 achieved an AA-Briefcase Elo of 1543, trailing only Claude Fable 5 and outperforming GPT-5.6 Sol and Claude Opus 4.8.
- The model recorded a 51% rubric pass rate and strong analytical quality, though its presentation quality remains lower than several competing frontier models.
- Operational costs average $10.57 per task, driven by high token usage and an average of 83 turns per task.
- Kimi K3 requires an average of 56.4 minutes to complete a task, which is significantly slower than other top-tier models on the benchmark.
- The model represents a substantial performance leap over its predecessor, Kimi K2.6, which previously scored an Elo of 816.