Kimi K3 Benchmarks
Where K3 lands against the frontier — and where it still trails.
Headline scores
| Benchmark | Kimi K3 | Nearest competitor |
|---|---|---|
| AA Intelligence Index (v4.1) | 57.1 (#4 of 189) | Claude Fable 5 (#1) |
| arena.ai Code WebDev | 1,679 (#1) | Claude Fable 5 (#2) |
| BrowseComp | 91.2 | Claude Fable 5 (88.6) |
| DeepSWE | 67.5 | FrontierSWE 81.2 (Fable 5) |
Kimi K3 vs GPT-5.6 Sol
K3 performs better in some reported tests, but Moonshot states it still trails GPT-5.6 Sol overall. The gap depends on task, harness, and reasoning effort. On Arena's frontend WebDev leaderboard, K3 ranks first — ahead of both Fable 5 and GPT-5.6 Sol.
Kimi K3 vs Claude Fable 5
K3 leads Fable 5 in Arena's preliminary WebDev ranking, but Fable 5 remains stronger overall on the broader Intelligence Index. No single benchmark establishes universal superiority.
Related: Kimi K3 pricing, specs & architecture.
Frequently Asked Questions
Is Kimi K3 better than GPT-5.6 Sol?
On some tests K3 leads, but Moonshot states it still trails GPT-5.6 Sol overall. Results depend on task, tools, and reasoning effort.
Is Kimi K3 better than Claude Fable 5?
K3 leads Arena's preliminary WebDev ranking, but Fable 5 remains stronger overall per Moonshot and independent indexes.
What is Kimi K3's Intelligence Index score?
Artificial Analysis places Kimi K3 at about 57 on its Intelligence Index v4.1, ranking fourth among 189 models — on par with Claude Opus 4.8 and GPT-5.5.