Kimi K3 Benchmarks

Where K3 lands against the frontier — and where it still trails.

Headline scores

BenchmarkKimi K3Nearest competitor
AA Intelligence Index (v4.1)57.1 (#4 of 189)Claude Fable 5 (#1)
arena.ai Code WebDev1,679 (#1)Claude Fable 5 (#2)
BrowseComp91.2Claude Fable 5 (88.6)
DeepSWE67.5FrontierSWE 81.2 (Fable 5)

Kimi K3 vs GPT-5.6 Sol

K3 performs better in some reported tests, but Moonshot states it still trails GPT-5.6 Sol overall. The gap depends on task, harness, and reasoning effort. On Arena's frontend WebDev leaderboard, K3 ranks first — ahead of both Fable 5 and GPT-5.6 Sol.

Kimi K3 vs Claude Fable 5

K3 leads Fable 5 in Arena's preliminary WebDev ranking, but Fable 5 remains stronger overall on the broader Intelligence Index. No single benchmark establishes universal superiority.

Reasoning-effort caveat: K3 reasons at maximum effort by default, and reasoning tokens count toward output. Real per-task cost can run higher than the headline rate suggests.

Related: Kimi K3 pricing, specs & architecture.

KimiK3 Max is an independent fan and reference site. It is not affiliated with, endorsed by, or operated by Moonshot AI. “Kimi” is a trademark of Moonshot AI. All trademarks, model names, and benchmark data belong to their respective owners.

Frequently Asked Questions

Is Kimi K3 better than GPT-5.6 Sol?

On some tests K3 leads, but Moonshot states it still trails GPT-5.6 Sol overall. Results depend on task, tools, and reasoning effort.

Is Kimi K3 better than Claude Fable 5?

K3 leads Arena's preliminary WebDev ranking, but Fable 5 remains stronger overall per Moonshot and independent indexes.

What is Kimi K3's Intelligence Index score?

Artificial Analysis places Kimi K3 at about 57 on its Intelligence Index v4.1, ranking fourth among 189 models — on par with Claude Opus 4.8 and GPT-5.5.