All benchmarks
Moonshot AI · China · Released 2025-07
Kimi K2
Moonshot's open agentic MoE (1T total / 32B active) tuned for tool use and long-horizon tasks.
Access
open-weights
Context
128,000 tokens
Modalities
text
Languages
40+
Benchmark scores
MMLU
84.1%
Massive Multitask Language Understanding
GPQA
68.4%
Graduate-level science QA (Diamond)
HumanEval
89.6%
Python code completion
SWE-bench
65.8%
Real GitHub issue resolution (Verified)
LiveCodeBench
N/A
Contamination-free coding
AIME
N/A
American Invitational Mathematics Exam
MMMU
N/A
Multimodal college-level reasoning
MATH
N/A
Competition-level math word problems
Scores as reported by the vendor or leading public leaderboards. "N/A" means the score has not been publicly disclosed for this metric.
Explain like I'm 5
An open model built to be an agent — it uses tools and plans multi-step tasks.
Key features
- Agentic tuning
- MoE
- Tool use
Strengths
- Strong SWE-bench for open weights
- Great tool-use
Limitations
- Text-only
Best for
Open-source agentsCoding assistants
Related models
OpenAI
GPT-5
OpenAI's flagship reasoning + multimodal model with a unified router that switches between fast and deep-thinking modes.
Anthropic
Claude Sonnet 4.5
Anthropic's best coding + agentic model, purpose-built for long-horizon computer-use and software engineering tasks.
Google DeepMind
Gemini 2.5 Pro
Google's most capable model with native 1M-token context and full multimodal I/O across text, images, audio and video.
xAI
Grok 4
xAI's frontier reasoning model with real-time X data grounding and a heavy multi-agent variant (Grok 4 Heavy).