Benchmark intelligence

AI Benchmarks

Compare frontier models side by side. Scores are drawn from official model cards, vendor releases and leading public leaderboards. Blank cells mean the score has not been publicly reported.

China· Open weights
MiniMax-M2
MiniMax

MiniMax's open MoE tuned for coding and agentic workflows, positioned as a cost-efficient Claude alternative.

mmlu
82.0%
gpqa
N/A
humaneval
90.4%
swebench
55.7%
200K ctxcodingagents
Official
USA· Closed
Claude Sonnet 4.5
Anthropic

Anthropic's best coding + agentic model, purpose-built for long-horizon computer-use and software engineering tasks.

mmlu
88.5%
gpqa
83.4%
humaneval
95.4%
swebench
77.2%
200K ctxcodingreasoningagents
Official
China· Open weights
DeepSeek V3.2
DeepSeek

Ultra-efficient MoE with sparse attention. Delivers frontier-adjacent quality at a fraction of the cost.

mmlu
87.1%
gpqa
71.5%
humaneval
92.7%
swebench
49.2%
128K ctxcodingreasoning
Official
China· Open weights
Qwen3-Max
Alibaba Qwen

Alibaba's flagship trillion-parameter Qwen3 tier — strong multilingual, coding and agentic performance.

mmlu
86.0%
gpqa
70.0%
humaneval
90.2%
swebench
N/A
262K ctxcodingmultimodalagents
Official
USA· Closed
GPT-5
OpenAI

OpenAI's flagship reasoning + multimodal model with a unified router that switches between fast and deep-thinking modes.

mmlu
91.4%
gpqa
89.4%
humaneval
96.3%
swebench
74.9%
400K ctxcodingreasoningvision
Official
USA· Closed
Grok 4
xAI

xAI's frontier reasoning model with real-time X data grounding and a heavy multi-agent variant (Grok 4 Heavy).

mmlu
87.0%
gpqa
87.5%
humaneval
92.0%
swebench
N/A
256K ctxreasoningcodingagents
Official
China· Open weights
Kimi K2
Moonshot AI

Moonshot's open agentic MoE (1T total / 32B active) tuned for tool use and long-horizon tasks.

mmlu
84.1%
gpqa
68.4%
humaneval
89.6%
swebench
65.8%
128K ctxcodingagentsreasoning
Official
China· Open weights
GLM-4.5
Zhipu AI

Zhipu's unified reasoning + coding + agent model, positioned as a Claude-class open alternative.

mmlu
84.6%
gpqa
79.1%
humaneval
90.6%
swebench
64.2%
128K ctxcodingreasoningagents
Official
USA· Closed
Gemini 2.5 Pro
Google DeepMind

Google's most capable model with native 1M-token context and full multimodal I/O across text, images, audio and video.

mmlu
89.8%
gpqa
86.4%
humaneval
92.6%
swebench
63.8%
1M ctxreasoningvisionaudio
Official
India· Open weights
Sarvam-M
Sarvam AI

India's flagship 24B open model with strong Indic-language reasoning and math performance.

mmlu
74.7%
gpqa
N/A
humaneval
N/A
swebench
N/A
32K ctxreasoning
Official
USA· Open weights
Llama 4 Maverick
Meta AI

Meta's flagship MoE model — 400B total / 17B active — with native multimodal input and 1M-token context.

mmlu
85.5%
gpqa
69.8%
humaneval
88.0%
swebench
N/A
1M ctxmultimodalvisioncoding
Official
Canada· Open weights
Command A
Cohere

Cohere's enterprise-grade model tuned for RAG, agents and multilingual business workflows.

mmlu
85.0%
gpqa
N/A
humaneval
86.0%
swebench
N/A
256K ctxagentsreasoning
Official
China· Closed
ERNIE 4.5
Baidu

Baidu's flagship multimodal model with strong Chinese-language performance and enterprise tooling.

mmlu
82.0%
gpqa
N/A
humaneval
N/A
swebench
N/A
128K ctxmultimodalvision
Official
India· Open weights
Krutrim 2
Krutrim

Ola Krutrim's multilingual model covering all 22 official Indian languages.

mmlu
72.0%
gpqa
N/A
humaneval
N/A
swebench
N/A
128K ctxreasoning
Official
USA· Open weights
Phi-4
Microsoft

Microsoft's 14B small language model trained on curated + synthetic data — punches far above its weight on reasoning.

mmlu
84.8%
gpqa
56.1%
humaneval
82.6%
swebench
N/A
16K ctxreasoningcoding
Official
France· Open weights
Mistral Large 2
Mistral AI

Europe's flagship dense 123B model — strong multilingual + code performance with open weights for research.

mmlu
84.0%
gpqa
N/A
humaneval
92.0%
swebench
N/A
128K ctxcodingreasoning
Official
China· Open weights
Hunyuan-Large
Tencent

Tencent's 389B open MoE (52B active) with a 256K context window.

mmlu
82.8%
gpqa
N/A
humaneval
71.4%
swebench
N/A
256K ctxreasoning
Official
Israel· Open weights
Jamba 1.5 Large
AI21 Labs

Hybrid SSM-Transformer (Mamba + attention) with a very long 256K context and open weights.

mmlu
81.2%
gpqa
N/A
humaneval
71.4%
swebench
N/A
256K ctxreasoning
Official
USA· Open weights
Nemotron-4 340B
NVIDIA

NVIDIA's 340B open model family designed for synthetic data generation and enterprise fine-tuning.

mmlu
81.1%
gpqa
N/A
humaneval
73.2%
swebench
N/A
4K ctxreasoning
Official
India· Open
IndicTrans2 / AI4Bharat Suite
AI4Bharat

Open translation and understanding models from IIT Madras covering all 22 scheduled Indian languages.

mmlu
N/A
gpqa
N/A
humaneval
N/A
swebench
N/A
4K ctxmultimodal
Official

Sources: OpenAI, Anthropic, Google DeepMind, xAI, Meta, DeepSeek, Alibaba Qwen, Moonshot AI, Zhipu AI, Mistral, Cohere, AI21, Microsoft Research, NVIDIA, Baidu, Tencent, MiniMax, Sarvam AI, Krutrim, AI4Bharat — plus lmarena.ai, Artificial Analysis, SWE-bench, LiveCodeBench and the HuggingFace Open LLM Leaderboard.