AI Benchmarks
Compare frontier models side by side. Scores are drawn from official model cards, vendor releases and leading public leaderboards. Blank cells mean the score has not been publicly reported.
MiniMax's open MoE tuned for coding and agentic workflows, positioned as a cost-efficient Claude alternative.
Anthropic's best coding + agentic model, purpose-built for long-horizon computer-use and software engineering tasks.
Ultra-efficient MoE with sparse attention. Delivers frontier-adjacent quality at a fraction of the cost.
Alibaba's flagship trillion-parameter Qwen3 tier — strong multilingual, coding and agentic performance.
OpenAI's flagship reasoning + multimodal model with a unified router that switches between fast and deep-thinking modes.
xAI's frontier reasoning model with real-time X data grounding and a heavy multi-agent variant (Grok 4 Heavy).
Moonshot's open agentic MoE (1T total / 32B active) tuned for tool use and long-horizon tasks.
Zhipu's unified reasoning + coding + agent model, positioned as a Claude-class open alternative.
Google's most capable model with native 1M-token context and full multimodal I/O across text, images, audio and video.
India's flagship 24B open model with strong Indic-language reasoning and math performance.
Meta's flagship MoE model — 400B total / 17B active — with native multimodal input and 1M-token context.
Cohere's enterprise-grade model tuned for RAG, agents and multilingual business workflows.
Baidu's flagship multimodal model with strong Chinese-language performance and enterprise tooling.
Ola Krutrim's multilingual model covering all 22 official Indian languages.
Microsoft's 14B small language model trained on curated + synthetic data — punches far above its weight on reasoning.
Europe's flagship dense 123B model — strong multilingual + code performance with open weights for research.
Tencent's 389B open MoE (52B active) with a 256K context window.
Hybrid SSM-Transformer (Mamba + attention) with a very long 256K context and open weights.
NVIDIA's 340B open model family designed for synthetic data generation and enterprise fine-tuning.
Open translation and understanding models from IIT Madras covering all 22 scheduled Indian languages.
Sources: OpenAI, Anthropic, Google DeepMind, xAI, Meta, DeepSeek, Alibaba Qwen, Moonshot AI, Zhipu AI, Mistral, Cohere, AI21, Microsoft Research, NVIDIA, Baidu, Tencent, MiniMax, Sarvam AI, Krutrim, AI4Bharat — plus lmarena.ai, Artificial Analysis, SWE-bench, LiveCodeBench and the HuggingFace Open LLM Leaderboard.