All benchmarks
xAI · USA · Released 2025-07

Grok 4

xAI's frontier reasoning model with real-time X data grounding and a heavy multi-agent variant (Grok 4 Heavy).

Access
closed
Context
256,000 tokens
Modalities
text, vision
Languages
80+

Benchmark scores

MMLU
87.0%
Massive Multitask Language Understanding
GPQA
87.5%
Graduate-level science QA (Diamond)
HumanEval
92.0%
Python code completion
SWE-bench
N/A
Real GitHub issue resolution (Verified)
LiveCodeBench
N/A
Contamination-free coding
AIME
95.0%
American Invitational Mathematics Exam
MMMU
78.0%
Multimodal college-level reasoning
MATH
93.0%
Competition-level math word problems

Scores as reported by the vendor or leading public leaderboards. "N/A" means the score has not been publicly disclosed for this metric.

Explain like I'm 5

Very good at hard exams and knows what's happening on X right now.

Key features
  • Real-time X grounding
  • Multi-agent Heavy mode
  • Tool use
Strengths
  • Top-tier on Humanity's Last Exam
  • Fresh real-time knowledge
Limitations
  • Limited multimodal breadth
  • Closed source

Best for

Real-time researchReasoning-heavy tasks

Related models