All benchmarks
OpenAI · USA · Released 2025-08
GPT-5
OpenAI's flagship reasoning + multimodal model with a unified router that switches between fast and deep-thinking modes.
Access
closed
Context
400,000 tokens
Modalities
text, vision, audio
Languages
100+
Benchmark scores
MMLU
91.4%
Massive Multitask Language Understanding
GPQA
89.4%
Graduate-level science QA (Diamond)
HumanEval
96.3%
Python code completion
SWE-bench
74.9%
Real GitHub issue resolution (Verified)
LiveCodeBench
N/A
Contamination-free coding
AIME
94.6%
American Invitational Mathematics Exam
MMMU
84.2%
Multimodal college-level reasoning
MATH
96.7%
Competition-level math word problems
Scores as reported by the vendor or leading public leaderboards. "N/A" means the score has not been publicly disclosed for this metric.
Explain like I'm 5
The smartest ChatGPT so far. It thinks longer on hard problems and can see, hear and write really well.
Key features
- Unified reasoning router
- 400K context
- Native vision + audio
- Tool use / agents
- Long agentic tasks
Strengths
- State-of-the-art reasoning
- Strong coding & math
- Very low hallucination rate vs GPT-4o
Limitations
- Closed source
- Rate limits on the highest reasoning tier
Best for
Agentic codingResearch assistantsComplex analysisMultimodal apps
Related models
Anthropic
Claude Sonnet 4.5
Anthropic's best coding + agentic model, purpose-built for long-horizon computer-use and software engineering tasks.
Google DeepMind
Gemini 2.5 Pro
Google's most capable model with native 1M-token context and full multimodal I/O across text, images, audio and video.
xAI
Grok 4
xAI's frontier reasoning model with real-time X data grounding and a heavy multi-agent variant (Grok 4 Heavy).
Meta AI
Llama 4 Maverick
Meta's flagship MoE model — 400B total / 17B active — with native multimodal input and 1M-token context.