All benchmarks
Google DeepMind · USA · Released 2025-06

Gemini 2.5 Pro

Google's most capable model with native 1M-token context and full multimodal I/O across text, images, audio and video.

Access
closed
Context
1,000,000 tokens
Modalities
text, vision, audio, video
Languages
100+

Benchmark scores

MMLU
89.8%
Massive Multitask Language Understanding
GPQA
86.4%
Graduate-level science QA (Diamond)
HumanEval
92.6%
Python code completion
SWE-bench
63.8%
Real GitHub issue resolution (Verified)
LiveCodeBench
N/A
Contamination-free coding
AIME
88.0%
American Invitational Mathematics Exam
MMMU
82.0%
Multimodal college-level reasoning
MATH
91.5%
Competition-level math word problems

Scores as reported by the vendor or leading public leaderboards. "N/A" means the score has not been publicly disclosed for this metric.

Explain like I'm 5

Reads whole books at once and understands pictures, sound, and video together.

Key features
  • 1M token context
  • Native video understanding
  • Deep Think mode
  • Grounded search
Strengths
  • Best long-context
  • Strong multimodal reasoning
  • Fast on Google infra
Limitations
  • Occasional refusals on edge cases

Best for

Long-doc analysisVideo Q&AMultimodal agents

Related models