All benchmarks
Anthropic · USA · Released 2025-09

Claude Sonnet 4.5

Anthropic's best coding + agentic model, purpose-built for long-horizon computer-use and software engineering tasks.

Access
closed
Context
200,000 tokens
Modalities
text, vision
Languages
95+

Benchmark scores

MMLU
88.5%
Massive Multitask Language Understanding
GPQA
83.4%
Graduate-level science QA (Diamond)
HumanEval
95.4%
Python code completion
SWE-bench
77.2%
Real GitHub issue resolution (Verified)
LiveCodeBench
N/A
Contamination-free coding
AIME
87.0%
American Invitational Mathematics Exam
MMMU
77.8%
Multimodal college-level reasoning
MATH
91.0%
Competition-level math word problems

Scores as reported by the vendor or leading public leaderboards. "N/A" means the score has not been publicly disclosed for this metric.

Explain like I'm 5

Great at writing code and running many steps in a row without getting lost.

Key features
  • Extended thinking
  • Computer use
  • Code editing tools
  • Long agent runs (>30h)
Strengths
  • Best-in-class SWE-bench
  • Superb long-context coding
  • Reliable tool use
Limitations
  • No native audio
  • Closed source

Best for

Autonomous coding agentsRefactoring large codebasesAnalyst workflows

Related models