All benchmarks
Alibaba Qwen · China · Released 2025-09

Qwen3-Max

Alibaba's flagship trillion-parameter Qwen3 tier — strong multilingual, coding and agentic performance.

Access
open-weights
Context
262,144 tokens
Modalities
text, vision
Languages
100+

Benchmark scores

MMLU
86.0%
Massive Multitask Language Understanding
GPQA
70.0%
Graduate-level science QA (Diamond)
HumanEval
90.2%
Python code completion
SWE-bench
N/A
Real GitHub issue resolution (Verified)
LiveCodeBench
N/A
Contamination-free coding
AIME
85.0%
American Invitational Mathematics Exam
MMMU
N/A
Multimodal college-level reasoning
MATH
88.0%
Competition-level math word problems

Scores as reported by the vendor or leading public leaderboards. "N/A" means the score has not been publicly disclosed for this metric.

Explain like I'm 5

A very big open model that speaks tons of languages and codes well.

Key features
  • MoE trillion-scale
  • Vision variant
  • Agent tools
  • Multilingual
Strengths
  • Multilingual coverage
  • Open ecosystem
  • Strong coding
Limitations
  • Vision only in specific variants

Best for

Multilingual assistantsEnterprise RAGAgentic apps

Related models