All benchmarks
Microsoft · USA · Released 2024-12

Phi-4

Microsoft's 14B small language model trained on curated + synthetic data — punches far above its weight on reasoning.

Access
open-weights
Context
16,000 tokens
Modalities
text
Languages
English-focused

Benchmark scores

MMLU
84.8%
Massive Multitask Language Understanding
GPQA
56.1%
Graduate-level science QA (Diamond)
HumanEval
82.6%
Python code completion
SWE-bench
N/A
Real GitHub issue resolution (Verified)
LiveCodeBench
N/A
Contamination-free coding
AIME
N/A
American Invitational Mathematics Exam
MMMU
N/A
Multimodal college-level reasoning
MATH
80.4%
Competition-level math word problems

Scores as reported by the vendor or leading public leaderboards. "N/A" means the score has not been publicly disclosed for this metric.

Explain like I'm 5

Small but surprisingly smart — runs on modest hardware.

Key features
  • 14B params
  • Synthetic-data training
  • Open weights
Strengths
  • Strong reasoning for size
  • Cheap to run
Limitations
  • Small context
  • English-focused

Best for

Edge deploymentFine-tuning base

Related models