🏃 Keep Up

Leaderboards

How the frontier models stack up across the benchmarks that matter. Rows are ranked by average standing; click any column to sort by that benchmark.

Last updated July 16, 2026.

Model LMArena Elo SWE-bench Verified % GPQA Diamond % AIME 2025 % τ-bench (Airline) %
Claude Opus 4.8
Anthropic
1451 #1 76.4% #1 85.1% #2 92% #2 68.9% #1
GPT-5.5
OpenAI
1447 #2 74.1% #2 84% #3 93.1% #1 66% #2
Gemini 3 Pro
Google DeepMind
1439 #3 72.5% #3 86.2% #1 90.5% #4 62.7% #4
Grok 5
xAI
1430 #4 69% #5 82.6% #4 91.2% #3 59.5% #5
Claude Sonnet 5
Anthropic
1421 #5 71.8% #4 81.3% #5 88.4% #5 64.2% #3
Qwen3.6 235B
Alibaba
1405 #6 64.7% #6 78.4% #6 84.6% #6 54.1% #6
Llama 5 405B
Meta
1398 #7 61.2% #7 76.8% #7 79% #7 51.3% #7

Benchmarks

LMArena
Human-preference Elo from anonymous pairwise battles.
SWE-bench Verified
Resolution rate on real-world GitHub issues.
GPQA Diamond
Graduate-level, Google-proof science QA.
AIME 2025
Competition mathematics.
τ-bench (Airline)
Tool-use / agentic task completion in dialogue.