Leaderboards
How the frontier models stack up across the benchmarks that matter. Rows are ranked by average standing; click any column to sort by that benchmark.
Last updated July 16, 2026.
| Model | LMArena Elo | SWE-bench Verified % | GPQA Diamond % | AIME 2025 % | τ-bench (Airline) % |
|---|---|---|---|---|---|
| Claude Opus 4.8 Anthropic | 1451 #1 | 76.4% #1 | 85.1% #2 | 92% #2 | 68.9% #1 |
| GPT-5.5 OpenAI | 1447 #2 | 74.1% #2 | 84% #3 | 93.1% #1 | 66% #2 |
| Gemini 3 Pro Google DeepMind | 1439 #3 | 72.5% #3 | 86.2% #1 | 90.5% #4 | 62.7% #4 |
| Grok 5 xAI | 1430 #4 | 69% #5 | 82.6% #4 | 91.2% #3 | 59.5% #5 |
| Claude Sonnet 5 Anthropic | 1421 #5 | 71.8% #4 | 81.3% #5 | 88.4% #5 | 64.2% #3 |
| Qwen3.6 235B Alibaba | 1405 #6 | 64.7% #6 | 78.4% #6 | 84.6% #6 | 54.1% #6 |
| Llama 5 405B Meta | 1398 #7 | 61.2% #7 | 76.8% #7 | 79% #7 | 51.3% #7 |
Benchmarks
- LMArena
- Human-preference Elo from anonymous pairwise battles.
- SWE-bench Verified
- Resolution rate on real-world GitHub issues.
- GPQA Diamond
- Graduate-level, Google-proof science QA.
- AIME 2025
- Competition mathematics.
- τ-bench (Airline)
- Tool-use / agentic task completion in dialogue.