TL;DR
Data snapshot: July 10 - Aug 6, 2026, cross-checked across boards. The overall #1 is Anthropic's Claude Fable 5 (released Jun 9, 2026; AA Intelligence Index 64.9, GDPval-AA Elo 1932), and the LMArena text top 5 is all Claude. Coding #1 is Kimi K3 (LMArena Coding Arena 1529). By real usage, DeepSeek V4 Flash is now the world's most-called model (~84% of one day's OpenRouter tokens). Chinese models have narrowed the gap to the frontier tier to within ~5 points (SuperCLUE May 2026).
LMArena Text Arena (2026-07-10)
| Rank | Model | Score (Elo) | Notes |
|---|---|---|---|
| 1 | Claude Fable 5 | 1509±9 | Released Jun 9, 2026; API $10/$50 per M tokens |
| 2–5 | Claude family (Opus 4.8 / Opus 4.7 etc.) | — | All top 5 are Claude models; strong human-preference lead |
| China #1 (Jan 2026 snapshot) | ERNIE-5.0-0110 (Baidu) | 1460 | Global #8, China #1; math #2 globally (2026-01-15) |
Coding & General Intelligence Boards (2026-07/08)
| Board (data date) | Leader | Score / Value |
|---|---|---|
| LMArena Coding Arena (2026-07-16) | Kimi K3 (Moonshot) | Elo 1529 |
| Artificial Analysis Intelligence Index (2026-07) | Claude Fable 5 | 64.9, ~5 pts ahead of GPT-5.5; Opus 4.8 #2 |
| AA GDPval-AA (real-world tasks, 2026-07) | Claude Fable 5 | Elo 1932 (previous record: Opus 4.8) |
| SWE-bench Verified (official, Jun 2026) | Claude Fable 5 | 95.0% (Opus 4.8: 88.6%) |
| OpenRouter call volume (2026-08-06) | DeepSeek V4 Flash | #1 globally; ~6.3T of ~7.5T daily tokens (~84%) |
Change vs. Earlier Snapshot (Jan 2026 → Aug 2026)
| Change | Jan 2026 snapshot | Aug 2026 snapshot |
|---|---|---|
| LMArena text #1 | GLM-4.7; ERNIE-5.0 in top 10 | Claude Fable 5; top 5 all Claude |
| Chinese models on LMArena text | ERNIE-5.0-0110 global #8 | No Chinese model in top 10 |
| Coding board | Claude/GPT led | Kimi K3 (1529) #1 |
| Call volume #1 | — | DeepSeek V4 Flash (price-war payoff) |
| Chinese general (SuperCLUE May 2026) | — | Gemini 3.1 Pro #1; DeepSeek-V4-Pro / Qwen3.7-Max / Doubao Seed-2.0-pro within <5 pts of frontier |
How to Choose (Aug 2026)
| Use case | First pick | Alternative | Why |
|---|---|---|---|
| General intelligence / hard reasoning | Claude Fable 5 | GPT-5.5 / Opus 4.8 | #1 on AA index; dominant on long-horizon agent tasks (GDPval-AA 1932) |
| Coding / software engineering | Kimi K3 / Claude family | DeepSeek V4 | Coding arena #1; Fable 5 hits 95.0% on SWE-bench Verified |
| Chinese-language tasks | Gemini 3.1 Pro | DeepSeek V4 / Qwen3.7-Max | SuperCLUE Chinese #1; domestic gap <5 pts |
| Cost / high-volume calls | DeepSeek V4 Flash | V4 Pro | Output ¥2/M tokens (Aug 6, 2026); #1 in real call volume |
| Self-hosted / open weights | DeepSeek V4 / Qwen3.7-Max | Kimi K3 | Open weights, deployable on-prem, data stays in-country |
Takeaways
- Anthropic leads both human-preference (LMArena) and benchmark (AA) boards; Fable 5 and Mythos 5 share the same core — Mythos 5 (no safety classifiers) is limited to Project Glasswing partners;
- Chinese models are splitting roles: Kimi K3 tops coding, DeepSeek V4 Flash tops call volume, but no Chinese model currently makes the LMArena text top 10 (ERNIE-5.0 was global #8 in Jan 2026);
- Board methodology differs: LMArena = human-preference Elo, AA = benchmark index, OpenRouter = real usage, SuperCLUE = Chinese-specific. Always cite the board + data date;
- DeepSeek announced an API price increase on Aug 6 — V4 Flash's cost advantage may narrow; re-check official pricing before choosing.
FAQ
Which AI model is the strongest overall in 2026?
Per the July 2026 Artificial Analysis Intelligence Index, the overall #1 is Anthropic's Claude Fable 5 (64.9, ~5 pts ahead of GPT-5.5); it is also #1 on the LMArena text board (Elo 1509±9). For Chinese-language tasks, SuperCLUE's #1 is Gemini 3.1 Pro.
Which model is best at coding?
LMArena Coding Arena (2026-07-16) #1 is Kimi K3 (Elo 1529); on official benchmarks, Claude Fable 5 reaches 95.0% on SWE-bench Verified and 80.0% on SWE-bench Pro, best for long-horizon code engineering. Pick by task: algorithm problems → LMArena Coding; real-repo bug fixing → SWE-bench.
How good are Chinese models right now?
They are at the frontier's edge: Kimi K3 tops the coding board, DeepSeek V4 Flash tops real call volume, and SuperCLUE (May 2026) puts DeepSeek-V4-Pro, Qwen3.7-Max and Doubao Seed-2.0-pro within ~5 points of the global frontier. But no Chinese model currently ranks in the LMArena text top 10.
Why is DeepSeek V4 Flash the world's most-called model?
Cost-performance: as of Aug 6, 2026, output is only ¥2/M tokens (¥4 peak) — tens of times cheaper than Claude/GPT flagships — with near-frontier quality. On Aug 5, V4 Flash accounted for ~6.3T of ~7.5T daily OpenRouter tokens (~84%). Note DeepSeek announced a price increase on Aug 6; the new rates are pending.
Where should I check AI model rankings?
Five commonly used sources: LMArena (human-preference Elo), Artificial Analysis (benchmark index), SuperCLUE (Chinese-specific), OpenRouter (real call volume), SWE-bench (code engineering). The same model can rank differently across boards, so always cite "board + data date", e.g. "LMArena text, 2026-07-10".
Sources
- LMArena Text/Coding Arena (2026-07-10 / 2026-07-16)
- Artificial Analysis Intelligence Index & official evaluation article (Claude Fable 5, Jul 2026)
- Anthropic official model docs & launch page (Fable 5 / Mythos 5 pricing and benchmarks, Jun 9, 2026)
- Kuai Technology (ERNIE-5.0-0110, LMArena 1460, global #8, 2026-01-15)
- OpenRouter usage data via NetEase/Zhihu coverage (DeepSeek V4 Flash #1, 2026-08-06)
- SuperCLUE May 2026 Chinese LLM benchmark (SegmentFault summary, Jun 2026)
- DeepSeek official notice & STAR Market Daily (API price-hike notice, 2026-08-06)