5dive Pulse
Labs game the benchmarks. We don't.
Independent reasoning scores, real token usage, and an honest read on what developers actually think.
5dive model ranking
Ranked by the 5dive score: 80% the Artificial Analysis Intelligence Index, an independent capability benchmark, 20% our developer-sentiment read from Hacker News and Reddit— because labs game the benchmarks.
Open a row's for both halves of its score, the developer read and the threads it came from.
Artificial Analysis Intelligence Index v4.3 · read from Artificial Analysis 28 Sept 2026
- 1
Claude Opus 5.5new
by Anthropic · $4/$20
Middling23 opinions — a thin sample, under this board's floor of 30 · mostly on coding, price, agents and tools. - 2
Claude Fable 5.1
by Anthropic · $10/$50
↗ usage rising this weekWell received55 opinions · Tops the benchmark and devs broadly agree; the complaint is the bill, not the output. - 3
GPT-6 Astra
by OpenAI · $10/$50
↘ usage falling this weekWell received23 opinions — a thin sample, under this board's floor of 30 · mostly on price, coding, speed. - 4
GPT-6 Solnew
by OpenAI · $2/$10
Best received22 opinions — a thin sample, under this board's floor of 30▲ turning more positive than the 30-day tally · mostly on price, coding, speed. - 5
Muse Spark 1.3
by Meta · $1.3/$4.3
↗ usage rising this weekMiddling5 opinions — a thin sample, under this board's floor of 30▲ turning more positive than the 30-day tally · mostly on price, self-hosting, speed. - 6
Qwen3.8 Max
by Qwen · $2/$6
↘ usage falling this weekWell received47 opinions▲ turning more positive than the 30-day tally · Strongest multilingual read of the group; trusted more than its benchmark rank. - 7
MiMo-V2.6-Pronew
by Xiaomi · $0.4/$0.9
No developer read yet - 8
GLM-5.3
by z-ai · $0.2/$2.8
Lukewarm24 opinions — a thin sample, under this board's floor of 30▼ turning more negative than the 30-day tally · Punches above its cost; widely used as the cheap second opinion. - 9
Gemini 3.8 Flash
by Google · $0.8/$3.8
Best received13 opinions — a thin sample, under this board's floor of 30 · Huge context and genuinely fast; the recurring gripe is drift on long agent runs. - 10
Kimi K3
by Moonshot AI · $3/$15
Lukewarm23 opinions — a thin sample, under this board's floor of 30 · The open-weight favourite — rated well above its price bracket by people running it. - 11
Grok 4.7
by xAI · $1.6/$4.8
Least liked9 opinions — a thin sample, under this board's floor of 30▼ turning more negative than the 30-day tally · mostly on price, speed, coding. - 12
GPT-5.6 Sol
by OpenAI · $2/$1050% off
↘ usage falling this weekLeast liked4 opinions — a thin sample, under this board's floor of 30▼ turning more negative than the 30-day tally · Strong at code, but devs flag verbose output and refusals on benign prompts. - 13
GPT-5.6 Terra
by OpenAI · $2/$12
No developer read yet · Solid generalist; the usual split is Sol for code, Terra for everything else. - 14
Claude Opus 4.8
by Anthropic · $5/$25
↘ usage falling this weekLukewarm11 opinions — a thin sample, under this board's floor of 30▲ turning more positive than the 30-day tally · The one teams actually keep in production — boring in the way that ships. - 15
Claude Sonnet 5
by Anthropic · $2/$10
Middling5 opinions — a thin sample, under this board's floor of 30 · Best speed-to-quality ratio in the line; the default choice inside agent loops. - 16
DeepSeek V4.1 Flash
by DeepSeek · $0.3/$1.2
↗ usage rising this weekLukewarm7 opinions — a thin sample, under this board's floor of 30 · mostly on price, coding, speed. - 17
GLM 5.3 Flash
by z-ai · $0.2/$0.5
↗ usage rising this weekLeast liked34 opinions▼ turning more negative than the 30-day tally · Very cheap and surprisingly coherent — the batch-job default for a lot of teams. - 18
GPT-6 Lunanew
by OpenAI · $0.1/$0.5
Middling15 opinions — a thin sample, under this board's floor of 30 · mostly on price, coding, agents and tools. - 19
Qwen3.8 27B
by Qwen · $0.4/$3
↗ usage rising this weekBest received29 opinions — a thin sample, under this board's floor of 30 · The size devs actually self-host — it runs where the flagships simply cannot. - 20
MiMo-V2.6-Flashnew
by Xiaomi · $0.1/$0.3
No developer read yet
Glass-box methodology
Four views. Nothing hidden.
One blended 5dive score, and each of the three signals under it on its own. Same models every time — a tab re-sorts the board, it never filters it. All live, no vendor marketing.
What the 5dive score is made of
80% the Artificial Analysis Intelligence Index — an independent capability benchmark, not a pure reasoning one: it weighs agentic work, coding and general knowledge alongside scientific reasoning — and 20% the developer-sentiment read. Open a row's (i) to see both parts and the arithmetic between them. A model we have no read on yet is still ranked, with its sentiment held at the median of the reads we have collected rather than dropped or scored on the benchmark alone — so it lands exactly where a model with a typical read and the same benchmark would, and our silence costs it nothing. Its (i) names the number it was anchored on. A fixed midpoint would not do that: the reads we collect sit well above the middle of the scale, so anchoring there would quietly rank every unread model last.
Beside each row sits the collected tone, our own one-line take on the model, and where we have it the real 7-day direction of token usage. The direction is context; it is not part of the score.
Where the intelligence score comes from
The benchmark half is the Artificial Analysis Intelligence Index v4.3, and we now read it from Artificial Analysis itself, not from a copy. It used to come from OpenRouter's mirror of the same index, and a mirror lags: for several days this board ranked Muse Spark 1.2 while Artificial Analysis had already published 1.3, and refreshing more often could not fix it, because our copy was current and the source of the copy was not.
OpenRouter still answers for everything OpenRouter owns — the model catalog, prices, and the real-usage board. If Artificial Analysis does not answer a refresh, the scores fall back to OpenRouter's copy of their index and the line above the board says so.
Intelligence scores by Artificial Analysis · Artificial Analysis (2025). LLM benchmarks dataset. https://artificialanalysis.ai · Artificial Analysis terms of use ↗
This read: 24 models scored from Artificial Analysis, 1 of them missing from OpenRouter's copy. 3 more are scored by Artificial Analysis but not listed on OpenRouter at all, so we cannot price them and do not show them.
Where the developer-sentiment read comes from
The score and the tone are collected from public developer discussion. Every comment in the window that names a model is counted once, as positive or negative or neither, and the score is the balance of those votes. One comment, one vote — so a heavily-discussed model and a quiet one sit on the same scale. A row's (i) carries the tally and links every thread it was read from.
The headline is the last 7 days, with the 30-day tally kept beside it and the direction between the two marked on the row. A single flat month cannot show a turn: a model's launch-week praise is thirty days of denominator, against which a two-day backlash barely moves the number. Where the recent window holds too few comments to clear the same floors, the row falls back to the longer one and its (i) says so — a thin recent sample is not promoted to a headline, and a model that has a read is never handed the median instead.
Every rated row prints how many opinionated comments it rests on — in amber below 30, which is thin. Most comments that name a model express no view; the ones that do are the real sample, and where there are only a handful of them a single commenter moves the score by several points. The number still stands and the row still ranks — we mark it rather than hide it.
The label on each row is comparative — how the model landed against the others here, not against an absolute bar. Developers write “fast” and “cheap” about almost anything they bothered to try, so an absolute cut puts nearly every model in the top band and tells you nothing. The order is what the comments actually support; the level is not.
The sentence after the tone is ours, not the commenters'. It is our own take on the model, written by hand for 39 of them and not derived from anything above; models without one show the subjects the comments actually kept returning to. We would rather say which half is which than blur them.
Captured 28 Sept 2026 · headline last 7 days, 30-day tally beside it · Hacker News, Reddit · 17 models read, 3 not yet · floor 5 comments.
Overall
The 5dive score
80% benchmark, 20% developer read. The board's default order, and the only view where the two are mixed. Benchmark and pricing refresh hourly; the sentiment half is collected daily.
Sentiment
What developers say
The collected read on its own, 0–100, from Hacker News and Reddit, captured 28 Sept 2026. A model we have no read on shows “—” and ranks last here rather than being dropped.
Intelligence
The benchmark, unmixed
Intelligence scores by Artificial Analysis (Intelligence Index v4.3), with no sentiment in it. The line above the board names which source answered this run and when we read it.
Usage
Real token volume
The tokens developers actually spend on each model, from OpenRouter's public rankings. Batch and free tiers are excluded. This is the one view with its own row set, so it lists what OpenRouter reports rather than the benchmark's slice.
Token maxxing
Most value per subscription $
API-equivalent token value per dollar of list price. Raw value, no quality adjustment.
- 1
Claude Pro
$20/mo · Claude Opus 5
~20× - 2
ChatGPT Plus
$20/mo · GPT-5.6 Sol
~17× - 3
Claude Max (5×)
$100/mo · Claude Opus 5
~17× - 4
ChatGPT Pro (5×)
$100/mo · GPT-5.6 Sol
~16× - 5
ChatGPT Pro (20×)
$200/mo · GPT-5.6 Sol
~16×
Data provenance
Intelligence scores by Artificial Analysis— an independent third party, read from their own index. Usage and pricing are real token volume and list prices, live from OpenRouter.