intelligence, usage & sentiment · refreshed hourlytoken maxxing: value per $ →

5dive Pulse

Labs game the benchmarks. We don't.

Independent reasoning scores, real token usage, and an honest read on what developers actually think.

5dive model ranking

Ranked by the 5dive score: 80% the Artificial Analysis Intelligence Index, an independent capability benchmark, 20% our developer-sentiment read from Hacker News and Reddit— because labs game the benchmarks.

Open a row's for both halves of its score, the developer read and the threads it came from.

Artificial Analysis Intelligence Index v4.3 · read from Artificial Analysis 28 Sept 2026

  1. 1

    Claude Opus 5.5new

    by Anthropic · $4/$20

    Middling23 opinions — a thin sample, under this board's floor of 30 · mostly on coding, price, agents and tools.

    Claude Opus 5.5

    rank #1 · Anthropic

    62.75dive score
    Score split
    AA 57.6 · sentiment 83 → 0.8×57.6 + 0.2×83
    Developer read
    Middling · 74 comments read over the last 7 days, compared with the other models on this board
    Window
    Last 7 days: 83 from 74 comments · 30-day tally: 83 from 74 · in line with the longer window
    Tally
    Clearly positive — 20 positive to 3 negative of 74 recent Hacker News and Reddit comments in the last 7 days; recurring subjects: coding, price, agents and tools.
    Subjects
    coding, price, agents and tools
    Captured
    28 Sept 2026
    Price
    $4 in / $20 out per 1M tokens
    Added
    in the last 7 days

    The read is comparative — how this model landed against the others on this board, not against an absolute bar. The tally under it is the raw count it was derived from.

  2. 2

    Claude Fable 5.1

    by Anthropic · $10/$50

    ↗ usage rising this weekWell received55 opinions · Tops the benchmark and devs broadly agree; the complaint is the bill, not the output.

    Claude Fable 5.1

    rank #2 · Anthropic

    59.55dive score
    Score split
    AA 53.4 · sentiment 84 → 0.8×53.4 + 0.2×84
    Developer read
    Well received · 114 comments read over the last 7 days, compared with the other models on this board
    Window
    Last 7 days: 84 from 114 comments · 30-day tally: 80 from 377 · in line with the longer window
    Tally
    Clearly positive — 46 positive to 9 negative of 114 recent Hacker News and Reddit comments in the last 7 days; recurring subjects: price, coding, speed.
    Subjects
    price, coding, speed
    Captured
    28 Sept 2026
    Trajectory
    ↗ usage rising this week
    Price
    $10 in / $50 out per 1M tokens

    The read is comparative — how this model landed against the others on this board, not against an absolute bar. The tally under it is the raw count it was derived from.

  3. 3

    GPT-6 Astra

    by OpenAI · $10/$50

    ↘ usage falling this weekWell received23 opinions — a thin sample, under this board's floor of 30 · mostly on price, coding, speed.

    GPT-6 Astra

    rank #3 · OpenAI

    59.25dive score
    Score split
    AA 52.7 · sentiment 85 → 0.8×52.7 + 0.2×85
    Developer read
    Well received · 75 comments read over the last 7 days, compared with the other models on this board
    Window
    Last 7 days: 85 from 75 comments · 30-day tally: 88 from 403 · in line with the longer window
    Tally
    Clearly positive — 19 positive to 4 negative of 75 recent Hacker News and Reddit comments in the last 7 days; recurring subjects: price, coding, speed.
    Subjects
    price, coding, speed
    Captured
    28 Sept 2026
    Trajectory
    ↘ usage falling this week
    Price
    $10 in / $50 out per 1M tokens

    The read is comparative — how this model landed against the others on this board, not against an absolute bar. The tally under it is the raw count it was derived from.

  4. 4

    GPT-6 Solnew

    by OpenAI · $2/$10

    Best received22 opinions — a thin sample, under this board's floor of 30▲ turning more positive than the 30-day tally · mostly on price, coding, speed.

    GPT-6 Sol

    rank #4 · OpenAI

    56.05dive score
    Score split
    AA 47.5 · sentiment 90 → 0.8×47.5 + 0.2×90
    Developer read
    Best received · 55 comments read over the last 7 days, compared with the other models on this board
    Window
    Last 7 days: 90 from 55 comments · 30-day tally: 84 from 61 · turning more positive than the longer window
    Tally
    Clearly positive — 20 positive to 2 negative of 55 recent Hacker News and Reddit comments in the last 7 days; recurring subjects: price, coding, speed.
    Subjects
    price, coding, speed
    Captured
    28 Sept 2026
    Price
    $2 in / $10 out per 1M tokens
    Added
    in the last 7 days

    The read is comparative — how this model landed against the others on this board, not against an absolute bar. The tally under it is the raw count it was derived from.

  5. 5

    Muse Spark 1.3

    by Meta · $1.3/$4.3

    ↗ usage rising this weekMiddling5 opinions — a thin sample, under this board's floor of 30▲ turning more positive than the 30-day tally · mostly on price, self-hosting, speed.

    Muse Spark 1.3

    rank #5 · Meta

    55.15dive score
    Score split
    AA 48.1 · sentiment 83 → 0.8×48.1 + 0.2×83
    Developer read
    Middling · 10 comments read over the last 7 days, compared with the other models on this board
    Window
    Last 7 days: 83 from 10 comments · 30-day tally: 77 from 45 · turning more positive than the longer window
    Tally
    Clearly positive — 5 positive to 0 negative of 10 recent Hacker News and Reddit comments in the last 7 days; recurring subjects: price, self-hosting, speed.
    Subjects
    price, self-hosting, speed
    Captured
    28 Sept 2026
    Trajectory
    ↗ usage rising this week
    Price
    $1.3 in / $4.3 out per 1M tokens

    The read is comparative — how this model landed against the others on this board, not against an absolute bar. The tally under it is the raw count it was derived from.

  6. 6

    Qwen3.8 Max

    by Qwen · $2/$6

    ↘ usage falling this weekWell received47 opinions▲ turning more positive than the 30-day tally · Strongest multilingual read of the group; trusted more than its benchmark rank.

    Qwen3.8 Max

    rank #6 · Qwen

    53.35dive score
    Score split
    AA 45.4 · sentiment 85 → 0.8×45.4 + 0.2×85
    Developer read
    Well received · 108 comments read over the last 7 days, compared with the other models on this board
    Window
    Last 7 days: 85 from 108 comments · 30-day tally: 79 from 348 · turning more positive than the longer window
    Tally
    Clearly positive — 39 positive to 8 negative of 108 recent Hacker News and Reddit comments in the last 7 days; recurring subjects: self-hosting, speed, coding.
    Subjects
    self-hosting, speed, coding
    Captured
    28 Sept 2026
    Trajectory
    ↘ usage falling this week
    Price
    $2 in / $6 out per 1M tokens

    The read is comparative — how this model landed against the others on this board, not against an absolute bar. The tally under it is the raw count it was derived from.

  7. 7

    MiMo-V2.6-Pronew

    by Xiaomi · $0.4/$0.9

    No developer read yet

    MiMo-V2.6-Pro

    rank #7 · Xiaomi

    51.85dive score
    Score split
    AA 46.3 · sentiment 74 (our median collected read — we have no read of this model yet) → 0.8×46.3 + 0.2×74
    Developer read
    7 mentions but only 1 expressed an opinion (floor is 4)
    Price
    $0.4 in / $0.9 out per 1M tokens
    Added
    in the last 7 days
  8. 8

    GLM-5.3

    by z-ai · $0.2/$2.8

    Lukewarm24 opinions — a thin sample, under this board's floor of 30▼ turning more negative than the 30-day tally · Punches above its cost; widely used as the cheap second opinion.

    GLM-5.3

    rank #8 · z-ai

    50.65dive score
    Score split
    AA 44.8 · sentiment 74 → 0.8×44.8 + 0.2×74
    Developer read
    Lukewarm · 70 comments read over the last 7 days, compared with the other models on this board
    Window
    Last 7 days: 74 from 70 comments · 30-day tally: 80 from 336 · turning more negative than the longer window
    Tally
    Clearly positive — 20 positive to 4 negative of 70 recent Hacker News and Reddit comments in the last 7 days; recurring subjects: coding, price, self-hosting.
    Subjects
    coding, price, self-hosting
    Captured
    28 Sept 2026
    Price
    $0.2 in / $2.8 out per 1M tokens

    The read is comparative — how this model landed against the others on this board, not against an absolute bar. The tally under it is the raw count it was derived from.

  9. 9

    Gemini 3.8 Flash

    by Google · $0.8/$3.8

    Best received13 opinions — a thin sample, under this board's floor of 30 · Huge context and genuinely fast; the recurring gripe is drift on long agent runs.

    Gemini 3.8 Flash

    rank #9 · Google

    50.55dive score
    Score split
    AA 40.9 · sentiment 89 → 0.8×40.9 + 0.2×89
    Developer read
    Best received · 39 comments read over the last 7 days, compared with the other models on this board
    Window
    Last 7 days: 89 from 39 comments · 30-day tally: 89 from 194 · in line with the longer window
    Tally
    Clearly positive — 13 positive to 0 negative of 39 recent Hacker News and Reddit comments in the last 7 days; recurring subjects: price, coding, reasoning.
    Subjects
    price, coding, reasoning
    Captured
    28 Sept 2026
    Price
    $0.8 in / $3.8 out per 1M tokens

    The read is comparative — how this model landed against the others on this board, not against an absolute bar. The tally under it is the raw count it was derived from.

  10. 10

    Kimi K3

    by Moonshot AI · $3/$15

    Lukewarm23 opinions — a thin sample, under this board's floor of 30 · The open-weight favourite — rated well above its price bracket by people running it.

    Kimi K3

    rank #10 · Moonshot AI

    50.55dive score
    Score split
    AA 43.6 · sentiment 78 → 0.8×43.6 + 0.2×78
    Developer read
    Lukewarm · 60 comments read over the last 7 days, compared with the other models on this board
    Window
    Last 7 days: 78 from 60 comments · 30-day tally: 80 from 230 · in line with the longer window
    Tally
    Clearly positive — 20 positive to 3 negative of 60 recent Hacker News and Reddit comments in the last 7 days; recurring subjects: price, coding, speed.
    Subjects
    price, coding, speed
    Captured
    28 Sept 2026
    Price
    $3 in / $15 out per 1M tokens

    The read is comparative — how this model landed against the others on this board, not against an absolute bar. The tally under it is the raw count it was derived from.

  11. 11

    Grok 4.7

    by xAI · $1.6/$4.8

    Least liked9 opinions — a thin sample, under this board's floor of 30▼ turning more negative than the 30-day tally · mostly on price, speed, coding.

    Grok 4.7

    rank #11 · xAI

    49.55dive score
    Score split
    AA 46.4 · sentiment 62 → 0.8×46.4 + 0.2×62
    Developer read
    Least liked · 23 comments read over the last 7 days, compared with the other models on this board
    Window
    Last 7 days: 62 from 23 comments · 30-day tally: 74 from 29 · turning more negative than the longer window
    Tally
    Broadly positive — 6 positive to 3 negative of 23 recent Hacker News and Reddit comments in the last 7 days; recurring subjects: price, speed, coding.
    Subjects
    price, speed, coding
    Captured
    28 Sept 2026
    Price
    $1.6 in / $4.8 out per 1M tokens

    The read is comparative — how this model landed against the others on this board, not against an absolute bar. The tally under it is the raw count it was derived from.

  12. 12

    GPT-5.6 Sol

    by OpenAI · $2/$1050% off

    ↘ usage falling this weekLeast liked4 opinions — a thin sample, under this board's floor of 30▼ turning more negative than the 30-day tally · Strong at code, but devs flag verbose output and refusals on benign prompts.

    GPT-5.6 Sol

    rank #12 · OpenAI

    48.65dive score
    Score split
    AA 47.0 · sentiment 55 → 0.8×47.0 + 0.2×55
    Developer read
    Least liked · 19 comments read over the last 7 days, compared with the other models on this board
    Window
    Last 7 days: 55 from 19 comments · 30-day tally: 84 from 113 · turning more negative than the longer window
    Tally
    Mixed — 3 positive to 1 negative of 19 recent Hacker News and Reddit comments in the last 7 days; recurring subjects: price, coding, reasoning.
    Subjects
    price, coding, reasoning
    Captured
    28 Sept 2026
    Trajectory
    ↘ usage falling this week
    Price
    $2 in / $10 out per 1M tokens50% off — a promotional price, not this model's list price

    The read is comparative — how this model landed against the others on this board, not against an absolute bar. The tally under it is the raw count it was derived from.

  13. 13

    GPT-5.6 Terra

    by OpenAI · $2/$12

    No developer read yet · Solid generalist; the usual split is Sol for code, Terra for everything else.

    GPT-5.6 Terra

    rank #13 · OpenAI

    48.55dive score
    Score split
    AA 42.1 · sentiment 74 (our median collected read — we have no read of this model yet) → 0.8×42.1 + 0.2×74
    Developer read
    9 mentions but only 2 expressed an opinion (floor is 4)
    Our take
    Solid generalist; the usual split is Sol for code, Terra for everything else.
    Price
    $2 in / $12 out per 1M tokens
  14. 14

    Claude Opus 4.8

    by Anthropic · $5/$25

    ↘ usage falling this weekLukewarm11 opinions — a thin sample, under this board's floor of 30▲ turning more positive than the 30-day tally · The one teams actually keep in production — boring in the way that ships.

    Claude Opus 4.8

    rank #14 · Anthropic

    47.25dive score
    Score split
    AA 41.8 · sentiment 69 → 0.8×41.8 + 0.2×69
    Developer read
    Lukewarm · 16 comments read over the last 7 days, compared with the other models on this board
    Window
    Last 7 days: 69 from 16 comments · 30-day tally: 61 from 127 · turning more positive than the longer window
    Tally
    Broadly positive — 8 positive to 3 negative of 16 recent Hacker News and Reddit comments in the last 7 days; recurring subjects: coding, price, self-hosting.
    Subjects
    coding, price, self-hosting
    Captured
    28 Sept 2026
    Trajectory
    ↘ usage falling this week
    Price
    $5 in / $25 out per 1M tokens

    The read is comparative — how this model landed against the others on this board, not against an absolute bar. The tally under it is the raw count it was derived from.

  15. 15

    Claude Sonnet 5

    by Anthropic · $2/$10

    Middling5 opinions — a thin sample, under this board's floor of 30 · Best speed-to-quality ratio in the line; the default choice inside agent loops.

    Claude Sonnet 5

    rank #15 · Anthropic

    47.25dive score
    Score split
    AA 38.2 · sentiment 83 → 0.8×38.2 + 0.2×83
    Developer read
    Middling · 18 comments read over the last 7 days, compared with the other models on this board
    Window
    Last 7 days: 83 from 18 comments · 30-day tally: 81 from 69 · in line with the longer window
    Tally
    Clearly positive — 5 positive to 0 negative of 18 recent Hacker News and Reddit comments in the last 7 days; recurring subjects: coding, price, speed.
    Subjects
    coding, price, speed
    Captured
    28 Sept 2026
    Price
    $2 in / $10 out per 1M tokens

    The read is comparative — how this model landed against the others on this board, not against an absolute bar. The tally under it is the raw count it was derived from.

  16. 16

    DeepSeek V4.1 Flash

    by DeepSeek · $0.3/$1.2

    ↗ usage rising this weekLukewarm7 opinions — a thin sample, under this board's floor of 30 · mostly on price, coding, speed.

    DeepSeek V4.1 Flash

    rank #16 · DeepSeek

    46.05dive score
    Score split
    AA 39.5 · sentiment 72 → 0.8×39.5 + 0.2×72
    Developer read
    Lukewarm · 16 comments read over the last 7 days, compared with the other models on this board
    Window
    Last 7 days: 72 from 16 comments · 30-day tally: 76 from 54 · in line with the longer window
    Tally
    Clearly positive — 6 positive to 1 negative of 16 recent Hacker News and Reddit comments in the last 7 days; recurring subjects: price, coding, speed.
    Subjects
    price, coding, speed
    Captured
    28 Sept 2026
    Trajectory
    ↗ usage rising this week
    Price
    $0.3 in / $1.2 out per 1M tokens

    The read is comparative — how this model landed against the others on this board, not against an absolute bar. The tally under it is the raw count it was derived from.

  17. 17

    GLM 5.3 Flash

    by z-ai · $0.2/$0.5

    ↗ usage rising this weekLeast liked34 opinions▼ turning more negative than the 30-day tally · Very cheap and surprisingly coherent — the batch-job default for a lot of teams.

    GLM 5.3 Flash

    rank #17 · z-ai

    45.85dive score
    Score split
    AA 41.8 · sentiment 62 → 0.8×41.8 + 0.2×62
    Developer read
    Least liked · 69 comments read over the last 7 days, compared with the other models on this board
    Window
    Last 7 days: 62 from 69 comments · 30-day tally: 84 from 292 · turning more negative than the longer window
    Tally
    Broadly positive — 21 positive to 13 negative of 69 recent Hacker News and Reddit comments in the last 7 days; recurring subjects: price, speed, coding.
    Subjects
    price, speed, coding
    Captured
    28 Sept 2026
    Trajectory
    ↗ usage rising this week
    Price
    $0.2 in / $0.5 out per 1M tokens

    The read is comparative — how this model landed against the others on this board, not against an absolute bar. The tally under it is the raw count it was derived from.

  18. 18

    GPT-6 Lunanew

    by OpenAI · $0.1/$0.5

    Middling15 opinions — a thin sample, under this board's floor of 30 · mostly on price, coding, agents and tools.

    GPT-6 Luna

    rank #18 · OpenAI

    45.65dive score
    Score split
    AA 37.3 · sentiment 79 → 0.8×37.3 + 0.2×79
    Developer read
    Middling · 35 comments read over the last 7 days, compared with the other models on this board
    Window
    Last 7 days: 79 from 35 comments · 30-day tally: 79 from 35 · in line with the longer window
    Tally
    Clearly positive — 13 positive to 2 negative of 35 recent Hacker News and Reddit comments in the last 7 days; recurring subjects: price, coding, agents and tools.
    Subjects
    price, coding, agents and tools
    Captured
    28 Sept 2026
    Price
    $0.1 in / $0.5 out per 1M tokens
    Added
    in the last 7 days

    The read is comparative — how this model landed against the others on this board, not against an absolute bar. The tally under it is the raw count it was derived from.

  19. 19

    Qwen3.8 27B

    by Qwen · $0.4/$3

    ↗ usage rising this weekBest received29 opinions — a thin sample, under this board's floor of 30 · The size devs actually self-host — it runs where the flagships simply cannot.

    Qwen3.8 27B

    rank #19 · Qwen

    45.25dive score
    Score split
    AA 33.7 · sentiment 91 → 0.8×33.7 + 0.2×91
    Developer read
    Best received · 56 comments read over the last 7 days, compared with the other models on this board
    Window
    Last 7 days: 91 from 56 comments · 30-day tally: 87 from 214 · in line with the longer window
    Tally
    Clearly positive — 27 positive to 2 negative of 56 recent Hacker News and Reddit comments in the last 7 days; recurring subjects: self-hosting, speed, coding.
    Subjects
    self-hosting, speed, coding
    Captured
    28 Sept 2026
    Trajectory
    ↗ usage rising this week
    Price
    $0.4 in / $3 out per 1M tokens

    The read is comparative — how this model landed against the others on this board, not against an absolute bar. The tally under it is the raw count it was derived from.

  20. 20

    MiMo-V2.6-Flashnew

    by Xiaomi · $0.1/$0.3

    No developer read yet

    MiMo-V2.6-Flash

    rank #20 · Xiaomi

    45.15dive score
    Score split
    AA 37.9 · sentiment 74 (our median collected read — we have no read of this model yet) → 0.8×37.9 + 0.2×74
    Developer read
    10 mentions but only 3 expressed an opinion (floor is 4)
    Price
    $0.1 in / $0.3 out per 1M tokens
    Added
    in the last 7 days

Glass-box methodology

Four views. Nothing hidden.

One blended 5dive score, and each of the three signals under it on its own. Same models every time — a tab re-sorts the board, it never filters it. All live, no vendor marketing.

What the 5dive score is made of

80% the Artificial Analysis Intelligence Index — an independent capability benchmark, not a pure reasoning one: it weighs agentic work, coding and general knowledge alongside scientific reasoning — and 20% the developer-sentiment read. Open a row's (i) to see both parts and the arithmetic between them. A model we have no read on yet is still ranked, with its sentiment held at the median of the reads we have collected rather than dropped or scored on the benchmark alone — so it lands exactly where a model with a typical read and the same benchmark would, and our silence costs it nothing. Its (i) names the number it was anchored on. A fixed midpoint would not do that: the reads we collect sit well above the middle of the scale, so anchoring there would quietly rank every unread model last.

Beside each row sits the collected tone, our own one-line take on the model, and where we have it the real 7-day direction of token usage. The direction is context; it is not part of the score.

Where the intelligence score comes from

The benchmark half is the Artificial Analysis Intelligence Index v4.3, and we now read it from Artificial Analysis itself, not from a copy. It used to come from OpenRouter's mirror of the same index, and a mirror lags: for several days this board ranked Muse Spark 1.2 while Artificial Analysis had already published 1.3, and refreshing more often could not fix it, because our copy was current and the source of the copy was not.

OpenRouter still answers for everything OpenRouter owns — the model catalog, prices, and the real-usage board. If Artificial Analysis does not answer a refresh, the scores fall back to OpenRouter's copy of their index and the line above the board says so.

Intelligence scores by Artificial Analysis · Artificial Analysis (2025). LLM benchmarks dataset. https://artificialanalysis.ai · Artificial Analysis terms of use ↗

This read: 24 models scored from Artificial Analysis, 1 of them missing from OpenRouter's copy. 3 more are scored by Artificial Analysis but not listed on OpenRouter at all, so we cannot price them and do not show them.

Where the developer-sentiment read comes from

The score and the tone are collected from public developer discussion. Every comment in the window that names a model is counted once, as positive or negative or neither, and the score is the balance of those votes. One comment, one vote — so a heavily-discussed model and a quiet one sit on the same scale. A row's (i) carries the tally and links every thread it was read from.

The headline is the last 7 days, with the 30-day tally kept beside it and the direction between the two marked on the row. A single flat month cannot show a turn: a model's launch-week praise is thirty days of denominator, against which a two-day backlash barely moves the number. Where the recent window holds too few comments to clear the same floors, the row falls back to the longer one and its (i) says so — a thin recent sample is not promoted to a headline, and a model that has a read is never handed the median instead.

Every rated row prints how many opinionated comments it rests on — in amber below 30, which is thin. Most comments that name a model express no view; the ones that do are the real sample, and where there are only a handful of them a single commenter moves the score by several points. The number still stands and the row still ranks — we mark it rather than hide it.

The label on each row is comparative — how the model landed against the others here, not against an absolute bar. Developers write “fast” and “cheap” about almost anything they bothered to try, so an absolute cut puts nearly every model in the top band and tells you nothing. The order is what the comments actually support; the level is not.

The sentence after the tone is ours, not the commenters'. It is our own take on the model, written by hand for 39 of them and not derived from anything above; models without one show the subjects the comments actually kept returning to. We would rather say which half is which than blur them.

Captured 28 Sept 2026 · headline last 7 days, 30-day tally beside it · Hacker News, Reddit · 17 models read, 3 not yet · floor 5 comments.

Claude Opus 5.5
83 · 74 comments · 23 opinionated (thin)thread 1↗thread 2↗thread 3↗thread 4↗
Claude Fable 5.1
84 · 114 comments · 55 opinionatedthread 1↗thread 2↗thread 3↗thread 4↗
GPT-6 Astra
85 · 75 comments · 23 opinionated (thin)thread 1↗thread 2↗thread 3↗thread 4↗
GPT-6 Sol
90 · 55 comments · 22 opinionated (thin)thread 1↗thread 2↗thread 3↗thread 4↗
Muse Spark 1.3
83 · 10 comments · 5 opinionated (thin)thread 1↗thread 2↗thread 3↗thread 4↗
Qwen3.8 Max
85 · 108 comments · 47 opinionatedthread 1↗thread 2↗thread 3↗thread 4↗
GLM-5.3
74 · 70 comments · 24 opinionated (thin)thread 1↗thread 2↗thread 3↗thread 4↗
Gemini 3.8 Flash
89 · 39 comments · 13 opinionated (thin)thread 1↗thread 2↗thread 3↗thread 4↗
Kimi K3
78 · 60 comments · 23 opinionated (thin)thread 1↗thread 2↗thread 3↗thread 4↗
Grok 4.7
62 · 23 comments · 9 opinionated (thin)thread 1↗thread 2↗thread 3↗thread 4↗
GPT-5.6 Sol
55 · 19 comments · 4 opinionated (thin)thread 1↗thread 2↗thread 3↗thread 4↗
Claude Opus 4.8
69 · 16 comments · 11 opinionated (thin)thread 1↗thread 2↗thread 3↗thread 4↗
Claude Sonnet 5
83 · 18 comments · 5 opinionated (thin)thread 1↗thread 2↗thread 3↗thread 4↗
DeepSeek V4.1 Flash
72 · 16 comments · 7 opinionated (thin)thread 1↗thread 2↗thread 3↗thread 4↗
GLM 5.3 Flash
62 · 69 comments · 34 opinionatedthread 1↗thread 2↗thread 3↗thread 4↗
GPT-6 Luna
79 · 35 comments · 15 opinionated (thin)thread 1↗thread 2↗thread 3↗thread 4↗
Qwen3.8 27B
91 · 56 comments · 29 opinionated (thin)thread 1↗thread 2↗thread 3↗thread 4↗

Overall

The 5dive score

80% benchmark, 20% developer read. The board's default order, and the only view where the two are mixed. Benchmark and pricing refresh hourly; the sentiment half is collected daily.

Sentiment

What developers say

The collected read on its own, 0–100, from Hacker News and Reddit, captured 28 Sept 2026. A model we have no read on shows “—” and ranks last here rather than being dropped.

Intelligence

The benchmark, unmixed

Intelligence scores by Artificial Analysis (Intelligence Index v4.3), with no sentiment in it. The line above the board names which source answered this run and when we read it.

Usage

Real token volume

The tokens developers actually spend on each model, from OpenRouter's public rankings. Batch and free tiers are excluded. This is the one view with its own row set, so it lists what OpenRouter reports rather than the benchmark's slice.

Token maxxing

Most value per subscription $

API-equivalent token value per dollar of list price. Raw value, no quality adjustment.

view more →
  1. 1

    Claude Pro

    $20/mo · Claude Opus 5

    ~20×

    Claude Pro

    rank #1 · Claude Opus 5

    ~20×value per $ (headline)
    Maxxed range
    20×
    Evidence
    Modelled · n=0
    Time-sensitive
    TEMPORARY, THROUGH 2026-09-13: Anthropic users report a +50% weekly-limit boost running to that date, so any Claude measurement taken inside that window overstates what the plan sustains afterwards. This row is MODELLED and does not incorporate the boost, and the community figures it cites were logged before it — but treat any Claude number you read anywhere this week, including your own, as inflated until the cut lands.
    Binding limit
    weekly + 5h windowA rolling 5-hour window plus a weekly ceiling on the newest model. Which one binds depends on whether the work is bursty or continuous, and neither is published as a number.
    Maxxed ceiling
    modelled (no quota published)No published quota to saturate, so nothing better than the workload model exists (derivation 4). The figure here is the modelled sustained-agent ceiling: an UPPER BOUND on a modelled token mix, not a measurement of saturation. Read it as “we cannot see this plan’s ceiling”, not as “this plan has no headroom”.
    Plan terms
    High
    Throughput
    Low · high volatility
    List price
    $20/mo
    Token value
    $410–$410
    Sources
    [1][2]
    Last checked
    checked Sep 5

    Rolling 5h window + weekly cap, ~5× the free tier; absolute messages/window no longer published. Entry tier — the 1× reference for Anthropic. Modelled: Anthropic publishes no ceiling we can apply and we do not meter this plan, so the throughput is the generic workload band.

  2. 2

    ChatGPT Plus

    $20/mo · GPT-5.6 Sol

    ~17×

    ChatGPT Plus

    rank #2 · GPT-5.6 Sol

    ~17×value per $ (headline)
    Maxxed range
    17×
    Observed
    ~17× · $348/mo API-equivalentbeside the modelled headline, not in place of it
    Evidence
    Observed · n=1 · seen Sep 7
    Surface ranked
    Codex (coding agent)This row ranks OpenAI's CODEX surface only, and the five-hour message ranges it publishes there. OpenAI meters Chat, Codex and Work on SEPARATE allowances, which cannot be added, averaged or blended into one number — doing that would invent a quota OpenAI does not publish. Codex is the surface this board is about (the same scope test that removed SuperGrok Lite and Google AI Plus), so it is the one ranked, and the chat allowances below are recorded as OUT OF SCOPE rather than folded in. ATTESTED, NOT READ: the chat figures come from the external audit's reading of help.openai.com articles that return HTTP 403 to every automated read from this host, so we publish them as somebody else's reading, not ours. On chat, Plus is the entry tier; the tie the audit breaks is between the two Pro tiers.
    Binding limit
    weekly + 5h windowOpenAI publishes local-message estimates per FIVE-HOUR period, per model and per tier (Plus / Pro 5x / Pro 20x), and states that weekly limits may also apply — so the window structure IS published, contrary to what this row said before. They are RANGES, explicitly disclaimed as “not fixed message limits” with the live numbers held in the usage dashboard, so the shape is published even though the quota is not.
    Maxxed ceiling
    metered cycleDerivation 1, and the only LIVE one on this board: a cycle we metered ourselves against OpenAI’s own quota gauge. An 18-minute agent task on the Codex CLI drew 14,016,606 tokens (98.2% cache reads) and the ChatGPT analytics page moved from 100% to 91% of the weekly limit remaining, so that one task was 9% of a Plus week. Published in DOLLARS at OpenAI’s own published rates — $80/week, ~$348/mo — because the measured cycle carries its own token mix and must never be re-valued through the board’s modelled one. NO TOKEN CEILING IS PUBLISHED FOR THIS ROW ON PURPOSE: extrapolating the tokens gives ~155.7M/week on the raw count and ~3.17M/week on fresh tokens only, 49x apart, because OpenAI meters this allowance in credits and requests rather than tokens. The dollar value of the cycle does not depend on that choice, which is why it is the number we sign.
    Plan terms
    Med
    Throughput
    Med · high volatility
    List price
    $20/mo
    Token value
    $348–$348
    Last checked
    checked Sep 6

    Standard text unlimited; only reasoning/tools metered, caps published only as per-five-hour message ranges on developers.openai.com/codex/pricing, disclaimed as not fixed. Entry paid tier — the 1× reference for OpenAI. Terms graded medium: openai.com and help.openai.com both refuse an automated read (HTTP 403), so the $20 price is cross-checked against secondary trackers, not read off the vendor page. THE RANK NO LONGER COMES FROM THAT MODEL: since 2026-09-07 this row is ranked on a cycle we metered on our own Plus account against OpenAI’s own weekly gauge — see Observed above. The published ranges stay on the row as what the vendor says, not as what the row is scored on.

  3. 3

    Claude Max (5×)

    $100/mo · Claude Opus 5

    ~17×

    Claude Max (5×)

    rank #3 · Claude Opus 5

    ~17×value per $ (headline)
    Maxxed range
    14–20×Headline is the centre: both ends are quota-saturated and equally likely.
    Evidence
    Community-observed · n=1 · date not establishedthe third-party log behind this row is undated, so we cannot say how old it is — treat the band, not the date, as what it supports
    Time-sensitive
    TEMPORARY, THROUGH 2026-09-13: Anthropic users report a +50% weekly-limit boost running to that date, so any Claude measurement taken inside that window overstates what the plan sustains afterwards. This row is MODELLED and does not incorporate the boost, and the community figures it cites were logged before it — but treat any Claude number you read anywhere this week, including your own, as inflated until the cut lands.
    Binding limit
    weekly + 5h windowA rolling 5-hour window plus a weekly ceiling on the newest model. Which one binds depends on whether the work is bursty or continuous, and neither is published as a number.
    Maxxed ceiling
    modelled (no quota published)No published quota to saturate, so nothing better than the workload model exists (derivation 4). This row tops out at the modelled sustained-agent ceiling: an UPPER BOUND on a modelled token mix, not a measurement of saturation. Read it as “we cannot see this plan’s ceiling”, not as “this plan has no headroom”. This row keeps its BAND rather than collapsing to that ceiling, because the band is not a workload guess — it is how much this tier delivers relative to Claude Pro, which is observed with real spread and which saturating the plan does not resolve.
    Plan terms
    High
    Throughput
    Low · high volatility
    List price
    $100/mo
    Token value
    $1,435–$2,049
    Sources
    [1][2]
    Last checked
    checked Sep 5

    “5×” is the 5-hour-session figure, not monthly; a June 2026 class action alleges ~3.5× Pro in practice. Modelled at 3.5× Pro at the floor and at the advertised 5× at the ceiling. Absolute caps unpublished. Community-observed rather than vendor-derived: the 3.5× floor comes from third-party allegation and user logs, not from anything Anthropic publishes.

  4. 4

    ChatGPT Pro (5×)

    $100/mo · GPT-5.6 Sol

    ~16×

    ChatGPT Pro (5×)

    rank #4 · GPT-5.6 Sol

    ~16×value per $ (headline)
    Maxxed range
    16×
    Evidence
    Modelled · n=0
    Surface ranked
    Codex (coding agent)This row ranks OpenAI's CODEX surface only, and the five-hour message ranges it publishes there. OpenAI meters Chat, Codex and Work on SEPARATE allowances, which cannot be added, averaged or blended into one number — doing that would invent a quota OpenAI does not publish. Codex is the surface this board is about (the same scope test that removed SuperGrok Lite and Google AI Plus), so it is the one ranked, and the chat allowances below are recorded as OUT OF SCOPE rather than folded in. ATTESTED, NOT READ: the chat figures come from the external audit's reading of help.openai.com articles that return HTTP 403 to every automated read from this host, so we publish them as somebody else's reading, not ours. ON CHAT, THE TIERS ARE NOT TIED: the audit reports Pro $100 at 50 GPT-6 Pro messages per week against Pro $200's 200 per week — 4x the weekly allowance for 2x the price, so on THAT surface $200 is the better value per dollar, the reverse of a tie. It does not move this row, because this row is not ranking that surface.
    Binding limit
    weekly + 5h windowOpenAI publishes local-message estimates per FIVE-HOUR period, per model and per tier (Plus / Pro 5x / Pro 20x), and states that weekly limits may also apply — so the window structure IS published, contrary to what this row said before. They are RANGES, explicitly disclaimed as “not fixed message limits” with the live numbers held in the usage dashboard, so the shape is published even though the quota is not.
    Maxxed ceiling
    modelled (no quota published)OpenAI publishes per-five-hour message RANGES rather than a token quota, and disclaims them as not fixed, so there is still no cycle this board can saturate arithmetically (derivation 4). The figure here is the modelled sustained-agent ceiling: an UPPER BOUND on a modelled token mix, not a measurement of saturation. Read it as “we cannot see this plan’s ceiling”, not as “this plan has no headroom”.
    Plan terms
    Med
    Throughput
    Low · high volatility
    List price
    $100/mo
    Token value
    $1,640–$1,640
    Tied tier
    unmeasured; tied by assumptionit prints the same rate as its other OpenAI tiers because neither is measured and the model scales the dearer one proportional to price
    Last checked
    checked Sep 6

    OpenAI sells Pro at two usage tiers, $100 (5× Plus) and $200 (20× Plus), same models and features — the price buys headroom only. Modelled proportional to price (5× Plus), not at the advertised 5× session label; they happen to coincide here. Terms graded medium: vendor pages return 403, cross-checked against secondary trackers.

  5. 5

    ChatGPT Pro (20×)

    $200/mo · GPT-5.6 Sol

    ~16×

    ChatGPT Pro (20×)

    rank #5 · GPT-5.6 Sol

    ~16×value per $ (headline)
    Maxxed range
    16×
    Evidence
    Modelled · n=0
    Surface ranked
    Codex (coding agent)This row ranks OpenAI's CODEX surface only, and the five-hour message ranges it publishes there. OpenAI meters Chat, Codex and Work on SEPARATE allowances, which cannot be added, averaged or blended into one number — doing that would invent a quota OpenAI does not publish. Codex is the surface this board is about (the same scope test that removed SuperGrok Lite and Google AI Plus), so it is the one ranked, and the chat allowances below are recorded as OUT OF SCOPE rather than folded in. ATTESTED, NOT READ: the chat figures come from the external audit's reading of help.openai.com articles that return HTTP 403 to every automated read from this host, so we publish them as somebody else's reading, not ours. ON CHAT, THE TIERS ARE NOT TIED: the audit reports this tier at 200 GPT-6 Pro messages per week (with GPT-5.6 Sol Pro at 170/day and both Pro models capped together at 200/day) against Pro $100's 50 per week — 4x the weekly allowance for 2x the price. On the chat surface this tier is therefore BETTER value per dollar than Pro $100, not tied with it; on the Codex surface ranked here, neither of them is measured at all.
    Binding limit
    weekly + 5h windowOpenAI publishes local-message estimates per FIVE-HOUR period, per model and per tier (Plus / Pro 5x / Pro 20x), and states that weekly limits may also apply — so the window structure IS published, contrary to what this row said before. They are RANGES, explicitly disclaimed as “not fixed message limits” with the live numbers held in the usage dashboard, so the shape is published even though the quota is not.
    Maxxed ceiling
    modelled (no quota published)OpenAI publishes per-five-hour message RANGES rather than a token quota, and disclaims them as not fixed, so there is still no cycle this board can saturate arithmetically (derivation 4). The figure here is the modelled sustained-agent ceiling: an UPPER BOUND on a modelled token mix, not a measurement of saturation. Read it as “we cannot see this plan’s ceiling”, not as “this plan has no headroom”.
    Plan terms
    Med
    Throughput
    Low · high volatility
    List price
    $200/mo
    Token value
    $3,279–$3,279
    Tied tier
    unmeasured; tied by assumptionit prints the same rate as its other OpenAI tiers because neither is measured and the model scales the dearer one proportional to price
    Last checked
    checked Sep 6

    The advertised “20× Plus” is a per-FIVE-HOUR rate-limit figure, not a sustained monthly one: OpenAI’s Codex pricing page scales its five-hour message ranges 5× and 20× off Plus, then disclaims them as not fixed limits, and publishes no monthly quota at all. A disclaimed burst range is not a ceiling this board can rank on — the same test applied to Google’s unperiodised 20× — so this stays modelled proportional to price (10× Plus). Terms graded medium: vendor pages return 403, cross-checked against secondary trackers.

Data provenance

Intelligence scores by Artificial Analysis— an independent third party, read from their own index. Usage and pricing are real token volume and list prices, live from OpenRouter.

OpenRouter benchmarks ↗

From signal to worker

Hire the model that is holding up.

Run it as an always-on AI teammate on a server you own.