Methodology

How the figures on the coding-subscription pages are produced, and what each one is allowed to claim.

Confidence grades

Every allowance carries a grade. It describes the strength of our evidence, not the size of the number.

official
Stated on the vendor's own pricing page or documentation. Requires a source URL and the vendor's wording.
measured
Measured by us or by the community. Requires a citable source and a description of the method.
estimated
Derived from other known figures. Requires the reasoning to be written down.

A derived figure takes the weakest grade among its inputs. This comparison page shows only official and measured figures; estimated ones appear in the planner, where you can see and change the assumptions.

What the vendor discloses

Separate from confidence, and often confused with it. Anthropic's "5x more usage than Pro" is official-grade evidence — their own words — of a plan whose allowance is still not a number.

published
An absolute figure, or an explicit "unlimited".
relative
Only a ratio to another plan, whose own allowance may be unpublished.
undisclosed
No figure at all.

An undisclosed allowance is not a gap in our research. Withholding the number is what lets a vendor tighten the allowance later without announcing a change or breaking a stated promise, so we mark it as a risk rather than an empty cell.

Capability tiers

Tiers label the models whose vendor published a comparable score. They appear on the model chips and nowhere else: a tier is a coarse bucket, never a score and never a ranking within itself, and nothing on this site filters or ranks by one.

A tier comes from a vendor-published Terminal-bench 2.1 score where one exists: frontier at 80 and above, strong from 70 to 79.9, efficient below 70. Where no score is published but the vendor states an unambiguous ordering about its own model, that statement sets the bucket and is recorded verbatim. Where neither exists, the model stays untiered.

Untiered does not mean weak. It means nobody published a figure we can compare. It used to cost the model its place under a capability floor — that floor is gone, because what it rejected was overwhelmingly vendors who publish no benchmark rather than models that fall short. A value figure names the model it was priced on instead.

Two things we deliberately do not use: scores from different benchmark versions, which move a model a whole bucket on nothing but the version; and distributors' billing tiers, which track price closely enough that ranking value-per-price on them would be circular.

Current coverage: 9 of 19 base models carry a tier.

ModelVendorTierEvidence
Claude Fable 5AnthropicfrontierNo Terminal-bench score published; Anthropic reports benchmarks as chart images. Tiered on Anthropic's own ordering: the Claude Opus 5 launch post (2026-07-24) describes Opus 5 as coming "close to the frontier intelligence of Claude Fable 5 at half the price", placing Fable 5 at the vendor's capability ceiling. https://www.anthropic.com/news/claude-opus-5
Claude Haiku 4.5Anthropicuntiered
Claude Opus 4.6Anthropicuntiered
Claude Opus 4.8Anthropicuntiered
Claude Opus 5AnthropicfrontierNo Terminal-bench score published. Tiered on Anthropic's own claim in the launch post (2026-07-24): "on Frontier-Bench v0.1, Opus 5 surpasses all other models, and more than doubles Opus 4.8's performance at a lower cost per task", and "It's the new default model on Claude Max, and the strongest model on Claude Pro." https://www.anthropic.com/news/claude-opus-5
Claude Sonnet 4.6Anthropicuntiered
Claude Sonnet 5AnthropicfrontierTerminal-bench 2.1 (Terminus-2 harness) 80.4%, as reported by Google in the Gemini Flash model table, retrieved 2026-07-27. At or above the 80.0 frontier threshold. https://deepmind.google/models/gemini/flash/
Composer 2.5Cursoruntiered
Gemini 3.1 ProGooglestrongTerminal-bench 2.1 (Terminus-2 harness) 73.8%, published by Google in its own Gemini Flash model table, retrieved 2026-07-27. The same model scores 68.5% on Terminal-Bench 2.0 in Google's Gemini 3.1 Pro table — the two versions are three points apart on one model, which is why they are not mixed. https://deepmind.google/models/gemini/flash/
Gemini 3.5 FlashGooglestrongTerminal-bench 2.1 (Terminus-2 harness) 76.2%, published by Google in its own Gemini Flash model table, retrieved 2026-07-27. https://deepmind.google/models/gemini/flash/
Gemini 3.6 FlashGooglestrongTerminal-bench 2.1 (Terminus-2 harness) 78.0%, published by Google in its own Gemini Flash model table, retrieved 2026-07-27. Within the 70.0-79.9 strong band. https://deepmind.google/models/gemini/flash/
Gemini 3 FlashGoogleuntiered
GPT-5.3 CodexOpenAIuntiered
GPT-5.5OpenAIfrontierTerminal-bench 2.1 83.4%, as reported by Anthropic in the Claude Opus 4.8 launch post's footnotes, retrieved 2026-07-27. Harness caveat recorded because it matters: the figure is OpenAI's self-reported score using the Codex CLI harness, not the Terminus-2 harness the other tiered figures use, so it is not strictly comparable with them. https://www.anthropic.com/news/claude-opus-4-8
GPT-5.6 LunaOpenAIfrontierTerminal-bench 2.1 (Terminus-2 harness) 84.7%, as reported by Google in the Gemini Flash model table, retrieved 2026-07-27 — the highest figure in that table. Note the tension worth keeping visible: OpenAI positions Luna as the lightweight, high-volume member of the GPT-5.6 family, while the benchmark places it above both flagships that have scores. https://deepmind.google/models/gemini/flash/
GPT-5.6 SolOpenAIuntiered
GPT-5.6 TerraOpenAIuntiered
Grok 4.5xAIfrontierTerminal-bench 2.1 (Terminus-2 harness) 83.3%, as reported by Google in the Gemini Flash model table, retrieved 2026-07-27. xAI's own docs publish no benchmark figures. https://deepmind.google/models/gemini/flash/
Grok Build 0.1xAIuntiered

Assumptions

Converting an allowance into dollars needs a model of how a coding session consumes tokens. These are editorial assumptions, not measurements, and they are adjustable in the planner.

ParameterValueBasis
Input tokens per turninputTokensPerTurn12,000Average context of one agent coding turn, including file reads
Cache hit ratiocacheHitRatio0.7Typical prompt-cache hit rate over a long session
Output tokens per turnoutputTokensPerTurn1,500Average generation per turn
Turns per hourturnsPerHour20Interaction rate during active coding
Hours per dayhoursPerDay6Hours a full-time developer actually spends coding
Days per monthdaysPerMonth22Working days

How a value figure is built

A plan's API-equivalent value comes from one allowance, priced on one model. Three steps decide which allowance and which model.

  1. 1Within one model, the rows that govern it are constraints. A plan-wide allowance and a row naming that model are not two offers to choose between — the second narrows the first — so the model's ceiling is the tighter of them.
  2. 2Across models, the same rows are alternatives. Models sharing a five-hour window spend it on whichever one you pick, so the group is worth the best choice in it. The model that wins is the one the figure is priced on.
  3. 3Across groups, every row binds. A five-hour window and a weekly cap are both hit, so the plan is worth the tightest of them.

Value multiple = that figure ÷ the monthly price. Above 1.00 means the allowance buys more, at published API rates, than the subscription costs.

The model the figure was priced on is named beside it, and has to be read with the multiple. Because the best choice wins across models, a plan can carry its whole multiple on a generous allowance of its cheapest model — the figure is correct and describes a model you may never use. Naming it is the only guard here. We do not filter models by capability: that would need benchmark coverage this catalogue does not have, and the section below says how little there is.

Four cases produce no figure rather than a soft one: a plan priced only on request; a published allowance we cannot convert to dollars anywhere in the chain; a free plan, which has no price to divide by; and allowances that reach none of the models we list.

Which model for which kind of work

The planner routes each category of work to a model. This is the weakest evidence on the site, and the coverage below is the first thing to read.

Coverage

CategoryBenchmarkModels with a score
Architecture & designnone published0
Hard debuggingDeepSWE v1.16
Backend implementationSWE-Bench Pro (Public)7
Frontend UInone published0
Refactoring & testsnone published0
DevOps & scriptingTerminal-bench 2.1 (Terminus-2)6

Where a benchmark exists, a model's band is set relative to the best published score in that same category: preferred at 92% of it or above, capable at 70% or above, unfit below. Relative rather than absolute because these benchmarks are not equally hard — the field tops out around 85% on one and 65% on another, and a fixed cut would grade the same model differently on nothing but the difficulty of the test.

Where no benchmark covers a category, a model inherits `capable` from its capability tier and carries no success rate. It can still be routed, but it can never win on a cost comparison it was never measured in, and the planner marks that row as chosen by tier.

Three kinds of evidence, in order of strength: benchmarks a vendor publishes in its own comparison table; aggregate behaviour from third-party leaderboards, which is not yet used because each one's redistribution licence has to be cleared first; and reader feedback on this site, which can raise a proposed change for review and never edits anything by itself.

Architecture and design has no mature public benchmark and is not expected to acquire one, so the planner shows the capability tier to look for and names no model for it.

From token price to cost per task

A stronger model costs more per token and needs fewer attempts, so 'dearer per token' and 'cheaper per task' are routinely true of the same model. The planner compares the expected cost of finishing one task: the category's baseline tokens, adjusted for the model, divided by its solve rate.

Baseline tokens per task

Architecture & design120,000
Hard debugging180,000
Backend implementation90,000
Frontend UI70,000
Refactoring & tests45,000
DevOps & scripting35,000

Every result is a range, produced by varying token use by ±30% and solve rate by ±10 percentage points. A single figure built on an editorial baseline times a benchmark solve rate would be the most confident-looking number on this site and the least defensible.

Feedback thresholds

A disagreement about which model suits a category becomes a proposed change once 5 readers have said the same thing and they are at least 70% of the votes on that pair. The proposal goes to a person, who checks it against a published benchmark; feedback never edits the data by itself.

Reports of hitting a plan's limit are grouped by plan, period and usage intensity, never pooled across intensities — heavy users cluster on the larger plans, and pooling would show the expensive plans running out sooner as an artefact of who bought them. A rate is shown only once a group holds at least 20 reports. Below that the cell says so.

How the figures are kept current

A watcher re-reads each vendor's pages every three days and compares what it finds against what we hold. Nothing it finds is published until a person confirms it against the source. When a vendor's page stops being readable, that shows on the plan rather than being absorbed quietly.

Last checked: 2026-07-29

All coding subscriptions