方法论
开发套餐相关页面上的数字是怎么来的,以及每个数字允许声称到什么程度。
可信度分级
每条额度都带一个等级。它说的是我们证据的强度,不是数字的大小。
- official
- 厂商定价页或文档写明。必须有来源链接和厂商原话。
- measured
- 我们或社区实测。必须有可查出处,并写清测法。
- estimated
- 由其他已知量推算。必须把推算依据写下来。
派生指标取输入里最弱的一档。本对比页只出现 official 和 measured;estimated 只出现在规划器里,那里你能看见并改动假设。
厂商公开到什么程度
这跟可信度是两码事,但经常被混为一谈。Anthropic 那句“比 Pro 多 5 倍用量”是 official 级证据——厂商自己的话——而它描述的额度依然不是一个数字。
- published
- 给了绝对数字,或明确写了“不限量”。
- relative
- 只给了相对另一档的倍数,而那一档自己的额度可能也没公开。
- undisclosed
- 什么数字都没有。
能力档
档位只是给厂商公布过可比分数的模型贴的标签,出现在模型 chip 上,别处不用。档是粗筛,不是打分,档内也不排序,站内没有任何一处按档筛选或排序。
有厂商公布的 Terminal-bench 2.1 分数就按分数划:80 及以上为 frontier,70 到 79.9 为 strong,低于 70 为 efficient。没有分数但厂商对自家模型给出了明确排序的,按那句话定档并原样记录。两者都没有的,留空不划。
未划档不等于弱,只是没人公布过可比的数字。它曾经的代价是过不了能力下限——那个下限已经撤掉了,因为被它挡住的绝大多数不是能力不够的模型,而是不公布基准分的厂商。现在价值数字改为标明它是按哪个模型算出来的。
两样东西我们刻意不用:跨版本的基准分数,它能仅凭版本差异就把一个模型挪一整档;以及分销商的计费分档,它跟价格贴得太紧,拿它划档再按“价值除以价格”排名就成了循环论证。
当前覆盖:19 个 base model 中有 9 个已划档。
| Model | Vendor | Tier | Evidence |
|---|---|---|---|
| Claude Fable 5 | Anthropic | frontier | No Terminal-bench score published; Anthropic reports benchmarks as chart images. Tiered on Anthropic's own ordering: the Claude Opus 5 launch post (2026-07-24) describes Opus 5 as coming "close to the frontier intelligence of Claude Fable 5 at half the price", placing Fable 5 at the vendor's capability ceiling. https://www.anthropic.com/news/claude-opus-5 |
| Claude Haiku 4.5 | Anthropic | 未划档 | — |
| Claude Opus 4.6 | Anthropic | 未划档 | — |
| Claude Opus 4.8 | Anthropic | 未划档 | — |
| Claude Opus 5 | Anthropic | frontier | No Terminal-bench score published. Tiered on Anthropic's own claim in the launch post (2026-07-24): "on Frontier-Bench v0.1, Opus 5 surpasses all other models, and more than doubles Opus 4.8's performance at a lower cost per task", and "It's the new default model on Claude Max, and the strongest model on Claude Pro." https://www.anthropic.com/news/claude-opus-5 |
| Claude Sonnet 4.6 | Anthropic | 未划档 | — |
| Claude Sonnet 5 | Anthropic | frontier | Terminal-bench 2.1 (Terminus-2 harness) 80.4%, as reported by Google in the Gemini Flash model table, retrieved 2026-07-27. At or above the 80.0 frontier threshold. https://deepmind.google/models/gemini/flash/ |
| Composer 2.5 | Cursor | 未划档 | — |
| Gemini 3.1 Pro | strong | Terminal-bench 2.1 (Terminus-2 harness) 73.8%, published by Google in its own Gemini Flash model table, retrieved 2026-07-27. The same model scores 68.5% on Terminal-Bench 2.0 in Google's Gemini 3.1 Pro table — the two versions are three points apart on one model, which is why they are not mixed. https://deepmind.google/models/gemini/flash/ | |
| Gemini 3.5 Flash | strong | Terminal-bench 2.1 (Terminus-2 harness) 76.2%, published by Google in its own Gemini Flash model table, retrieved 2026-07-27. https://deepmind.google/models/gemini/flash/ | |
| Gemini 3.6 Flash | strong | Terminal-bench 2.1 (Terminus-2 harness) 78.0%, published by Google in its own Gemini Flash model table, retrieved 2026-07-27. Within the 70.0-79.9 strong band. https://deepmind.google/models/gemini/flash/ | |
| Gemini 3 Flash | 未划档 | — | |
| GPT-5.3 Codex | OpenAI | 未划档 | — |
| GPT-5.5 | OpenAI | frontier | Terminal-bench 2.1 83.4%, as reported by Anthropic in the Claude Opus 4.8 launch post's footnotes, retrieved 2026-07-27. Harness caveat recorded because it matters: the figure is OpenAI's self-reported score using the Codex CLI harness, not the Terminus-2 harness the other tiered figures use, so it is not strictly comparable with them. https://www.anthropic.com/news/claude-opus-4-8 |
| GPT-5.6 Luna | OpenAI | frontier | Terminal-bench 2.1 (Terminus-2 harness) 84.7%, as reported by Google in the Gemini Flash model table, retrieved 2026-07-27 — the highest figure in that table. Note the tension worth keeping visible: OpenAI positions Luna as the lightweight, high-volume member of the GPT-5.6 family, while the benchmark places it above both flagships that have scores. https://deepmind.google/models/gemini/flash/ |
| GPT-5.6 Sol | OpenAI | 未划档 | — |
| GPT-5.6 Terra | OpenAI | 未划档 | — |
| Grok 4.5 | xAI | frontier | Terminal-bench 2.1 (Terminus-2 harness) 83.3%, as reported by Google in the Gemini Flash model table, retrieved 2026-07-27. xAI's own docs publish no benchmark figures. https://deepmind.google/models/gemini/flash/ |
| Grok Build 0.1 | xAI | 未划档 | — |
假设参数
把额度折算成金额,需要一个“编码会话怎么消耗 token”的模型。这些是编辑部假设而非实测,在规划器里可以调。
| 参数 | 取值 | 依据 |
|---|---|---|
| 单轮输入 tokeninputTokensPerTurn | 12,000 | Agent 编码单轮平均上下文,含文件读取 |
| 缓存命中率cacheHitRatio | 0.7 | 长会话下典型的 prompt cache 命中率 |
| 单轮输出 tokenoutputTokensPerTurn | 1,500 | 单轮平均生成量 |
| 每小时轮次turnsPerHour | 20 | 活跃编码时的交互频率 |
| 每日小时数hoursPerDay | 6 | 全职开发者日均实际编码时长 |
| 每月天数daysPerMonth | 22 | 工作日 |
价值数字是怎么算出来的
一个套餐的 API 等效价值,来自一条额度、按一个模型折算。三步决定是哪一条、哪一个。
- 1对同一个模型,各行是约束。一条整体额度和一条点名该模型的行不是二选一,后者是把前者收窄,所以这个模型的上限取两者中更紧的那个。
- 2跨模型,各行是备选。共用一个五小时窗口的几个模型,窗口花在你挑中的那个上,所以这一组值最好的那个选择。胜出的模型就是这个数字的计价模型。
- 3跨组,各行同时生效。五小时窗口和周上限你都会撞上,所以套餐值其中最紧的一条。
价值倍数 = 这个数字 ÷ 月费。大于 1.00 表示按公开 API 单价算,这份额度值的钱超过订阅费本身。
四种情况我们不给数字,而不是给一个软化过的数字:需联系销售的套餐;链路上任何一处有已公布却折算不出美元的额度;免费套餐,没有可作除数的价格;公布的额度没有对应到我们列出的任何模型。
哪个模型干哪类活
成本测算页把每一类活路由到一个模型。这是全站可信度最弱的一层,下面的覆盖率应当先读。
覆盖情况
| 品类 | 基准 | 有分数的模型 |
|---|---|---|
| 架构设计 | 无公开基准 | 0 |
| 疑难调试 | DeepSWE v1.1 | 6 |
| 后端实现 | SWE-Bench Pro (Public) | 7 |
| 前端 UI | 无公开基准 | 0 |
| 重构与测试 | 无公开基准 | 0 |
| DevOps 与脚本 | Terminal-bench 2.1 (Terminus-2) | 6 |
有基准的品类,档位按同一品类内已公布最高分的相对位置划定:达到 92% 以上判 preferred,70% 以上判 capable,低于则判 unfit。用相对值而非绝对值,是因为这些基准难度并不相同——有的满场最高约 85%,有的只有 65%,固定阈值会仅因试题难度就把同一个模型划到不同档。
没有基准覆盖的品类,模型从能力档继承 capable,且不带解决率。它仍可被路由,但永远不会在一场它从未被测量过的成本比较中胜出,页面会把那一行标为「按档位」选出。
三类证据,按强度排列:厂商在自家对比表中公布的基准;第三方榜单反映的聚合行为,因需逐个核对再分发许可,目前尚未启用;以及本站读者反馈,它只能生成一条待审改动,本身不会改动任何数据。
架构设计没有成熟的公开基准,也不预期会有,因此这一品类只给出应当寻找的能力档,不点名模型。
从 token 单价到每任务成本
更强的模型每 token 更贵、需要的尝试更少,所以「每 token 更贵」和「每任务更便宜」在同一个模型上经常同时成立。测算比较的是完成一个任务的期望成本:品类基线 token 数按模型调整后,除以它的解决率。
各品类单任务基线 token
| 架构设计 | 120,000 |
| 疑难调试 | 180,000 |
| 后端实现 | 90,000 |
| 前端 UI | 70,000 |
| 重构与测试 | 45,000 |
| DevOps 与脚本 | 35,000 |
所有结果都是区间,由 token 消耗 ±30%、解决率 ±10 个百分点传播得到。用一个编辑部基线乘以一个基准解决率再算出的单点数字,会是全站看起来最确定、实际最站不住的数。
反馈阈值
对某个品类该用哪个模型的异议,需累计 5 条同向、且占该(模型,品类)全部投票的 70% 以上,才生成一条待审改动。改动交人工,对照公开基准核定;反馈本身永远不改库。
撞限报告按套餐、周期、使用强度分组,绝不跨强度合并——重度用户天然集中在更大的套餐上,合并会得出「越贵的套餐越容易撞限」这种由购买人群造成的假象。单个格子累计满 20 条才显示比率,不足则如实标注样本不足。
数字怎么保持新鲜
监测每 3 天重读一次各家页面,把读到的和库里的比对。它发现的任何改动,都要有人核对原文确认之后才会上线。厂商页面读不动了,这件事会显示在套餐上,而不是被悄悄吸收掉。
最近核实: 2026-07-29
全部开发套餐