Check your sub’s worth.
SubWorth is an open benchmark measuring what AI coding subscriptions actually deliver — in dollars of equivalent API usage, derived from real usage rather than from marketing pages.
Last reading 2026-09-25 08:00 UTC. Follow the experiment →
The question
What is a Worth Multiple?
Providers publish multipliers — “5× Pro”, “20× Pro” — but never amounts. SubWorth supplies the amount from two things anyone can observe locally: what a stretch of usage costs at the official API rate card, and how much of the quota window it moved.
The binding window is whichever quota window runs out first — on Claude plans, the weekly one — and it is the only window that may be extrapolated. A month is 30.4375 Julian days (365.25/12), so a weekly window regenerates 30.4375/7 ≈ 4.35 times per month, not 4. Every published figure carries a fixed scope: theoretical monthly ceiling, weekly-capped.
Leaderboard
What will be published, and when
The table exists before the numbers do. Each row lists the plan, the signal it is measured from, and the confidence grade it can realistically reach.
| Plan | Role | Quota signal | Weekly value | Worth Multiple | Confidence |
|---|---|---|---|---|---|
| Claude Max 20x | Primary | statusline rate_limits, integer % | Measuring | Measuring | High (target) |
| Codex Pro 20x | Secondary | /wham/usage endpoint via hook, integer % | Measuring | Measuring | Low → Medium |
| Kimi Allegro | Secondary | official /usages endpoint via hook, integer % | Collecting | Collecting | Medium (ceiling) |
| Grok Heavy | Secondary | unified.jsonl billing events imported, integer % | Collecting | Collecting | Low |
Confidence shows the grade each plan can realistically reach, not one it has earned. Grok Heavy carries a caveat that will ship with any number it produces: the plan-exclusive Heavy model is web-only with no API price and invisible to the CLI, so what is measurable is this account’s CLI coding usage, excluding whatever the web-side Heavy model consumes.
Pipeline
How a number would be produced
1 · Collect
A statusline hook appends each quota reading — window, percentage used, reset time — to a local JSONL file. Nothing is uploaded; nothing but those fields is ever written.
2 · Price
The local token log for the same interval is priced item by item with the provider’s API rate card as it stood that day. Rate cards are versioned in the repository, so a price cut never reads as a quota cut.
3 · Aggregate
Each pair of readings becomes one delta per window, quality-scored on its own terms, reduced per contributor and then across contributors by median — never a flat average, which one heavy user could swing.
Commitments
What this project holds itself to
- Raw conversations never leave the device. Collection is local by default and uploading is an explicit, separate command. See exactly what would be sent →
- Every number carries its provenance. Measured, community measured, derived, or estimated — plus a confidence grade. An estimate is never dressed up as an official quota.
- Nothing ships before the criteria pass. The validation experiment has a written pass/fail table, decided before the data came in. The criteria →