Methodology
How a subscription gets measured
Every choice below exists to make one claim defensible: that a plan’s quota, if you had bought the same usage from the API, would have cost this much. Nothing here claims to reconstruct how a provider meters internally.
The unit
Dollars, not tokens
Token counts are not comparable across models or providers. Output tokens cost roughly five times input, a cache read roughly a tenth, and every provider weighs them differently — so “82M tokens” means one thing on a small model and another on a large one. Making tokens comparable requires choosing a set of weights, and the official API rate card is a set of weights: public, authoritative, and already maintained by the provider.
The deeper difference: a token-denominated metric is a guess about how the provider meters. A dollar-denominated one is a statement about what the user avoided paying. The second is true regardless of what happens inside the provider, which is why it is the metric of record here.
The formula
From one interval to a monthly ceiling
Prices are versioned. Each interval is priced with the rate card in force on the day it was recorded, kept in the repository as a data file with an effective_from date — otherwise a price cut and a quota cut look identical in a historical chart.
Only the binding window may be extrapolated. The multiplier is 30.4375 Julian days in a month divided by the window length in days — 30.4375/7 ≈ 4.35 for weekly. Extrapolating from the 5-hour window produces a figure the weekly cap makes unreachable — that number is barred from every published result.
The published entity is a plan, not a model. The same plan measured through a small model and through a large one must converge on the same dollar figure; the gap between models is folded into the confidence grade.
The data model
Snapshot, observation, delta
- Quota Snapshot
- One moment’s reading of every window: percentage used and reset timestamp.
- Observation
- A pair of snapshots plus the local token detail recorded between them.
- WindowDelta
- What one observation implies for one window. Quality is scored here, not on the observation: a single interval can be worthless for the 5-hour window because it straddles a reset, and perfectly good for the weekly one.
A delta counts as high quality when
- The interval does not cross a window reset, judged by the official reset timestamp.
- The interval is short.
- The usage delta is large — with an integer percentage field, a bigger delta means less rounding error.
- The local token log is continuous across the interval.
- The model in use is unambiguous.
- Client state did not change mid-interval.
- Ideally: one device, one client.
Low-quality deltas are kept and down-weighted rather than deleted — they still carry information about drift, and discarding data silently is its own kind of dishonesty.
Known noise
What makes this hard
- A subscription’s quota is shared with the web, desktop, and mobile clients. Usage recorded nowhere locally still moves the percentage.
- Users have reported quota states that do not add up — a daily window at 0% while the weekly window throttles — and display values that disagree with the actual cut-off.
- The percentage field is quantized to whole numbers, so every small delta carries up to half a point of rounding error. This is why large deltas are preferred over frequent ones.
- Readings are cached per session: two live sessions on one account can report different percentages at the same instant. Deltas therefore have to come from a single session’s own series. Measured on day 0 →
- Capacity itself moves with context size, effort level, and cache hits. No provider promises a fixed number of prompts.
Provenance
Where each number comes from
| Grade | Meaning |
|---|---|
| Measured | Read from a probe account we own and operate. |
| Community measured | Derived from anonymous telemetry contributed by users. |
| Derived | Computed from a measured plan via an official multiplier, inside one window only. |
| Estimated | Everything else, and labelled as such. An estimate is never presented as an official quota. |
Confidence (High / Medium / Low) is graded on sample count, number of independent contributors, agreement between models, and whether any model in the plan lacks a public API price to convert against.
Publication gate
When numbers ship
The current experiment normalises the same set of high-quality deltas twice — once in dollars, once in tokens — and compares the variance. The decision table was written before the data arrived:
| Result | Consequence |
|---|---|
| Dollar-denominated variance clearly below token-denominated variance | Metering tracks cost. The method holds, and this page says so. |
| Dollar-denominated variance within ±10% | Good enough to build on: continue with the CLI and the aggregation pipeline. |
| Both denominations vary by ±50% | Something in the metering is unmodelled. Understand the unit before shipping any benchmark. |
Until one of the first two rows holds, the leaderboard stays empty. Current progress →