Methodology

How a subscription gets measured

Every choice below exists to make one claim defensible: that a plan’s quota, if you had bought the same usage from the API, would have cost this much. Nothing here claims to reconstruct how a provider meters internally.

The unit

Dollars, not tokens

Token counts are not comparable across models or providers. Output tokens cost roughly five times input, a cache read roughly a tenth, and every provider weighs them differently — so “82M tokens” means one thing on a small model and another on a large one. Making tokens comparable requires choosing a set of weights, and the official API rate card is a set of weights: public, authoritative, and already maintained by the provider.

The deeper difference: a token-denominated metric is a guess about how the provider meters. A dollar-denominated one is a statement about what the user avoided paying. The second is true regardless of what happens inside the provider, which is why it is the metric of record here.

The claim is API-equivalent value — what this usage would have cost at list price. It is not a claim about the provider’s internal accounting, which may weigh things differently.

The formula

From one interval to a monthly ceiling

consumed ($)=ninPin+noutPout+ncrPcr+ncwPcwAEV(w)=consumed ($)% of window consumedMonthly ceiling=AEV(wbinding)×30.43757Worth Multiple=Monthly ceilingmonthly price\begin{aligned} \text{consumed}\ (\$) &= n_{\text{in}}P_{\text{in}} + n_{\text{out}}P_{\text{out}} + n_{\text{cr}}P_{\text{cr}} + n_{\text{cw}}P_{\text{cw}} \\[8pt] \mathrm{AEV}(w) &= \frac{\text{consumed}\ (\$)}{\text{\% of window consumed}} \\[8pt] \text{Monthly ceiling} &= \mathrm{AEV}(w_{\text{binding}}) \times \frac{30.4375}{7} \\[8pt] \text{Worth Multiple} &= \frac{\text{Monthly ceiling}}{\text{monthly price}} \end{aligned}

Prices are versioned. Each interval is priced with the rate card in force on the day it was recorded, kept in the repository as a data file with an effective_from date — otherwise a price cut and a quota cut look identical in a historical chart.

Only the binding window may be extrapolated. The multiplier is 30.4375 Julian days in a month divided by the window length in days — 30.4375/7 ≈ 4.35 for weekly. Extrapolating from the 5-hour window produces a figure the weekly cap makes unreachable — that number is barred from every published result.

The published entity is a plan, not a model. The same plan measured through a small model and through a large one must converge on the same dollar figure; the gap between models is folded into the confidence grade.

The data model

Snapshot, observation, delta

Quota Snapshot
One moment’s reading of every window: percentage used and reset timestamp.
Observation
A pair of snapshots plus the local token detail recorded between them.
WindowDelta
What one observation implies for one window. Quality is scored here, not on the observation: a single interval can be worthless for the 5-hour window because it straddles a reset, and perfectly good for the weekly one.

A delta counts as high quality when

Low-quality deltas are kept and down-weighted rather than deleted — they still carry information about drift, and discarding data silently is its own kind of dishonesty.

Known noise

What makes this hard

The output is therefore an estimate with a confidence interval, not a precise figure. A benchmark that reported four significant digits here would be lying about its own precision.

Provenance

Where each number comes from

GradeMeaning
MeasuredRead from a probe account we own and operate.
Community measuredDerived from anonymous telemetry contributed by users.
DerivedComputed from a measured plan via an official multiplier, inside one window only.
EstimatedEverything else, and labelled as such. An estimate is never presented as an official quota.

Confidence (High / Medium / Low) is graded on sample count, number of independent contributors, agreement between models, and whether any model in the plan lacks a public API price to convert against.

Publication gate

When numbers ship

The current experiment normalises the same set of high-quality deltas twice — once in dollars, once in tokens — and compares the variance. The decision table was written before the data arrived:

ResultConsequence
Dollar-denominated variance clearly below token-denominated varianceMetering tracks cost. The method holds, and this page says so.
Dollar-denominated variance within ±10%Good enough to build on: continue with the CLI and the aggregation pipeline.
Both denominations vary by ±50%Something in the metering is unmodelled. Understand the unit before shipping any benchmark.

Until one of the first two rows holds, the leaderboard stays empty. Current progress →