Cache read share: how much input was cached

Cache read share is the part of a model's input served from the prompt cache; DevFlow's planned export pools it per model and role over a trailing 90-day window, from 0 to 1.

The definition

Every request to a model has an input side. Part of that input may be served from the provider's prompt cache, and the rest is processed fresh. The share is the served part divided by the whole input side.

A value near 0 means almost nothing was reused. A value near 1 means most of what the model took in was already stored.

The exact formula

DevFlow's planned export adds up two columns over all runs of one model in one role: cache_read_tokens and the uncached input_tokens. The first sum divided by the total of both gives the share, and when both sums are zero the export reports 0.

Because the sums are pooled, a long run with many turns counts more than a short one. That weighting matches how spend accumulates.

The cache_creation_tokens column sits outside the formula: it counts writes, which fill the prompt cache rather than draw on it.

A worked example

Suppose one role made two runs with the same model. The first run was long and took most of its tokens from the prompt cache; the second was short and processed everything fresh. Pooling puts the result close to the first run's own ratio, because the first run moved far more tokens.

A plain average of the two per-run ratios would give the short run equal weight. That is why the export pools before dividing.

Outside the formula

Reasoning tokens live on separate <model>:reasoning rows, which the export skips. Output tokens play no part either, since the share describes only what the model took in.

The figure also says nothing about price. Providers bill cached and fresh tokens at their own rates, so two models with the same share can still differ in spend.

Window and thresholds

The export will look at a trailing window of 90 days. A model and role pair with fewer than 30 runs will be published as empty, and a model with fewer than 100 runs overall will be omitted.

Before writing, the exporter validates every value and stops on anything outside the range 0 to 1. A broken export fails at that check instead of reaching the site.

Interpreting it

A higher share usually means cheaper input for the same work, since cached tokens tend to be billed below fresh ones. Compare the share alongside median spend, though: a role can reuse a lot and still consume a great deal.

Differences between roles are expected. Long exploring roles resend more history than a quick one-shot role, so long roles have more prefix to reuse.

FAQ

Is cache read share an average of per-run ratios?

No. DevFlow's planned export pools all runs in the window before dividing, so long runs weigh more than short ones.

What range can cache read share take?

Between 0 and 1. DevFlow's exporter rejects any value outside that range before a snapshot is written.