Median and p90: reading agent run statistics
The median is the value half of all runs fall below and p90 the value 90 percent fall below; DevFlow plans to publish both for spend per run, over a trailing 90-day window.
Two summaries of one distribution
Sort a set of runs by some figure. The median sits in the middle: half of the runs are below it and half above. The 90th percentile, written p90, is the value that 90 percent of runs stay under.
These two statistics describe a distribution better than any single mean. The median says what to expect from a normal run; p90 says how bad an unlucky run gets.
Why agent runs need both
Agent workloads are skewed. Most runs finish quickly, but some explore a large codebase, loop over failing tests or produce a long file. Those few runs can be many times the typical size, and long test logs or large diffs push a run further toward the tail.
A mean blends the outliers in and describes no real run. Reporting the middle and the tail separately keeps both visible.
A worked example
Take ten runs of one role, sorted by spend. Nine are close together and the tenth is a long exploration. The median is set by the fifth and sixth runs and does not move however large the tenth one becomes.
With percentile_cont, the p90 of those ten runs falls between the ninth and the tenth, so the outlier pulls it up. The gap between the two figures is the signal that a long tail exists.
What DevFlow plans to publish
For each model and role, the planned export computes both statistics for spend per run, and only the median for uncached input, generated output and wall-clock duration. Values come from PostgreSQL's percentile_cont, which interpolates between the two nearest runs when the position falls between them.
Cost per merged pull request uses a plain median over merged tasks, computed in Go.
Minimum sample sizes
Percentiles from a few points are unstable, so thresholds apply. A model and role pair needs at least 30 runs in the trailing 90 days, otherwise the pair will be published as empty. A model with under 100 runs in total will not appear at all.
Keeping the empty entry lets a page say the data is thin instead of silently dropping the role. The snapshot also stores both thresholds, 30 and 100, next to the data, so a page can state the rule it was built under.
What percentiles do not tell you
Percentiles do not add up. Summing the figures of several roles does not give the figure of a whole task, and multiplying a typical run by a run count does not give total spend.
Percentiles also hide order and timing. Two models with equal figures can differ in how often their expensive runs happen in a row.
Reading a gap
A p90 close to the median means predictable runs. A wide gap means occasional expensive outliers, which matter more for a monthly budget than the typical run does.
FAQ
Why use the median instead of an average for agent runs?
One runaway agent run can pull an average far from what a typical run looks like. The median ignores how extreme the outliers are.
Why will DevFlow publish no p90 for token counts?
The planned snapshot keeps the p90 for spend per run, where the tail matters for budgeting. Token counts and duration are published without a p90.