Cost per merged PR: what shipped work costs

Cost per merged PR is the total agent spend of one task whose pull request was merged; DevFlow will publish the median per implementer model once 30 such tasks exist.

Why measure per shipped change

Spend per call is easy to compute but says little about value. A cheap model that needs three rework rounds can cost more than an expensive one that gets the change right first time. Counting only work that shipped puts failed and abandoned attempts where they belong: in the bill of the change that finally landed.

How the poller sets the status

After mark-done, DevFlow opens a pull request on GitHub and polls its state every 60 seconds. Once none of a task's pull requests is still open, its status becomes merged if at least one of them merged, and pr_closed otherwise. A task that spans several repositories can have several pull requests, and it still counts once, with one total.

How the figure is built

For every merged task, the planned export sums cost_usd over all of its agent runs: planning, decomposition, implementation, review, summary and any rework. Every role counts, not only the code writer.

The total is credited to whichever model ran the implementer role most often on that task, with ties broken by spend and then by name; tasks with no implementer run are skipped, and the ranking ignores the <model>:reasoning bookkeeping rows.

A worked example

Suppose most implementer runs on one task ran on one model and a single rework subtask on another, while planning and review used a third. The whole total, rework included, goes to the first of the three.

What is not counted

Failed tasks, tasks still waiting for review and pr_closed outcomes are not counted, while failed attempts inside a task that merged do count. Spend with no task attached, such as team chat, is outside the metric too. Work that never reached mark-done has no pull request at all, so it never enters the figure.

Window and threshold

A merged task falls in the window when its last agent run is inside the trailing 90 days. All of its spend then counts, including runs from before the window opened. The export reads only the timestamps of agent runs, never the moment GitHub reports the pull request as closed.

The export will publish, for each implementer model, the median of these totals and the count of merged tasks behind it, and will withhold the figure below 30 merged tasks. A median from a handful would move with every new one.

Where the figure can appear

The figure is attached only to a model that the snapshot already lists. A model with fewer than 100 recorded runs in total is left out of the snapshot, so it shows no shipped-change median even if enough of its changes landed.

With an even count, the median is the mean of the two middle values; with an odd count, it is the middle value itself. Either way, one unusually expensive change cannot drag it far.

Reading it next to evidence

A low number means little if checks were skipped. The merge evidence report on each pull request shows which gates ran: the models used, the verifier outcome per subtask, the build gates, the security scan summary and the human approvals. Read together, the two figures answer what a change cost and how well it was verified.

A related tool, the per-task cost comparison, re-prices recorded tokens at the OpenRouter rates of another model, which answers what one finished change would have cost elsewhere.

FAQ

Are abandoned tasks included in cost per merged PR?

No. Only tasks in the merged state count; a task whose pull requests were all closed without merging ends in pr_closed and is left out.

Why does cost per merged PR credit only the implementer model?

One task can use several models across roles. Crediting the one that ran most implementer runs gives a single comparable figure per model.