Agent invocations: one run of one pipeline role

An agent invocation is one run of one pipeline role, such as the planner or the verifier; DevFlow gives each invocation its own process and a hard cap of 90 minutes.

The unit of work

A coding pipeline is a chain of model calls, each with its own instructions and tools. An agent invocation is one link in that chain: a single role, given a prompt, working through as many turns as it needs until it returns a result. Planning, decomposing, implementing and reviewing are separate invocations.

How DevFlow runs one

Every invocation is a fresh subprocess of the agent CLI, started in the task's worktree. Its output is streamed to the browser line by line and parsed when the subprocess exits, and the next pipeline step builds on that result.

Two clocks guard the subprocess. A hard cap stops it after 90 minutes, and an idle watchdog kills it after 45 minutes without output. The idle limit is the real liveness check; the hard cap is a backstop.

Retries stay inside one run

When a provider reports a transient error, DevFlow re-issues the same call inside the invocation instead of starting a new one. The backoff honours a Retry-After header and shows a retrying notice in the browser while it waits. Agent errors that are not transient are not retried: the task fails and its worktree is kept for inspection.

The operator sets the attempt limit with AGENT_RETRY_MAX, and a value of 1 turns these retries off; the base and maximum delays between attempts have settings of their own.

Two other narrow exceptions do start a second invocation. An implementer that returns without any file change is re-run once, and an OpenCode run whose final answer fails JSON validation gets one reformatting invocation with every tool denied, which repairs the format of the finished work without redoing it.

What gets recorded

When an invocation completes, DevFlow writes rows to the agent_invocations table, one per model the run used. Each row stores the role, the backend, the model id, the reported spend in US dollars, input and output token counts, cache reads and cache writes, and the duration in milliseconds.

Reasoning tokens are kept under a separate <model>:reasoning key, so they never inflate the input count. Analytics and the planned public snapshot skip those bookkeeping rows.

A typical task

One task produces at least one invocation for planning, since planning can loop through rounds of questions, one for decomposition, implementer invocations for its subtasks, one review pass at the end and a summary. Each of those becomes at least one row, so the spend of the whole task is the sum of its rows.

Why it is the base metric

Every spend figure in DevFlow starts from these rows. The monthly budget sums them, the model performance view groups them by role and model, and the planned public snapshot will compute its medians from them.

That snapshot will cover a trailing 90-day window and show a role's figures for a model only once they rest on at least 30 invocations. A model with fewer than 100 invocations in total will be left out entirely.

FAQ

Is a retry after a rate limit a new agent invocation in DevFlow?

No. When a provider answers 429, 503 or 529, DevFlow re-issues the same call with bounded backoff, at most 4 attempts by default, inside the same invocation.

Do team chat replies count toward DevFlow's spend records?

Yes. Team chat runs have no task, so their spend rows carry the workspace directly, which keeps team chat inside the monthly spend total.