Prompt caching: reusing prompts in AI agents
Prompt caching lets a model provider reuse an already processed prompt prefix instead of reading it again; DevFlow tracks it with 2 counters per run, cache reads and cache writes.
How it works in general
Many requests to a model begin with the same long text: system instructions, tool definitions, earlier turns of a conversation. A provider that supports caching stores the processed form of that opening part. When the next request starts with the identical prefix, the provider reads the stored version instead of processing that text again.
Reads from the cache are typically billed below normal input, and writing to the cache may carry a surcharge. The saving depends on how much of each request repeats exactly.
Why agents benefit
An agent turn resends everything before it, so later turns share a growing prefix with earlier ones. Long runs with many tool calls repeat the most. A change near the start of the prompt breaks the match for everything after it.
What DevFlow records
Each run's row in agent_invocations has a cache_read_tokens and a cache_creation_tokens column, next to uncached input and generated output. The OpenCode backend fills both columns from the step reports of the CLI, and the Claude Code path fills them from the CLI's per-model usage block. Either way the counts reflect what the provider actually returned, and the Claude Code backend is off unless the operator enables it, so on a default installation every figure comes from OpenCode.
Keeping the prefix stable
Stable instructions first and changing details last is the general rule, and a timestamp or a reordered list near the top is the typical way to break it. System instructions and tool definitions open every request, so they are the part most worth keeping identical from one turn to the next. Appending new material at the end, as an agent does after every tool call, keeps everything before it reusable.
The compression proxy
A workspace can route Anthropic and OpenAI traffic through an optional context-compression proxy, set in Workspace Settings → Execution and off by default. The setting applies only to the OpenCode backend; the Claude Code path is unaffected.
The proxy only covers those two providers. OpenRouter has no route through it, so for the default model, openrouter/minimax/minimax-m3, switching the proxy on changes nothing and the provider's prefix reuse is left as it was.
Measuring it over time
The planned public model pages will summarise caching as one figure per model and role, the cache read share, computed over a trailing 90-day window. The planned export computes it as cache reads divided by the sum of uncached input and cache reads, so cache writes stay out of the ratio.
As an example, a share of 0.75 means that of all uncached input and reads together, three tokens in four were served from a stored prefix instead of being processed afresh. Compare the figure between models for the same role, since a planner and a summarizer send very different prompts.
FAQ
Do I have to switch on prompt caching in DevFlow?
No. Prompt caching belongs to the model provider and the agent CLI, not to a DevFlow setting; DevFlow records what the provider reported for each run.
Can a context-compression proxy hurt prompt caching?
It can. Lossy compression rewrites the start of the request, which may turn a cache hit into a miss, so net spend per task is the number to compare.