# The Useful Metric Is Cost Per Finished Task

Source: https://www.bobzhu.tech/cost-per-finished-task/
Markdown: https://www.bobzhu.tech/assets/agents/cost-per-finished-task.md
Published: 2026-08-12T01:05:00.000Z
Tags: Essays, AI, Agents, Systems

Summary: The July 2026 model releases are argued in cost, tokens, and tool calls per task rather than leaderboard position. The harder problem is that most workflows can't say what finished means.

The number I care about now is what it costs to finish one real task.

Not the score. Not the price per million tokens. The whole cost of moving a specific piece of work from asked to done: input tokens, output tokens, tool calls, retries, wall-clock time, and the minutes I spend repairing whatever comes back.

The labs seem to have landed in the same place. Read the July releases from Anthropic, Google, and OpenAI and they are all arguing efficiency per task rather than position on a chart. The charts have a cost axis now.

That's a real change in what's being competed over. It also creates a problem that none of those posts can solve for you, which is the word *finished*.

## The release notes are already written this way

[Claude Opus 5](https://www.anthropic.com/news/claude-opus-5) arrived on 24 July with a pitch that is almost entirely economic. Anthropic describes it as coming close to the frontier intelligence of Claude Fable 5 at half the price, and keeps Opus pricing unchanged at $5 per million input tokens and $25 per million output tokens. The specific claims are per-task claims: on CursorBench at max effort, within 0.5% of Fable 5's peak score at half the cost per task. On Zapier's AutomationBench, a pass rate around 1.5 times the next-best model for the same cost per task. On OSWorld 2.0, better than every other model at any given cost.

[Gemini 3.6 Flash](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/) landed three days earlier with the same shape of argument from the cheap end. Google reports 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index, fewer reasoning steps and tool calls on multi-step workflows, and a lower price per token at $1.50 input and $7.50 output. Then it says the point out loud: the combination "reduces the overall cost per agentic task".

[GPT-5.6](https://openai.com/index/gpt-5-6/) had already framed its July launch as performance per dollar rather than raw capability. OpenAI's headline claim is state-of-the-art results with fewer tokens and at lower estimated cost, and the supporting numbers are ratios: Sol at medium reasoning beating Fable 5 on Agents' Last Exam by 11.4 points at roughly a quarter of the estimated cost, and Terra and Luna beating it at around one-sixteenth. Three weeks later OpenAI [cut Luna's price by 80% and Terra's by 20%](https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/), taking Luna to $0.20 input and $1.20 output per million tokens.
<div class="reader-artifact reader-artifact-notion" data-reader-artifact="notion">
  <div class="reader-notion-header">
    <span class="reader-notion-view">Table view</span>
    <h3 class="reader-notion-title">Three July releases, one argument</h3>
    <p class="reader-notion-summary">Each post leads with efficiency per task rather than a single leaderboard position.</p>
  </div>
  <div class="reader-notion-table">
    <div class="reader-notion-row is-header">
      <span class="reader-notion-cell">Release</span>
      <span class="reader-notion-cell">Unit</span>
      <span class="reader-notion-cell">Headline claim</span>
    </div>
    <div class="reader-notion-row">
      <span class="reader-notion-cell" data-reader-label="Release"><strong>Claude Opus 5</strong> (24 Jul)</span>
      <span class="reader-notion-cell" data-reader-label="Unit"><span class="reader-notion-tag">cost per task</span></span>
      <span class="reader-notion-cell" data-reader-label="Headline claim">Within 0.5% of Fable 5's peak CursorBench score at half the cost per task.</span>
    </div>
    <div class="reader-notion-row">
      <span class="reader-notion-cell" data-reader-label="Release"><strong>Gemini 3.6 Flash</strong> (21 Jul)</span>
      <span class="reader-notion-cell" data-reader-label="Unit"><span class="reader-notion-tag">tokens per task</span></span>
      <span class="reader-notion-cell" data-reader-label="Headline claim">17% fewer output tokens than 3.5 Flash, with fewer reasoning steps and tool calls.</span>
    </div>
    <div class="reader-notion-row">
      <span class="reader-notion-cell" data-reader-label="Release"><strong>GPT-5.6</strong> (9 Jul)</span>
      <span class="reader-notion-cell" data-reader-label="Unit"><span class="reader-notion-tag">performance per dollar</span></span>
      <span class="reader-notion-cell" data-reader-label="Headline claim">Sol at medium reasoning beats Fable 5 on Agents' Last Exam by 11.4 points at roughly a quarter of the estimated cost.</span>
    </div>
  </div>
</div>
Three labs, three eval suites, one denominator.

## Effort became a dial you pay for

The other thing all three releases share is that the model is no longer a single price point.

Opus 5 ships with an effort setting that Anthropic describes as a way to "optimize for intelligence or conserve tokens", and its charts plot performance against that setting rather than reporting one number. GPT-5.6 has `max` above `xhigh`, plus an `ultra` mode that coordinates four agents in parallel by default and openly trades higher token use for a better result sooner. Gemini's Flash-Lite exposes thinking levels for the same reason.

So the first decision isn't which model. It's how hard to make the model you already chose think.

That decision has a bigger spread than most model comparisons. One legal-workflow customer in Anthropic's early-access notes reported holding quality at lower reasoning levels while generating 26% fewer tokens on average than Opus 4.8 at max reasoning. Same family of models, roughly a quarter of the tokens gone, and the thing that changed was the dial.

Latency has a sticker price now too. Opus 5 has a Fast mode that runs around 2.5 times the default speed at twice the base price, and OpenAI's Fast mode does roughly the same for Sol. I find that clarifying rather than annoying. When speed is a line item, you have to answer whether finishing sooner is worth double, which is a better question than assuming faster is free.

## Fewer tokens is the quieter improvement

Most of the efficiency claims in these posts aren't about the answer. They're about the loop that produced it.

Google's line about fewer reasoning steps and tool calls is the one I'd underline. OpenAI reports that Sol surpasses Opus 4.8 on OSWorld 2.0 while using 85% fewer output tokens, and that Programmatic Tool Calling lets tool-heavy work advance with fewer tokens, fewer model round trips, and less guidance. Lovable, quoted in the same post, reported roughly 25% fewer steps and 35 to 48% fewer tool calls than the model it replaced, with stuck runs down 15%. On Anthropic's side, a financial-modelling customer reported nine percentage points more accuracy with a third fewer turns and tool calls and 60% less time, and a trading firm reported roughly a seventh of the reasoning tokens and under half the latency of Opus 4.8.

Tool calls deserve more attention than tokens here. A token is a cost. A tool call is a cost, a round trip, a chance to drift, and something a person may eventually have to read. When an agent takes forty steps to do a nine-step job, the bill is the least of it.

This is also why "cheaper model" and "cheaper task" keep coming apart. A model at a fifth of the token price that needs three times the steps and two attempts is not a saving. Per-token pricing tells you almost nothing on its own once the work involves tools.

## The hard part is the word finished

Every benchmark in those posts has something your workflow probably doesn't: a grader.

AutomationBench knows whether the business task completed start to finish. Terminal-Bench knows whether the command-line workflow ran. That's what makes a cost-per-task number meaningful — there's a definition of done sitting underneath it, and the failed attempts are counted as failures rather than quietly absorbed.

In most real workflows, done is a person's judgement made after the fact, which means the cost of the attempts that didn't work never gets attributed to anything.

The arithmetic is unforgiving once you write it down. Say a cheap lane runs ten tasks and finishes six cleanly. The other four come back plausible and wrong, and each takes twenty minutes to diagnose and repair. You've bought eighty minutes of human attention with your token saving, and those are the expensive minutes: interrupted, context-switched, and spent on someone else's mistake. This is the [time-back problem](https://www.bobzhu.tech/ai-does-not-automatically-give-you-time-back/) arriving through the billing page instead of the calendar.

The reverse holds too, which is the part people running cost-cutting exercises tend to skip. A model that costs four times as much per token and finishes first time can be the cheap option, and it usually is for anything where the repair work lands on a senior person.

So the useful denominator isn't tasks attempted. It's tasks accepted. And accepted has to mean something a machine can check: the tests pass, the schema validates, the numbers reconcile, the required sections exist, the [review layer](https://www.bobzhu.tech/ai-evaluation-checklist/) doesn't send it back. If the only definition of finished is that someone eventually stopped complaining, you don't have a metric. You have a spend figure.

## Their cost per task is not your cost per task

The vendor numbers are useful for direction and almost useless as a budget.

OpenAI is straightforward about this in its own footnotes: latency and API cost are estimated by simulating production behaviour offline, and real-world results may vary substantially. [^ OpenAI's stated method is to estimate latency and cost from the production behaviour of its models and simulate offline, accounting for tool call details, sampled tokens, and input tokens, with latency simulated at fast API speeds and cost at regular API pricing.] Anthropic's Frontier-Bench footnote is similarly specific about the setup: an internal run on the mini-SWE-agent harness with a GKE backend, mean reward over five attempts per task, with Opus 4.8 serving as the fallback when safety classifiers refused. [^ Five attempts per task is a sensible way to reduce variance in a benchmark. It's not how anyone runs production work, where the first attempt is the one you pay for and the second one is a decision.]

None of that is dishonest. It's the normal cost of publishing comparable numbers. But it does mean the harness is doing a lot of the work, and your harness isn't theirs.

Your cost per task depends on your prompt length, your cache hit rate, how many tools you've exposed, how patient your retry policy is, whether context is trimmed between steps, and what your reviewers will accept. Two teams using the identical model on the identical task can land a long way apart on all of those. The cross-vendor comparisons are shakier still, since each lab reports its own eval suite — CursorBench and AutomationBench in one post, Agents' Last Exam and Terminal-Bench in another, DeepSWE and OSWorld-Verified in a third.

Read their numbers as a claim about the direction of the frontier. Measure your own for anything you're going to budget.

## A task has to be a thing before you can price it

There's a plumbing version of this problem, and it's being solved in a way that's easy to miss.

The [2026-07-28 MCP specification](https://modelcontextprotocol.io/specification/2026-07-28) includes a [Tasks extension](https://modelcontextprotocol.io/extensions/tasks/overview) for long-running work. A server can return a durable handle instead of blocking, the client polls it, and the task carries a status through `working`, `input_required`, `completed`, `failed`, or `cancelled`. It survives a disconnect. It can pause to ask a person something and then continue.

That's described as a transport fix, and it is one. It's also an accounting shape.

A conversation that ran for forty minutes can't really be costed, because there's no boundary around it and no recorded verdict at the end. A task with an ID, a lifecycle, and a terminal state can be. You can attach tokens, tool calls, elapsed time, human interruptions, and the final status to the same object. That's the difference between knowing what you spent last month and knowing what a finished task costs.

## What I'd measure

This doesn't need a platform. It needs a spreadsheet and some discipline.

1. Pick five to ten tasks you genuinely repeat. Not demos. The weekly ones.
2. Write down what finished means for each, as a check something else can run.
3. Log every run: input tokens, output tokens including reasoning tokens, tool calls, wall-clock time, retries, human minutes, and whether it was accepted.
4. Divide total spend by accepted results, not attempts.
5. Move the effort dial on the model you already use before you change models. It's the cheapest experiment available and it usually has the widest spread.
6. Re-run the comparison when prices move. They moved twice in July.

The logging is the part people skip, and it's the part that makes the rest work. Most of the durable artefacts you'd [build around a workflow](https://www.bobzhu.tech/build-something-that-stays/) — the checks, the templates, the cached context, the deterministic preprocessing — show up as a lower cost per accepted task, which is a much easier case to make than an argument about good engineering practice.

It also turns [model routing](https://www.bobzhu.tech/gpt-56-routing-layer-cerebras/) into something you can settle with evidence. Which lane should handle this job is a guess until you have your own numbers, and then it's arithmetic.

Right now most teams can tell you what they spent on AI last month and roughly which models they like. Very few can tell you what one finished task costs. Closing that gap doesn't need a better model. It needs a definition of done and the habit of writing the numbers down.

## Try this prompt

Take one task I run every week with AI and turn it into a cost-per-finished-task measurement. Define what finished means as a check that can run without me. List every cost I should log per run, including tokens, tool calls, retries, wall-clock time, and my own repair minutes. Then design the smallest comparison I could run this week across two effort levels of the same model, and tell me what result would justify moving to a cheaper or more expensive lane.

## Related on this site

- [The Interesting Part of GPT-5.6 Is the Routing Layer](https://www.bobzhu.tech/gpt-56-routing-layer-cerebras/) covers the lane-by-lane view of model choice that cost per finished task is meant to settle.
- [Build Something That Stays](https://www.bobzhu.tech/build-something-that-stays/) is the argument for the durable artefacts that quietly lower the denominator.
- [The Floor and Ceiling of AI](https://www.bobzhu.tech/the-floor-and-ceiling-of-ai/) tracks how far the open floor sits behind the frontier, which changes which lanes are even available to you.
- [AI Does Not Automatically Give You Time Back](https://www.bobzhu.tech/ai-does-not-automatically-give-you-time-back/) explains why the human repair minutes are the cost most workflows fail to count.
- [AI Evaluation Checklist](https://www.bobzhu.tech/ai-evaluation-checklist/) is the short review layer for deciding whether an output actually counts as accepted.

### Sources

- [Introducing Claude Opus 5](https://www.anthropic.com/news/claude-opus-5)
- [Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/)
- [GPT-5.6: Frontier intelligence that scales with your ambition](https://openai.com/index/gpt-5-6/)
- [Advancing the price-performance frontier with GPT-5.6](https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/)
- [MCP specification 2026-07-28](https://modelcontextprotocol.io/specification/2026-07-28)
- [MCP Tasks extension](https://modelcontextprotocol.io/extensions/tasks/overview)
