A year of billing data across nineteen client accounts, and the three things that actually moved cost per run. None of them was the model price.

We publish cost per run to every client every month, which means we have four thousand eight hundred and twelve rows of it to look at. The pattern is duller than the discourse suggests.
On the median workflow, fetching and shaping context costs more than the completion does. Teams that tune the prompt first are optimising the cheaper half.
The single biggest saving we have made on any account came from caching a lookup that ran four times per ticket and never changed inside a day.
Failed runs cost roughly what successful ones do, and a workflow with a five percent retry rate carries that overhead on every invoice. Fixing the connector timeout was worth more than switching model tier.
A run that stops early and routes to a person costs a fraction of one that reasons its way to a wrong answer somebody then has to unpick. The human gate pays for itself twice.
Four cents on an average run, across 1,406 runs on an average day. The figure has fallen every quarter since 2024 and almost none of that came from cheaper tokens.
Two weeks watching the real work, with the eval suite written before anything ships.
Two weeks. If we can’t find 500 hours a year to return, it’s free.
hello@agentiq.studio
+1 (415) 555-0138
410 Fremont Street, Suite 300
San Francisco, CA