From the assessments4 min read
What the bill does not say
Your model invoice knows tokens and models. It knows nothing about workflows, teams or outcomes. Which is why it cannot answer the one question the board asks.
by Maik Tran
The total is rarely the problem. Open your model provider's invoice and you see an amount, and usually you also see that it is growing. The problem starts with the next question: on what, exactly.
What the invoice contains
A billing export from any of the major providers is remarkably precise, as long as you stay inside its categories. It generally knows:
- the model and the model version
- input, output and cached tokens, priced separately
- the project or organisation key the call ran under
- the timestamp
- depending on the provider, the endpoint type: chat, embedding, batch
That reliably answers which model cost how much. It is a clean answer to a question nobody inside the company asked.
What it does not contain
Missing is everything people actually ask about:
- which workflow triggered the call
- which team or cost centre sits behind it
- whether a usable result came out at the end
- how many calls belonged to a single user action
- whether the same request had already been answered once
Those five fields are not missing through carelessness. They do not exist at the provider because the provider cannot know them. A workflow is a concept inside your company, not a field in a third party's API log.
Why the API key does not close the gap
The usual first attempt is the key: one per team, and attribution is solved. In practice that holds for roughly a quarter.
Keys get shared, because a deployment finishes faster than an approval process. They migrate into shared services working for several teams, and from then on the key measures the service rather than the originator. About the quality of the result it says nothing at all.
A key answers who called. The question in the room is what was paid for.
Agents widen the gap
As long as a human asks and a model answers, the gap is uncomfortable but manageable. Agentic workflows invert that. A single delegated task breaks into planning, tool calls, intermediate checks, correction loops and a summary, and every one of those steps is a billed call.
On the invoice they appear as independent lines. They were one operation. If you are not recording at the point of consumption, you cannot reassemble them afterwards, because the bracket that joined them only ever existed at runtime.
Cost per outcome
The figure that actually matters is not a price per token but a price per outcome: what a reviewed contract costs, an answered ticket, a merged pull request. It has two factors, and the invoice knows neither.
- Cost per attempt
- Model choice, context length, cache hit rate, number of steps in the workflow.
- Attempts per outcome
- Quality of the task definition, abandonments, repeats, discarded answers, silent retries in the code.
The second factor is both the more expensive and the more invisible one. A workflow that only lands on the third attempt costs three times one that lands on the first. On the invoice the two look identical.
Estimating and measuring are different things
There is still a lot to extract from billing data: the level and the trend, the model mix, the accounts nobody was tracking any more, and a reasoned estimate per workflow once you have had the three or four conversations that go with it. That is exactly what our Kostenbild is. It needs no access to your systems and lands in days.
What it cannot deliver is cost per outcome. That is not a weakness of the work, it is a property of the data source. Sell an estimate as a measurement and you hand over a number the customer roughly knew already, which burns the second conversation.
Measuring means recording at the point of consumption, giving every call the bracket that makes it part of one operation, and attaching the outcome to it. Inside your environment, because otherwise the prompts leave the building, and the question of where they flow occupies every vendor review for months.
The shortest test
You do not need a project to know where you stand. One question is enough:
How much are you currently spending on AI, and who in the building can break that number down?
If the second half comes back with a name, you are fine. If it comes back with a pause, the invoice is not the problem. The problem is that nobody can defend it the moment somebody asks.
can you break down your ai bill?
If not, the Kostenbild is the cheapest way to find out. I'll get back to you within 24 hours.

