The first time you use a cloud AI API, the price is intoxicating. You ask it something useful and it costs a fraction of a penny. Your brain does the obvious multiplication, decides AI is basically free, and moves on. That moment of arithmetic is where a lot of expensive decisions get made.
Here is what the demo does not show you.
The meter never stops
Per-token pricing is a meter. It runs on every single interaction, forever. That is fine when three people use it now and then. It is a very different thing when four hundred people use it all day, every working day, for the next five years, and the meter is still running the whole time. You are not buying a tool. You are renting a habit, and the rent is due every month whether the tool got more valuable or not.
Owned things have a different shape. You pay once, and then usage is close to free. A meter has the opposite shape. Usage is the cost, and the more your people rely on AI, the more you pay for their reliance. Success makes the bill go up. Think about how strange that incentive is.
The half of the bill nobody quotes
There is a detail in the pricing that quietly doubles or triples real costs. Output tokens, the text the model writes back, cost several times more than input tokens across every major provider. Look at the published rates from Anthropic or Google and you will see the same pattern: output is where the money is. A good assistant is one that writes a lot back. So the better the tool does its job, the more of the expensive kind of token it generates. The pricing punishes exactly the behaviour you wanted.
Then there is retrieval. The moment your assistant answers questions using your own documents, every question drags a big slab of context through the model, and you pay for all of it. A one-line question can carry ten thousand tokens of retrieved material behind it. And agentic workflows, where the model works in loops, can multiply the count tenfold. None of this shows up when you are poking at a demo with single questions.
Do the annual maths, once
I am not going to tell you the cloud is always the wrong choice, because it is not. For light, occasional use, the cheap model tiers are genuinely hard to beat, and you should use them. The problem is not the cloud. The problem is signing up for a per-token meter without ever working out what it costs at full tilt.
So do it once. Take a real heavy user, estimate their honest daily token use including retrieval and any agents, multiply by the people who will actually use the tool hard, apply the real input and output rates separately, add growth because usage of a good tool goes up, and annualise. We walk through the method in more detail in our note on estimating your real token bill. Most people never run that number. It is usually two to five times the figure the demo implied.
Once it is in front of you, one question does the rest of the work. Would you rather pay that number every year, growing, forever? Or own the thing that produces the same answers, on hardware in your building, at a fixed cost that does not care how many questions your people ask? For a lot of companies, seeing the annual figure written down is the entire decision.
