Where on-prem inference actually beats the API bill
A framework for deciding whether to buy GPUs — utilisation, KV cache maths, the hidden operational cost, and the three cases where the answer is obviously yes.
“Should we run our own inference?” is usually asked after a monthly bill arrives and it is usually answered with a spreadsheet that compares the wrong things.
Here is the framework we use, including the parts that make the answer no.
Three reasons to own hardware, in order of strength
1. The data cannot leave. If a regulator, a contract or a client’s security review says prompts and documents stay inside your network, this is not an economic decision. Price it, build it, move on. This is the strongest case and the one where we see the most genuine need.
2. Latency is in the user’s critical path. If a human is waiting on the first token and your users are far from the nearest region of your API provider, local inference can be the difference between a product that feels responsive and one that does not. Note that this argues for edge placement, not necessarily for ownership.
3. The bill exceeds the hardware. The weakest-looking reason and the most common trigger. It is also the one most often gotten wrong, because the comparison usually omits utilisation.
The utilisation problem
A GPU costs the same whether it is busy or idle. A hosted API costs nothing when you are not calling it.
So the real comparison is not “cost per million tokens on our hardware” against “cost per million tokens from the API”. It is:
(hardware + colocation + power + spares + engineering time) ÷ (tokens you actually push through it)
If your traffic is business-hours-shaped in one timezone, you are paying for roughly 24 hours and using 8. That triples your effective cost per token and frequently flips the answer.
The fix is to fill the idle window: evaluation runs, embedding backfills, batch classification, fine-tuning experiments. If you have that work, the economics improve sharply. If you do not, be honest about it rather than assuming you will find some.
KV cache, not parameter count
The second common error is sizing from model size alone.
Weights are the fixed cost. The variable cost — the thing that decides how many concurrent users a card serves — is the KV cache, and it scales with context length times concurrency. A model that runs comfortably for one user can refuse the twentieth because every active session holds attention state.
Practical consequences:
- Long-context workloads need far more memory than the parameter count suggests.
- Reducing maximum context length is often cheaper than buying another card.
- A 4-bit quantised model that meets your quality bar on one card beats an unquantised one needing four. Evaluate quantisation against your tasks, not against published benchmarks.
- Prefix caching is close to free throughput if your workload shares long system prompts, which internal tooling almost always does.
The cost nobody puts in the spreadsheet
Someone has to run this. Driver and CUDA version management, node failures, model rollouts, queue tuning, capacity planning, and being awake when a card falls off the bus at 2am.
That is a real fraction of an engineer, ongoing. Include it. If the business case only works when you value operational time at zero, it does not work.
What the answer usually is
Hybrid, and not as a compromise — as the correct design.
Steady baseline load runs on owned hardware. Overflow above a concurrency threshold routes to a hosted API rather than queuing. Frontier-quality work — the hard tail where open weights are genuinely not good enough — stays on a hosted frontier model permanently.
Make that possible by putting an OpenAI-compatible gateway in front of everything from day one. Application code points at a base URL and a model name and never learns where the request went. That indirection is what lets you change the split later without touching a dozen repositories.
The honest test
Before buying anything, measure for two weeks: requests per second at peak and median, prompt and completion length distribution, concurrency, and p95 latency tolerance. Then size from that.
If the resulting utilisation figure is below roughly a third, and no regulatory or latency constraint applies, stay on the API. That recommendation has cost us work and it has never cost us a client.
We run this setup ourselves — here is how it is built.
Keep reading
The RPKI and IRR checklist nobody hands you with your first prefix
You received an allocation. Here are the records, filters and tests that decide whether the internet accepts your routes or quietly drops them.
Anycast without the hand-waving
What anycast actually gives you, what it costs operationally, and the four failure modes that surprise teams building their first multi-POP edge.