Service
On-Prem AI Inference
Run your models on hardware you control — for cost, for latency, or because the data cannot leave.
- Zero egress
- Prompts and data never leave your network
- Per-token
- Cost modelled against hosted API pricing
- Dedicated
- Your GPUs, your queue, no noisy neighbours
The case for running inference yourself is rarely ideological. It is one of three things: the per-token bill has outgrown a hardware purchase, the round trip to a hosted API is too slow for the product, or the data is not allowed to leave.
We build the infrastructure for all three, and we tell you when none of them apply.
Sizing before buying
GPU purchases are the easiest way to waste a large amount of money quickly. The first phase is measurement, not procurement.
- Workload profile. Requests per second at peak and at median, prompt and completion length distribution, concurrency, and the p95 latency your product can tolerate.
- Model selection. Parameter count, context window, quantisation tolerance. A 4-bit quantised model that fits on one card and meets your quality bar beats an unquantised model that needs four.
- Memory maths. Weights plus KV cache at your concurrency and context length. KV cache is what actually determines how many concurrent sessions a card holds, and it is what people forget.
- Utilisation. If the cluster is idle sixteen hours a day, the business case is weaker than the spreadsheet suggests — unless you can fill those hours with batch work, fine-tuning or evaluation runs.
Output is a sizing document with a hardware specification, an expected cost per million tokens, and the break-even volume against the hosted API you use today.
Serving
# vLLM: continuous batching is the single biggest throughput lever.
from vllm import LLM, SamplingParams
llm = LLM(
model="/models/llama-3.3-70b-instruct-awq",
quantization="awq",
tensor_parallel_size=2,
gpu_memory_utilization=0.92,
max_model_len=16384,
enable_prefix_caching=True, # huge win for shared system prompts
)
We run vLLM for most text workloads — continuous batching and paged attention are what make dedicated hardware economical. TensorRT-LLM where the last increment of latency matters and the engineering cost is justified. Triton where you are serving a mix of model types. Ray Serve when the topology is genuinely multi-stage.
In front of that: an OpenAI-compatible gateway so application code does not care where inference happens, request routing by model and priority, token accounting per team, and overflow to a hosted provider above a threshold.
Multi-tenancy and scheduling
A shared GPU cluster with no scheduling discipline becomes one team’s cluster. We implement MIG partitioning or time-slicing where the cards support it, per-team quotas and queues, priority classes so interactive traffic preempts batch, and accounting so the cost is attributable.
Private RAG
Where the driver is data sensitivity, serving the model is only half the problem. Embeddings, the vector store, the document pipeline and the logs all contain the data you were trying to protect. We build the full path inside your boundary — ingestion, embedding, retrieval, serving and observability — with prompt and completion logging that respects the same classification as the source documents.
Who this is for
Teams with meaningful, steady inference volume; products where inference latency is in the user’s critical path; and organisations in regulated or sovereignty-constrained sectors where a hosted API is not an option regardless of price.
Questions
Asked often enough to answer here.
When does on-prem inference beat an API?
When volume is high and steady, when latency to your users matters more than model frontier quality, or when the data genuinely cannot leave your network. At low or spiky volume, hosted APIs win on cost and on operational simplicity — and we will say so.
Can we run frontier-quality models on our own hardware?
You can run very capable open-weight models. You cannot run the frontier closed models. Most production workloads — classification, extraction, summarisation, retrieval-augmented answering, code assistance — are well served by open weights, and that is where the economics work. Keep a hosted frontier model for the hard tail.
What hardware do we need?
It depends entirely on model size, context length, concurrency and your p95 latency target. That sizing exercise is the first deliverable — specifying hardware before measuring the workload is how people end up with idle GPUs.
Can this be hybrid?
Yes, and it usually is. Steady baseline load on owned GPUs, burst to a hosted API or rented cloud GPUs above a threshold. The routing layer is part of what we build.
Also
The rest of the estate.
- InfrastructureWe take the estate you already have and make it legible, automated and boring.
- NetworkRouting, peering and address space — designed on paper, built as code, proven by withdrawal tests.
- CDN & EdgeYour own content network on your own address space — or a sane configuration of someone else's.
- SecurityReduce the number of ways in, then prove what happened on the ones that remain.
- On-PremA private cloud that behaves like a public one — self-service, API-driven, and yours.
- Device & MDMEvery laptop and phone enrolled, encrypted, patched and accounted for — from unboxing to offboarding.
- ObservabilityKnow it broke before the customer does — and know which layer, in one click.
- All servicesOverview, delivery method and the things we will tell you not to buy.
Next step
Tell us what breaks at 3am.
A 30-minute call with the engineers who would do the work — not a sales desk. We will tell you whether this is a bolt.sh problem or something you can fix in-house.