Skip to content

Service

On-Prem AI Inference

Run your models on hardware you control — for cost, for latency, or because the data cannot leave.

vLLMTensorRT-LLMTritonRay ServeNVIDIA MIGKubernetesOllamaOpenStack
Zero egress
Prompts and data never leave your network
Per-token
Cost modelled against hosted API pricing
Dedicated
Your GPUs, your queue, no noisy neighbours

The case for running inference yourself is rarely ideological. It is one of three things: the per-token bill has outgrown a hardware purchase, the round trip to a hosted API is too slow for the product, or the data is not allowed to leave.

We build the infrastructure for all three, and we tell you when none of them apply.

Sizing before buying

GPU purchases are the easiest way to waste a large amount of money quickly. The first phase is measurement, not procurement.

  • Workload profile. Requests per second at peak and at median, prompt and completion length distribution, concurrency, and the p95 latency your product can tolerate.
  • Model selection. Parameter count, context window, quantisation tolerance. A 4-bit quantised model that fits on one card and meets your quality bar beats an unquantised model that needs four.
  • Memory maths. Weights plus KV cache at your concurrency and context length. KV cache is what actually determines how many concurrent sessions a card holds, and it is what people forget.
  • Utilisation. If the cluster is idle sixteen hours a day, the business case is weaker than the spreadsheet suggests — unless you can fill those hours with batch work, fine-tuning or evaluation runs.

Output is a sizing document with a hardware specification, an expected cost per million tokens, and the break-even volume against the hosted API you use today.

Serving

# vLLM: continuous batching is the single biggest throughput lever.
from vllm import LLM, SamplingParams

llm = LLM(
    model="/models/llama-3.3-70b-instruct-awq",
    quantization="awq",
    tensor_parallel_size=2,
    gpu_memory_utilization=0.92,
    max_model_len=16384,
    enable_prefix_caching=True,   # huge win for shared system prompts
)

We run vLLM for most text workloads — continuous batching and paged attention are what make dedicated hardware economical. TensorRT-LLM where the last increment of latency matters and the engineering cost is justified. Triton where you are serving a mix of model types. Ray Serve when the topology is genuinely multi-stage.

In front of that: an OpenAI-compatible gateway so application code does not care where inference happens, request routing by model and priority, token accounting per team, and overflow to a hosted provider above a threshold.

Multi-tenancy and scheduling

A shared GPU cluster with no scheduling discipline becomes one team’s cluster. We implement MIG partitioning or time-slicing where the cards support it, per-team quotas and queues, priority classes so interactive traffic preempts batch, and accounting so the cost is attributable.

Private RAG

Where the driver is data sensitivity, serving the model is only half the problem. Embeddings, the vector store, the document pipeline and the logs all contain the data you were trying to protect. We build the full path inside your boundary — ingestion, embedding, retrieval, serving and observability — with prompt and completion logging that respects the same classification as the source documents.

Who this is for

Teams with meaningful, steady inference volume; products where inference latency is in the user’s critical path; and organisations in regulated or sovereignty-constrained sectors where a hosted API is not an option regardless of price.

Questions

Asked often enough to answer here.

When does on-prem inference beat an API?

When volume is high and steady, when latency to your users matters more than model frontier quality, or when the data genuinely cannot leave your network. At low or spiky volume, hosted APIs win on cost and on operational simplicity — and we will say so.

Can we run frontier-quality models on our own hardware?

You can run very capable open-weight models. You cannot run the frontier closed models. Most production workloads — classification, extraction, summarisation, retrieval-augmented answering, code assistance — are well served by open weights, and that is where the economics work. Keep a hosted frontier model for the hard tail.

What hardware do we need?

It depends entirely on model size, context length, concurrency and your p95 latency target. That sizing exercise is the first deliverable — specifying hardware before measuring the workload is how people end up with idle GPUs.

Can this be hybrid?

Yes, and it usually is. Steady baseline load on owned GPUs, burst to a hosted API or rented cloud GPUs above a threshold. The routing layer is part of what we build.

Next step

Tell us what breaks at 3am.

A 30-minute call with the engineers who would do the work — not a sales desk. We will tell you whether this is a bolt.sh problem or something you can fix in-house.