Skip to content
All case studies

GeekyAnts internal platform

Standing up on-prem LLM inference next to the private cloud

Serving open-weight models on owned GPUs with vLLM — sizing against KV cache, multi-tenant scheduling, and an OpenAI-compatible gateway in front.

vLLMNVIDIA MIGKubernetesOpenStackPrometheusGrafanaTerraform

Internal AI usage grew the way it does everywhere: a few experiments, then a few products, then a monthly bill that someone in finance started asking about. Alongside that, a recurring constraint in client conversations — some documents are not allowed to leave the network, regardless of a vendor’s data-processing agreement.

Both pointed the same direction. We put GPUs next to the private cloud and served open-weight models ourselves.

Sizing is about KV cache, not parameters

The most common sizing mistake is multiplying parameter count by bytes per weight and stopping there. Weights are the fixed cost. The variable cost — the one that decides how many concurrent users a card holds — is the KV cache, and it scales with context length and concurrency.

A model that fits comfortably in memory with one user can refuse the twentieth, because every active session is holding attention state. We sized from the workload backwards:

  • Request rate at peak, not at average.
  • Prompt and completion length distribution, taken from real traffic rather than estimated.
  • Concurrency target and p95 latency the products could tolerate.
  • Quantisation tolerance, evaluated against our own task set rather than published benchmarks.

A 4-bit quantised model that met the quality bar on one card beat an unquantised model needing several. That decision changed the hardware specification more than any other.

Serving

from vllm import LLM

llm = LLM(
    model="/models/qwen3-32b-instruct-awq",
    quantization="awq",
    tensor_parallel_size=2,
    gpu_memory_utilization=0.92,
    max_model_len=32768,
    enable_prefix_caching=True,
)

enable_prefix_caching was the single biggest throughput win. Internal tools share long system prompts, and caching the prefix across requests removes work that was otherwise repeated on every call.

Continuous batching is the other. Static batching leaves the GPU idle while it waits to fill a batch; continuous batching keeps it fed. This is why dedicated hardware can be economical at all.

The gateway matters more than the model

In front of the serving layer sits an OpenAI-compatible gateway. Application code points at a base URL and a model name. It does not know, and must not care, whether that request was served on our GPUs or forwarded to a hosted provider.

That indirection bought three things:

  • Overflow. Above a concurrency threshold, requests route to a hosted API rather than queuing. Owned hardware covers the baseline; the tail is rented.
  • Accounting. Tokens per team, per model, per application. Without it, “is this worth it?” is unanswerable.
  • Model migration. Swapping the model behind a name is a configuration change, not a code change across a dozen repositories.

Multi-tenancy

A shared GPU cluster without scheduling becomes the busiest team’s cluster. MIG partitioning where the cards supported it, priority classes so interactive traffic preempts batch, and per-team queues with quotas.

Batch work — evaluation runs, embedding backfills, fine-tuning experiments — fills the overnight window that would otherwise be idle. Utilisation is what makes the business case real, and idle hours are where business cases quietly die.

Observability

Standard infrastructure monitoring does not watch the things that go wrong here. We added GPU memory pressure and fragmentation, queue depth and wait time per model, tokens per second in and out, prefix cache hit ratio, and request failures by cause — out-of-memory, context overflow, timeout.

Queue wait time turned out to be the number that correlated with user complaints. Not latency, not throughput. Waiting.

Where we landed

Steady internal workloads run on owned hardware with no data leaving the network. Frontier-quality work and overflow go to hosted APIs. That split is the honest answer for most teams, and it is the one we recommend: on-prem inference is a cost and sovereignty decision, not a capability decision.

This is the platform behind the on-prem AI inference practice.

Next step

Tell us what breaks at 3am.

A 30-minute call with the engineers who would do the work — not a sales desk. We will tell you whether this is a bolt.sh problem or something you can fix in-house.