Service
On-Prem & Private Cloud
A private cloud that behaves like a public one — self-service, API-driven, and yours.
- OpenStack
- Private cloud we design, build and operate
- Hybrid
- On-prem and public cloud on one network
- Capex
- Predictable cost per workload, no egress surprises
Public cloud is excellent at elasticity and expensive at steadiness. Once a workload runs at predictable utilisation for years, you are renting a machine at a large multiple of what the machine costs — plus egress.
Private cloud is not a rejection of that. It is a decision about which workloads belong where, made with numbers.
The model we build
We operate private clouds on OpenStack, which gives teams the thing they actually miss when they leave public cloud: self-service. Developers get an API, a quota and a console. They do not file a ticket to get a virtual machine.
- Compute. KVM under OpenStack Nova, or Proxmox where the estate is smaller and the operational overhead of OpenStack is not justified.
- Storage. Ceph for block and object, replicated across racks and power feeds. S3-compatible object storage so applications written for public cloud do not need rewriting.
- Network. Routed-first fabric, VXLAN where L2 adjacency is genuinely required, and the same BGP discipline we bring to the edge. Private cloud networking failures are network failures, and they are the ones that hurt.
- Automation. Terraform providers for the OpenStack layer so tenants provision the same way they do in public cloud, and Ansible for the hosts underneath.
Hybrid is the normal outcome
Almost nobody goes all the way. The common end state is steady-state compute, storage and inference on owned hardware, with public cloud retained for burst capacity, managed databases, and regions where you have no footprint.
That makes the connection between the two the critical path. Private interconnect or well-engineered tunnels, consistent addressing, shared identity, and monitoring that spans both so an incident does not require two dashboards and a guess about which side is broken.
Repatriation, modelled honestly
Before anything moves, we produce:
- Cost per workload today, including egress, support plans and the parts of the bill nobody reads.
- Cost per workload on owned hardware, including depreciation, colocation, remote hands, spares and the engineering time to run it.
- The crossover point, and the workloads that should never move.
- Migration sequence and rollback plan, because the first attempt will surface something the audit missed.
If the numbers say stay, we say stay. The audit is still worth having.
Who this is for
Teams with large, steady compute or storage footprints; anyone whose egress bill has become a line item leadership asks about; organisations with data-residency obligations that public cloud regions do not satisfy; and teams running GPU workloads, where owning the hardware changes the economics fastest.
Questions
Asked often enough to answer here.
Is leaving the cloud actually cheaper?
For steady-state, high-utilisation workloads with heavy egress — often substantially. For spiky, low-utilisation workloads — rarely. We model it against your actual usage before recommending anything, including the staffing cost of running hardware.
We do not want to run a datacentre.
You would not be. Colocation means someone else provides space, power, cooling and remote hands. You own the hardware and the configuration. We handle specification, buildout and ongoing operation.
What about hardware failure?
Design for it. N+1 on anything stateful, Ceph replication across failure domains, spares on the shelf, and a support contract with four-hour response where it is justified. A failed disk should be a ticket, not an incident.
Can we run Kubernetes on this?
Yes — that is the common pattern. OpenStack or bare metal underneath, Kubernetes on top, and the same manifests you run in public cloud. The goal is that workloads do not know or care where they are.
Proof
Where we have done this.
GeekyAnts internal platform
Standing up on-prem LLM inference next to the private cloud
Serving open-weight models on owned GPUs with vLLM — sizing against KV cache, multi-tenant scheduling, and an OpenAI-compatible gateway in front.
- vLLM
- Continuous batching and prefix caching
- No egress
- Prompts and documents stay in-network
- OpenAI API
- Drop-in gateway for existing code
- Quota
- Per-team token accounting
GeekyAnts internal platform
Running OpenStack as an internal private cloud
Self-service compute for engineering teams on hardware we operate — Nova, Ceph, Terraform tenancy and the operational lessons that only show up after month three.
- Self-serve
- Teams provision via API and quota
- Ceph
- Replicated block and S3-compatible object
- Terraform
- Same workflow as public cloud
- Hybrid
- Routed to public cloud regions
Also
The rest of the estate.
- InfrastructureWe take the estate you already have and make it legible, automated and boring.
- NetworkRouting, peering and address space — designed on paper, built as code, proven by withdrawal tests.
- CDN & EdgeYour own content network on your own address space — or a sane configuration of someone else's.
- SecurityReduce the number of ways in, then prove what happened on the ones that remain.
- AI InferenceRun your models on hardware you control — for cost, for latency, or because the data cannot leave.
- Device & MDMEvery laptop and phone enrolled, encrypted, patched and accounted for — from unboxing to offboarding.
- ObservabilityKnow it broke before the customer does — and know which layer, in one click.
- All servicesOverview, delivery method and the things we will tell you not to buy.
Next step
Tell us what breaks at 3am.
A 30-minute call with the engineers who would do the work — not a sales desk. We will tell you whether this is a bolt.sh problem or something you can fix in-house.