GeekyAnts internal platform
Running OpenStack as an internal private cloud
Self-service compute for engineering teams on hardware we operate — Nova, Ceph, Terraform tenancy and the operational lessons that only show up after month three.
The brief was simple to state and easy to get wrong: give engineering teams the ergonomics of public cloud on hardware we own.
Not a hypervisor cluster where someone files a ticket for a virtual machine. An API, a quota, a Terraform provider and a console.
Why OpenStack
We evaluated running plain Proxmox and building tenancy on top. For a small estate that is the right answer — less to operate, fewer moving parts, and the UI is genuinely good.
We chose OpenStack because the requirement was multi-tenant self-service with quotas and a mature Terraform provider, and because teams already wrote Terraform for public cloud. The transition cost for a developer should be a provider block, not a new mental model.
The trade is operational weight. OpenStack has many services and they all have opinions. Anyone telling you it is low-maintenance has not run it.
The shape of it
- Compute. Nova on KVM. Flavours deliberately few — a long flavour list is how capacity planning becomes impossible.
- Storage. Ceph for block (Cinder) and object (RADOS Gateway, S3-compatible). Replicated across racks and power feeds, not just across hosts. Applications written against S3 did not need changing.
- Network. Routed-first. VXLAN where tenants genuinely needed L2 adjacency, which was less often than anyone claimed. The same BGP discipline we apply at the edge applies here — the fabric is a network, and network failures are the ones that cascade.
- Images. Built with Packer, versioned, scanned, and rotated. No snowflake images, no “the one Ravi made”.
- Tenancy. Each team gets a project, a quota and credentials. Terraform state in a shared backend, provisioning reviewed in merge requests like everything else.
provider "openstack" {
cloud = "internal"
}
resource "openstack_compute_instance_v2" "worker" {
count = 4
name = "build-worker-${count.index}"
image_name = "ubuntu-24.04-hardened-2025.09"
flavor_name = "c4.m16"
key_pair = var.key_pair
network { name = "team-platform" }
# Capacity planning only works if everything is attributable.
metadata = {
team = "platform"
cost_centre = "eng-infra"
}
}
That metadata block looks like bureaucracy. It is the reason we can answer “what does this team actually consume” without a spreadsheet archaeology session.
What we learned after month three
Quotas are a social system, not a technical one. The first quota model was generous because nobody wanted to be the team that blocked work. Within a quarter, utilisation was high and nobody could tell whether that was real demand or abandoned instances. We now expire untagged instances and report consumption per team monthly. Reclaim is mostly automatic and nobody has complained.
Ceph tells you before it hurts, if you listen. Placement-group health, OSD commit latency and near-full ratios are leading indicators. We added them to the critical alert set after one near-full warning that nobody noticed for two days.
Upgrades need a rehearsal environment. OpenStack upgrades touch many services. A staging cloud that mirrors the production topology — smaller, same versions, same configuration management — turned upgrade weekends into upgrade afternoons.
Noisy neighbours are real. Steal time on oversubscribed hypervisors produced latency complaints that looked like application bugs for an embarrassingly long time. CPU pinning for latency-sensitive workloads and a dedicated flavour class fixed it; monitoring steal time per instance is what found it.
Hybrid, not replacement
This did not replace public cloud and was never intended to. Steady-state compute, build infrastructure and storage sit here. Managed databases, regions where we have no footprint, and genuinely spiky workloads stay where elasticity is cheaper than ownership.
The connection between them is the part that needed the most design attention — consistent addressing, shared identity, and monitoring that spans both so an incident does not start with an argument about which side is broken.
This platform is what the on-prem and private cloud practice is built from, and it is where the GPU inference work runs.
- Self-serve
- Teams provision via API and quota
- Ceph
- Replicated block and S3-compatible object
- Terraform
- Same workflow as public cloud
- Hybrid
- Routed to public cloud regions
Practices involved
- Infrastructure ManagementWe take the estate you already have and make it legible, automated and boring.
- Observability & SREKnow it broke before the customer does — and know which layer, in one click.
- On-Prem & Private CloudA private cloud that behaves like a public one — self-service, API-driven, and yours.