Service
Infrastructure Management
We take the estate you already have and make it legible, automated and boring.
- Day 1
- Written inventory of everything you run
- IaC
- No hand-edited hosts, routers or firewalls
- 24×7
- Named on-call rota with escalation path
Most infrastructure problems are not exotic. They are the accumulated result of nobody owning the whole picture: a VPC somebody built in 2021, a firewall rule nobody will delete, a backup job that has been failing silently since a credential rotated, three different ways to provision a host.
We start by making the estate legible, then we automate it, then we run it.
Where we usually start
The first two to three weeks are a fixed-scope audit. No tooling to install, no agents to deploy — read access and a few conversations.
- Inventory. Cloud accounts, regions, VPCs and peerings. Physical: racks, circuits, cross-connects, PDUs. Network: ASNs, prefixes, BGP sessions, DNS zones. Identity: who can reach what, and through which path.
- Failure domains. What is single-homed. What shares a power feed, a hypervisor, an availability zone or a DNS provider. Where a single expired certificate takes the business offline.
- Spend. Cost per workload rather than cost per service line. Egress is almost always the surprise.
- Operational debt. Patch lag, unsupported OS versions, backups that have never been restored, runbooks that reference people who left.
You get a written report with a prioritised remediation list. It is useful whether or not you continue with us.
Infrastructure as code, properly
“We use Terraform” and “our infrastructure is code” are different claims. The second one means a destroyed environment can be rebuilt from a Git repository, and that drift is detected rather than discovered.
# Every environment is the same module with different inputs.
module "edge_pop" {
source = "../../modules/edge-pop"
site = "blr1"
transit_peers = var.transit_peers["blr1"]
announce = ["203.0.113.0/24"]
rpki_validate = true
bird_config = file("${path.module}/bird/blr1.conf")
monitoring_tier = "tier-1"
}
Provisioning in Terraform, configuration in Ansible, images in Packer, everything reviewed in merge requests. Golden images rather than long-lived pets. Secrets in Vault or your cloud’s KMS — never in the repo, never in a shared vault spreadsheet.
Running it
A managed-run retainer covers monitoring and alert routing, patch windows with tested rollback, quarterly disaster-recovery drills against real backups, capacity review before you hit a wall rather than after, and a monthly service review with the engineers who hold the pager.
We publish SLOs, not uptime theatre. If an error budget is burning, you hear it from us first.
Who this is for
Teams between roughly ten and five hundred engineers who have outgrown “whoever set it up maintains it” but do not want a dedicated infrastructure department. Also teams running regulated workloads who need evidence — change records, access logs, restore tests — rather than assurances.
Questions
Asked often enough to answer here.
Do you replace our existing ops team?
Usually not. Most engagements run alongside an in-house team — we take the parts nobody has time for (network, edge, DR testing, patch cadence) and hand back documented, automated systems.
We are entirely on AWS. Is this still relevant?
Yes. A large share of the work is making single-cloud estates cheaper and more survivable — reserved capacity modelling, egress reduction, multi-AZ correctness and an exit path that is real rather than theoretical.
Can you work inside our change management process?
Yes. We work in your ticket queue, your change windows and your approval chain. If you do not have one, we will help you build the lightest version that passes audit.
Also
The rest of the estate.
- NetworkRouting, peering and address space — designed on paper, built as code, proven by withdrawal tests.
- CDN & EdgeYour own content network on your own address space — or a sane configuration of someone else's.
- SecurityReduce the number of ways in, then prove what happened on the ones that remain.
- On-PremA private cloud that behaves like a public one — self-service, API-driven, and yours.
- AI InferenceRun your models on hardware you control — for cost, for latency, or because the data cannot leave.
- Device & MDMEvery laptop and phone enrolled, encrypted, patched and accounted for — from unboxing to offboarding.
- ObservabilityKnow it broke before the customer does — and know which layer, in one click.
- All servicesOverview, delivery method and the things we will tell you not to buy.
Next step
Tell us what breaks at 3am.
A 30-minute call with the engineers who would do the work — not a sales desk. We will tell you whether this is a bolt.sh problem or something you can fix in-house.