Service
Observability & SRE
Know it broke before the customer does — and know which layer, in one click.
- One pane
- Cloud, on-prem and edge in one view
- SLOs
- Error budgets, not uptime theatre
- Runbooks
- Every alert links to what to do
Most teams have monitoring. Fewer have observability, and almost nobody has a pager that only fires for things worth waking up for.
The symptom is recognisable: hundreds of alerts, a channel everyone mutes, and incidents that are still discovered by a customer email.
What we fix first
Alert rationalisation. We pull the last six months of alerts and sort by how often each one was acted on. Alerts that nobody ever acted on are deleted — not tuned, deleted. What remains gets a severity, a routing rule and a runbook link. An alert without a runbook is a notification, and notifications do not belong on a pager.
Coverage gaps. The signals infrastructure actually fails on, which generic agents do not collect:
- BGP session state, received prefix count, and RPKI validity per peer
- Cache hit ratio and origin offload per POP
- Ceph placement-group health, OSD latency, near-full ratios
- Hypervisor memory ballooning and noisy-neighbour steal time
- GPU memory pressure, queue depth and tokens per second per model
- Certificate expiry across every endpoint, including the internal ones
Correlation. Metrics in Prometheus, logs in Loki, traces in Tempo, one Grafana. Exemplars linking a latency spike to the trace that caused it, and labels consistent enough that “show me this host” works across all three.
SLOs and error budgets
An SLO is only useful if someone is willing to change their behaviour when it burns. We run the sessions that get engineering and product to agree on a target, then wire the burn-rate alerts.
# Multi-window burn-rate alerting: fast burn pages, slow burn files a ticket.
- alert: EdgeAvailabilityFastBurn
expr: |
(slo:error_ratio:rate5m{service="edge"} > 14.4 * 0.001)
and
(slo:error_ratio:rate1h{service="edge"} > 14.4 * 0.001)
for: 2m
labels: { severity: page }
annotations:
runbook: "https://runbooks.internal/edge-availability"
Two windows, so a brief blip does not page and a sustained degradation does. One SLO per user-visible journey, not one per microservice.
Incident practice
On-call rota with humane rotation and a real escalation path. A declared incident has a commander, a scribe and a channel. Postmortems are blameless, written within a week, and produce tracked actions — otherwise the same incident returns next quarter wearing a different hat.
We will run this for you on retainer, or build it and hand it over. Handover includes training your engineers on the rota, not a document drop.
Who this is for
Teams who have been surprised by an outage they should have seen coming; teams whose observability bill has outgrown its usefulness; and anyone running across cloud, on-prem and edge who currently needs three dashboards and a guess.
Questions
Asked often enough to answer here.
We already have Datadog. Do we need this?
Possibly not — but most Datadog estates we see are expensive and noisy at the same time. The work is often rationalisation — cut cardinality, define SLOs, delete alerts nobody acts on. If cost is the driver, we will model a self-hosted Prometheus and Loki stack against your bill.
What makes infrastructure monitoring different from application monitoring?
The failure modes. APM tools do not watch BGP session state, prefix counts, RPKI validity, cache hit ratio per POP, Ceph placement-group health or GPU memory pressure. Those are the signals that tell you which layer is at fault.
Can you run our on-call?
Yes, on a managed-run retainer — named engineers, defined escalation, and a monthly review. Or we build the practice and your team holds the pager, which is the more common outcome.
How long until we see something useful?
First dashboards and the critical alert set typically land inside the first two to three weeks. SLO definition takes longer because it requires decisions from your side about what "good" means.
Proof
Where we have done this.
GeekyAnts internal platform
Standing up on-prem LLM inference next to the private cloud
Serving open-weight models on owned GPUs with vLLM — sizing against KV cache, multi-tenant scheduling, and an OpenAI-compatible gateway in front.
- vLLM
- Continuous batching and prefix caching
- No egress
- Prompts and documents stay in-network
- OpenAI API
- Drop-in gateway for existing code
- Quota
- Per-team token accounting
GeekyAnts internal platform
Running OpenStack as an internal private cloud
Self-service compute for engineering teams on hardware we operate — Nova, Ceph, Terraform tenancy and the operational lessons that only show up after month three.
- Self-serve
- Teams provision via API and quota
- Ceph
- Replicated block and S3-compatible object
- Terraform
- Same workflow as public cloud
- Hybrid
- Routed to public cloud regions
Also
The rest of the estate.
- InfrastructureWe take the estate you already have and make it legible, automated and boring.
- NetworkRouting, peering and address space — designed on paper, built as code, proven by withdrawal tests.
- CDN & EdgeYour own content network on your own address space — or a sane configuration of someone else's.
- SecurityReduce the number of ways in, then prove what happened on the ones that remain.
- On-PremA private cloud that behaves like a public one — self-service, API-driven, and yours.
- AI InferenceRun your models on hardware you control — for cost, for latency, or because the data cannot leave.
- Device & MDMEvery laptop and phone enrolled, encrypted, patched and accounted for — from unboxing to offboarding.
- All servicesOverview, delivery method and the things we will tell you not to buy.
Next step
Tell us what breaks at 3am.
A 30-minute call with the engineers who would do the work — not a sales desk. We will tell you whether this is a bolt.sh problem or something you can fix in-house.