Skip to content

Service

Observability & SRE

Know it broke before the customer does — and know which layer, in one click.

PrometheusGrafanaLokiTempoThanosAlertmanagerOpenTelemetryNetdata
One pane
Cloud, on-prem and edge in one view
SLOs
Error budgets, not uptime theatre
Runbooks
Every alert links to what to do

Most teams have monitoring. Fewer have observability, and almost nobody has a pager that only fires for things worth waking up for.

The symptom is recognisable: hundreds of alerts, a channel everyone mutes, and incidents that are still discovered by a customer email.

What we fix first

Alert rationalisation. We pull the last six months of alerts and sort by how often each one was acted on. Alerts that nobody ever acted on are deleted — not tuned, deleted. What remains gets a severity, a routing rule and a runbook link. An alert without a runbook is a notification, and notifications do not belong on a pager.

Coverage gaps. The signals infrastructure actually fails on, which generic agents do not collect:

  • BGP session state, received prefix count, and RPKI validity per peer
  • Cache hit ratio and origin offload per POP
  • Ceph placement-group health, OSD latency, near-full ratios
  • Hypervisor memory ballooning and noisy-neighbour steal time
  • GPU memory pressure, queue depth and tokens per second per model
  • Certificate expiry across every endpoint, including the internal ones

Correlation. Metrics in Prometheus, logs in Loki, traces in Tempo, one Grafana. Exemplars linking a latency spike to the trace that caused it, and labels consistent enough that “show me this host” works across all three.

SLOs and error budgets

An SLO is only useful if someone is willing to change their behaviour when it burns. We run the sessions that get engineering and product to agree on a target, then wire the burn-rate alerts.

# Multi-window burn-rate alerting: fast burn pages, slow burn files a ticket.
- alert: EdgeAvailabilityFastBurn
  expr: |
    (slo:error_ratio:rate5m{service="edge"} > 14.4 * 0.001)
    and
    (slo:error_ratio:rate1h{service="edge"} > 14.4 * 0.001)
  for: 2m
  labels: { severity: page }
  annotations:
    runbook: "https://runbooks.internal/edge-availability"

Two windows, so a brief blip does not page and a sustained degradation does. One SLO per user-visible journey, not one per microservice.

Incident practice

On-call rota with humane rotation and a real escalation path. A declared incident has a commander, a scribe and a channel. Postmortems are blameless, written within a week, and produce tracked actions — otherwise the same incident returns next quarter wearing a different hat.

We will run this for you on retainer, or build it and hand it over. Handover includes training your engineers on the rota, not a document drop.

Who this is for

Teams who have been surprised by an outage they should have seen coming; teams whose observability bill has outgrown its usefulness; and anyone running across cloud, on-prem and edge who currently needs three dashboards and a guess.

Questions

Asked often enough to answer here.

We already have Datadog. Do we need this?

Possibly not — but most Datadog estates we see are expensive and noisy at the same time. The work is often rationalisation — cut cardinality, define SLOs, delete alerts nobody acts on. If cost is the driver, we will model a self-hosted Prometheus and Loki stack against your bill.

What makes infrastructure monitoring different from application monitoring?

The failure modes. APM tools do not watch BGP session state, prefix counts, RPKI validity, cache hit ratio per POP, Ceph placement-group health or GPU memory pressure. Those are the signals that tell you which layer is at fault.

Can you run our on-call?

Yes, on a managed-run retainer — named engineers, defined escalation, and a monthly review. Or we build the practice and your team holds the pager, which is the more common outcome.

How long until we see something useful?

First dashboards and the critical alert set typically land inside the first two to three weeks. SLO definition takes longer because it requires decisions from your side about what "good" means.

Next step

Tell us what breaks at 3am.

A 30-minute call with the engineers who would do the work — not a sales desk. We will tell you whether this is a bolt.sh problem or something you can fix in-house.