bolt.sh platform
Building our own anycast CDN on a /24 we own
We acquired IPv4 space, got an ASN, announced the same prefix from POPs on AWS and Vultr with BIRD, and ran the edge ourselves. Here is what it took and what broke.
We did not build this for a client. We built it because every conversation about edge delivery eventually reaches a point where the honest answer is “it depends on what your prefix does under failure”, and we were tired of answering from theory.
The goal
One IPv4 prefix, announced from multiple points of presence across two different providers, serving cached content, with the ability to drain any single site by withdrawing its announcement.
Constraints we set ourselves: no provider-assigned address space, because the whole point is portability. No managed BGP abstraction — real sessions, real filters, real failure modes. And the entire thing reproducible from a Git repository, because an edge you cannot rebuild is an edge you cannot operate.
Getting the address space
This was the slow part, and it is worth saying plainly: the engineering took days, the paperwork took months.
An ASN and a /24 require RIR membership, a justification, and in practice an IPv4 transfer rather than a fresh allocation. Then the records that make the prefix usable by anyone running validation:
- IRR objects so networks filtering on route objects accept the announcement.
- RPKI ROAs binding the prefix to our ASN with a maximum prefix length, so validating networks treat it as valid rather than unknown.
A prefix with no ROA is increasingly treated as second-class. A prefix with a wrong ROA is dropped entirely. We got this wrong once in a staging announcement, which was a cheap and instructive way to learn it.
POPs
Each site is the same Terraform module with different inputs — region, transit peer, local cache sizing. The module brings up compute, attaches the prefix, configures BIRD, and deploys the NGINX edge via Ansible.
# Each POP announces the same /24. Health drives whether it announces at all.
protocol static edge_origin {
ipv4;
route 203.0.113.0/24 blackhole {
bgp_large_community.add((OURASN, 100, 1));
};
}
protocol bgp upstream {
local as OURASN;
neighbor 198.51.100.1 as 64500;
ipv4 {
import filter transit_in; # RPKI-validated, length-filtered
export where source = RTS_STATIC; # only our own prefix, nothing else
};
}
The export where source = RTS_STATIC line is the one that matters. Announce only what you originate. Every large routing incident in the last decade has an export filter at the centre of it.
What broke
Cold POPs hammered the origin. Bringing a new site into the announcement meant an empty cache serving live traffic, and the origin felt all of it at once. The fix was a mid-tier shield: new edges fill from the shield, not from origin, so a cold site costs the shield some work and the origin almost nothing.
Convergence is not instant. Withdrawing a prefix moves traffic in seconds, not milliseconds, and in-flight TCP connections to the drained site die. Anything long-lived — large downloads, streaming, websockets — needs connection draining before withdrawal, not instead of it.
Health checks that only test the process are useless. Our first version withdrew the announcement if NGINX stopped. It did not withdraw when NGINX was running and returning errors because the shield was unreachable. The check now exercises the full path — a real request for a real object, through the real cache tier.
Monitoring had a blind spot. Host and container metrics said everything was fine while one site had lost half its received prefixes. We now scrape BGP session state, received and accepted prefix counts, and ROA validity per peer, and alert on all three.
What we took from it
The operational properties are genuinely different from DNS-based steering. Draining a site is a routing change, not a TTL wait. But the debugging surface is larger: a problem can exist only for users whose packets take one path, and reproducing it requires looking from the right place on the internet.
We would build it again for a workload with high egress or residency constraints. We would not build it for a site serving a few terabytes a month — a commercial CDN with well-designed cache keys wins that comparison easily, and we tell clients so.
This is the system the CDN & edge practice is built on, and it is why the network engineering conversations start from operating experience rather than vendor documentation.
- /24
- IPv4 block we own and announce
- 2 clouds
- AWS and Vultr in one anycast fabric
- RPKI
- Signed ROAs on every announcement
- BIRD 2
- eBGP sessions we operate ourselves