Anycast without the hand-waving
What anycast actually gives you, what it costs operationally, and the four failure modes that surprise teams building their first multi-POP edge.
Anycast gets described as “the same IP in many places”, which is true and useless. The useful description is: you give up control of routing to the internet, and in exchange you get failover that does not wait for DNS.
Whether that trade is good depends entirely on what your traffic looks like.
What you actually get
Drain by withdrawal. To take a site out of service, stop announcing the prefix from it. Traffic moves to the next-closest site within BGP convergence — seconds, not a TTL. This is the single strongest argument for anycast and it is an operational argument, not a performance one.
No client-side logic. No geolocation database to keep current, no DNS steering service, no SDK. The routing system already solves “nearest” better than a GeoIP lookup does, because it optimises for the actual network topology rather than physical distance.
One address to document. Firewall allowlists, partner integrations and customer configuration reference one prefix forever, regardless of how many sites you add behind it.
What it costs you
You cannot pin a user to a site. Which POP a given user reaches is decided by networks you do not control, and it can change mid-session when an upstream reconverges. Anything stateful at the edge has to tolerate that.
Debugging requires vantage points. A problem that only exists for users whose packets traverse one particular transit path is invisible from your office. You need looking glasses, RIPE Atlas measurements, or probes in the networks that matter to you.
Cache coherence gets harder. Every site has its own cache. Purge has to reach all of them, and “all of them” includes the one that was down during the purge and came back with stale content.
Four things that surprise people
1. Convergence is not instant, and in-flight connections die. Withdrawing a prefix moves new flows quickly. Existing TCP connections to the drained site are gone. For short HTTP requests nobody notices; for large downloads, websockets or streaming, you need connection draining before withdrawal, with enough lead time for sessions to finish.
2. A cold POP will hurt your origin. A new site with an empty cache serving live traffic sends every request upstream. Without a mid-tier shield, bringing up a POP is a self-inflicted origin load test. Fill new edges from a shield, not from origin.
3. Process-level health checks lie. “Is the web server running” does not catch a web server that is running and returning errors because its upstream is unreachable. Health checks that drive route withdrawal must exercise the full path — a real request for a real object through the real cache tier. Otherwise you keep announcing from a site that cannot serve.
4. Your monitoring probably does not watch routing. Host and container metrics will show a perfectly healthy POP that has silently lost half its received prefixes. Scrape BGP session state, received and accepted prefix counts, and RPKI validity per peer. Alert on all three. We learned this one the slow way.
The export filter
One paragraph of this post matters more than the rest. Announce only what you originate.
protocol bgp upstream {
local as OURASN;
neighbor 198.51.100.1 as 64500;
ipv4 {
import filter transit_in;
export where source = RTS_STATIC; # only our own prefix
};
}
Every large routing incident of the last decade has a permissive export policy somewhere near the centre of it. An explicit originate list, a maximum prefix length, and a max-prefix limit on the session are not advanced configuration. They are the baseline.
When not to do this
If you serve a few terabytes a month, a commercial CDN with well-designed cache keys will beat anything you build, on cost and on your engineers’ time. Anycast starts to make sense at high egress volume, with unusual cache semantics, under residency constraints, or when you need the edge to run your own compute.
We built our own to find out where that line is. Here is what it took.
Keep reading
Where on-prem inference actually beats the API bill
A framework for deciding whether to buy GPUs — utilisation, KV cache maths, the hidden operational cost, and the three cases where the answer is obviously yes.
The RPKI and IRR checklist nobody hands you with your first prefix
You received an allocation. Here are the records, filters and tests that decide whether the internet accepts your routes or quietly drops them.