What is a CDN? How do CDNs work?
animeshgaitonde.medium.com
animeshgaitonde.medium.com
- How do they distribute a huge load between multiple servers in a DC (some kind of level 2/3 load balancing I don't know of probably?)
- How is the cached data distributed between POP's?
- The whole TLS termination thing is probably an interesting aspect as well. The sensitive key material is needed on every node / POP, but you probably don't want all keys just lying around in clear text on disks? Some kind of distributed HSM-thingy?
- Just practically, how do you manage all these machines. Something like Ansible, or something more like Nomad/K8S?
Perhaps the fly.io people could write something about this (if they haven't already :-))
For example, if you are in country X, then even though there are dozens or hundreds of POP addresses for the name you requested, you’ll be answered with one for country X.
DNS gives more control and precision to CDN provider than BGP. BGP might be good enough if you got few POP around the world but with the scale/distribution of Akamai, DNS is better suited.
One is an actual routing system that tells routers where to send data, the other is a name translation system with multiple layers of caches outside of your control.
DNS is a layer above as BGP is still used to actually navigate to the listed IP, and any large CDN will own that IP space and announce their own routes anyway.
Yeah, I'm not sure what parent is on about, BGP and DNS are not alternatives to each other, the internet relies on both of them but at different layers. Without BGP packets wouldn't know how to be routed and without DNS they wouldn't know where to be routing to, they are complementary.
Answer: Nothing. They are just IPs.
Using a combination of Anycast and DNS is going to give the best control over steering http traffic. Particularly if you own a few prefixes and can do clever addressing tricks.
With DNS and dynamic responses you are directing request to specific DC, even server, almost on every request. It may be dedicated for this traffic type (live stream different than static images etc). Your DNS server can take the hostname ("www.google.com") into consideration - BGP doesn't even know the hostname in the URL. If you wanted to do it with BGP you would need to place specific content to a dedicated /24 subnet and that is impossible considering how many IPv4 addresses are available.
BGP doesn't even consider network latency, current network load. CDN knows load on their machines, on their network link, where given content is placed. The bottleneck may be storage, network or CPU processing, different for different sites and content type. They need to direct traffic on request basis considering this and at least the hostname from the URL. That's why DNS is used first.
Is this really true? Last time I was dealing with load-balancing and DNS queries, DNS was simply "round-robin" the replies, giving you back a random record basically of the ones replied. So if you have three A records with different IPs, each query will give you back one of them, but not depending on the location.
Maybe things have changed since I last dealt with it, but the DNS ecosystem doesn't tend to move very fast so I'm doubtful...
https://easydns.com/features/geo-dns/
(Having said that, I tried my toy site on cloudflare free tier from the UK and it gave me San Francisco IPs, so presumably they only do this for large enough customers)
It can, but it tends to be a premium feature of specific DNS providers, not a global/by-default feature of DNS as efitz seems to be alluding to.
DNSimple supports it for example, but only on their "Professional" plan (and they call it "Regional Records") while others like Gandi don't support it at all.
Basically you announce the same IP address from multiple locations, and BGP chooses the best* Route to reach that address.
* Note that "best" does not nessecarily mean "fastest", not even "least hops" - BGP has many knobs where the routing decisions can be manipulated, but by default it choses the route where traffic has to traverse the smallest number of "Autonomous Systems" (roughly equivalent to "Organizational Entities").
[1] https://www.routledge.com/A-Practical-Guide-to-Content-Deliv...
[2] https://research.google/pubs/pub43438 https://www.youtube.com/watch?v=0W49z8hVn0k
[3] https://research.google/pubs/pub49065 https://research.google/people/JohnWilkes
I wonder if things have changed since then
Something some people seem to miss though, is to leverage ETag header properly, so the client and CDN can serve fresh content automatically when it exists, or serve the cached content otherwise. It's not that tricky, but somehow many seems to not even know about it.
Some CDN provides distributed key value stores where you can push data in seconds, you can also deploy your own code in minutes (for real time changes you would push data, not code).
Some providers expose cache purge APIs so you can clear cached data in seconds.
They also allow serving cached content while asynchronously revalidating with the Origin so following users get a fresh version.
With proper configuration/architecture they are a great way to scale your real-time website.
That’s a neat fact and perhaps an advantage usually reserved for new entrants into an entrenched market. Presumably Fastly benefited, at least in some part, from technological innovation that existing companies could not apply as easily.
Most of the CDN providers predate all those things. Most are custom implementations with varying degrees of documentation and features.
> How is the cached data distributed between POP's?
I've seen two implementations and both do roughly the same thing. A request comes in and if it's not in cache it checks a regional POP preconfigured for that region like it was filling from origin. If it's not there the regional server gets it from origin. That fills the cache for both the regional one and the local. One implementation the regional POP was just config. Meaning the regional POP was just a standard POP that also served the requests from other POPs. In another it was something different only serving regional requests.
There are good publications and papers from fastly and facebook circa 2012 to 2016 at events like usenix. Akamai had some seminal works in the early to mid 2000s. Oh, and google maglev and I think cloudlfare talk about IP fragmentation and similar problems where you need persistance but dont get a 5 tuple.
In short: DNS and edns client subnet tell the CDNs name servers the clients network location. Map the client subnet and network topology to a nearby POP. Return those addresses in the A record response.
If you dont get edns0 use a weighted result of the resolvers clients.
The A record IPs get the client to the best available POP considering load, latency, bandwidth, cost, customer domains, etc.
At the pop you play network games to get the tcp session to the first layer of CDN hosts. Ill call them “layer 1” or L1. You almost definitely DO NOT run a full proxy with traffic 100% in:out at this layer. Its simply too wasteful. Instead you can use some combination of ARP, ECMP, or “layer 3 switching” to get tcp flows distributed among the available L1 hosts. These hosts will almost always share those initial “virtual” IP addresses amongst them. The L1 hosts will terminate TCP/TLS and probably parse the HTTP request to determine the customer domain, customer cache rules, etc. The L1 howts will have a local “hot” cache of popular objects, lots of duplication between howts to distribute load. Probably a 50-80% hit rate here to immediately return the result to the client.
If the L1 is a miss they will use a consistent hash function to map the object (think URI plus any customer rules) to an L2 host. The L2 is a segmented cache, maximizing unique bytes stored for content in the “middle” of the popularity distribution. Implementation varies here; “L2” could map to a single L1 host which will have that object. It could map to a single L2 host in a dedicated fleet. It could map to another POP or a shared regional cache. Expect another 50-80% hit rate at this L2.
If your L2 doesnt have the object now you need to go back to the origin. You want to do connection pooling here to save the long distance tcp setup. Use more consistent hashing as necessary. Insert object in the L2 cache.
There are optimizations to make at every stage. Check on probabilistic structures, network encapsulation, cumulative distributions, and network mapping/latencies.
On the other side you can extend the “depth” of your cache hierarchy with additional layers, like L3 etc. Think of doing things like having regional caches that your local/edge caches read from. Or store very large media objects in fewer locations with more density. Or centralize all origin requests through 1 or 2 specific locations to minimize the number of requests to the origin. I seem to recall that Akamai made a TON of money on this “net storage” layer historically.
Also forgot to mention previously TCP fast open and precomputed/cached TCP session values like cwnd. Youll want those to avoid slow start & bandwidth probing. Again, lots of optimizations at every step.
To answer your questions:
The load is distributed by custom, special purpose load balancers. They route each request to the correct server based on information in the request. The reason requests are routed to a particular server instead of a random one is because we want requests for the same content to go to the same set of servers so that we don’t have to cache more copies of content than is necessary, allowing more total content to be cached in a pop. The server that actually serves the content will then use Direct Server Return to bypass the load balancer when returning the content to the client. Since the content returned is a lot larger than the request for the content, the load balancers are able to handle a lot more requests than if all the connection flow had to go through the load balancer.
Cached content can be distributed to each pop in a few ways. The simplest is just for each pop to individually request the content from origin. This means the origin will see one request for each piece of content per POP. The other main way is by having one POP be the gateway. Other pops use that gateway as their origin, so the content is cached once at the gateway and then served to all the other pops. This way, the customer origin only gets a single request per piece of content from the gateway pop.
TLS termination is a challenge because of what you say. The keys are encrypted for distribution, and then only decrypted and loaded into a servers memory if it gets requests for that domain name. It is complicated.
Managing the machines is a big part of the work a CDN does. I work on one of the teams that works on part of that management. Lots of custom software, asset management, config distribution, etc. I would write more but I need to get my daughter to school!
Edit: didn't even realize the F is Cloudflare has since been lowercased: https://blog.cloudflare.com/end-of-the-road-for-cloudflare-n....
In more general terms:
A distributed reverse proxy with sensible defaults for caching.
Even though it doesn’t sound like much, it’s pretty cool and useful.
One of the things in ops that can be hard and expensive is any kind of high availability. Say you’re restarting your origin server regularly after pushing new code.
Depending on your use case, it might just be enough to have exactly two well provisioned web servers (VM/dedicated) and maybe one DB server. Then put a CDN in front of it and voila, high availability reads.
The Cloudflare Workers Runtime∗, for instance, is built directly around V8; it does not use nginx or any other existing web server stack. Many new features of Cloudflare are in turn built on Workers, and much of the old stack build on nginx is gradually being migrated to Workers. https://workers.dev https://github.com/cloudflare/workerd
In another part of the stack, there is Pingora, another built-from-scratch web server focused on high-performance proxying and caching: https://blog.cloudflare.com/how-we-built-pingora-the-proxy-t...
Even when using nginx, Cloudflare has rewritten or added big chunks of code, such as implementing HTTP/3: https://github.com/cloudflare/quiche And of course there is a ton of business logic written in Lua on top of that nginx base.
Though arguably, Cloudflare's biggest piece of magic is the layer 3 network. It's so magical that people don't even think about it, it just works. Seamlessly balancing traffic across hundreds of locations without even varying IP addresses is, well, not easy.
I could go on... automatic SSL certificate provisioning? DDoS protection? etc. These aren't nginx features.
So while Cloudflare may have gotten started being more-or-less nginx-as-a-service I don't think you can really call it that anymore.
∗ I'm the tech lead for Cloudflare Workers.
https://wiki.opensourceecology.org/wiki/Coral_CDN
https://en.wikipedia.org/wiki/Coral_Content_Distribution_Net...
https://github.com/morganestes/coralcdn
Just a few years ago, you could append .nyud.net to any domain and fetch it through the Coral cache instantly. But it's apparently been quietly swept under the rug.
Caching is such a basic thing to do that I'm concerned that the current crop of CDNs will mostly be used for mass surveillance. I also worry that VPNs are used for similar purposes by spy agencies.
IMHO the static web should have been distributed from the start. It should have been https everywhere. We should have kept cookies instead of trying to wedge in security on the frontend with OAuth, which is like a leaky sieve in comparison. We should have had Subresource Integrity (SRI) and been able to load scripts and other sensitive files from these caches fearlessly.
That said CDNs are still very useful for a range of functionality such as geographic distribution of content, reducing load to origin servers, DDoS protection and many other applications.
I'm pretty sure all static hosts use a CDN for all their customers' files as well.
In the past devs would use a CDN delivered asset in the hope that it was already in the user's cache from a visit to a different website that uses the same CDN, meaning it wouldn't need to be downloaded again. That meant you got a little speed boost.
Today browsers partition their caches so you can be absolutely sure the user doesn't have any of your website's assets in their cache if they've not visited before. This actually makes CDNs more important rather than less, because you want to deliver everything from the closest possible location to the user. That means you need a CDN.
Essentially the focus of CDNs has shifted away from Content, and now it's all more about Network.
Note that the original post is a link to an article, not a request for help.