I spent a couple years in CDN & DNS land. Every meaningful CDN I know of is bespoke, but they all take very similar approaches to common problems.
There are good publications and papers from fastly and facebook circa 2012 to 2016 at events like usenix. Akamai had some seminal works in the early to mid 2000s. Oh, and google maglev and I think cloudlfare talk about IP fragmentation and similar problems where you need persistance but dont get a 5 tuple.
In short:
DNS and edns client subnet tell the CDNs name servers the clients network location. Map the client subnet and network topology to a nearby POP. Return those addresses in the A record response.
If you dont get edns0 use a weighted result of the resolvers clients.
The A record IPs get the client to the best available POP considering load, latency, bandwidth, cost, customer domains, etc.
At the pop you play network games to get the tcp session to the first layer of CDN hosts. Ill call them “layer 1” or L1. You almost definitely DO NOT run a full proxy with traffic 100% in:out at this layer. Its simply too wasteful. Instead you can use some combination of ARP, ECMP, or “layer 3 switching” to get tcp flows distributed among the available L1 hosts. These hosts will almost always share those initial “virtual” IP addresses amongst them.
The L1 hosts will terminate TCP/TLS and probably parse the HTTP request to determine the customer domain, customer cache rules, etc. The L1 howts will have a local “hot” cache of popular objects, lots of duplication between howts to distribute load. Probably a 50-80% hit rate here to immediately return the result to the client.
If the L1 is a miss they will use a consistent hash function to map the object (think URI plus any customer rules) to an L2 host. The L2 is a segmented cache, maximizing unique bytes stored for content in the “middle” of the popularity distribution. Implementation varies here; “L2” could map to a single L1 host which will have that object. It could map to a single L2 host in a dedicated fleet. It could map to another POP or a shared regional cache. Expect another 50-80% hit rate at this L2.
If your L2 doesnt have the object now you need to go back to the origin. You want to do connection pooling here to save the long distance tcp setup. Use more consistent hashing as necessary. Insert object in the L2 cache.
There are optimizations to make at every stage. Check on probabilistic structures, network encapsulation, cumulative distributions, and network mapping/latencies.