Cloudflare servers don't own IPs anymore so how do they connect to the internet?
blog.cloudflare.com
blog.cloudflare.com
It’s very interesting that they are essentially treating IP addresses as “data”. Once looking at the problem from a distributed system lens, the solution here can be mapped to distributed systems almost perfectly.
- Replicating a piece of data on every host in the fleet is expensive, but fast and reliable. The compromise is usually to keep one replica in a region; same as how they share a single /32 IP address in a region.
- “sending datagram to IP X” is no different than “fetching data X from a distributed system”. This is essentially the underlying philosophy of the soft-unicast. Just like data lives in a distributed system/cloud, you no longer know where is an IP address located.
It’s ingenious.
They said they don’t like stateful NAT, which is understandable. But the load balancer has to be stateful still to perform the routing correctly. It would be an interesting follow up blog post talking about how they coordinate port/data movements (moving a port from server A to server B), as it’s state management (not very different from moving data in a distributed system again).
The cost they are working around is the cost of IPv4 addresses, versus the combinatorial explosion in their allocation scheme (they need number of services * number of regions * whatever dimension they add next, because IP addresses are nothing like data).
I am not sure where you see data replication in this scheme?
Overall, it seems like they are treating ip addresses as data essentially, which becomes most obvious when they talk about soft-unicast.
Anyway, I just found it interesting to look at this through this lens.
Spot on!
In past:
* /24 per datacenter (BGP), /32 per server (local network) (all 64K ports)
New:
* /24 per continent (group of colos), /32 per colo, port-slice per server
This is totally hierarchical. All we did is build a tech to change the "assignment granularity". Now with this tech we can do... anything we want. We're not tied to BGP, or IP's belonging to servers, or adjacent IP's needing to be nearby.
The cost is the memory cost of global topology. We don't want a global shared-state NAT (each 2 or 4-tuple being replicated globally on all servers). We don't want zero-state (a machine knowing nothing about routing, just BGP does the job). We want to select a reasonable mix. Right now it's /32 per datacenter.... but we can change it if we want and be more, or less specific than that.
Incidentally, this is exactly how GCP Cloud NAT works.
I had to solve this exact problem a year ago when attempting to build an anycast forward proxy, quickly came to the conclusion that it'd be impossible without a massive infrastructure presence. Ironically I was using CF connections to debug how they might go about this problem, when I realized they were just using local unicast routes for egress traffic I stopped digging any deeper.
Maintaining a routing table in unimog to forward lopsided egress connections to the correct DC is brilliant and shows what is possible when you have a global network to play with, however I wonder if this opens up an attack vector where previously distributed connections are now being forwarded & centralized at a single DC, especially if they are all destined for the same port slice...
Slightly OT question, but why wouldn't this be a problem with ingress, too?
E.g. suppose I want to send a request to https://1.2.3.4. What I don't know is that 1.2.3.4 is an anycast address.
So my client sends a SYN packet to 1.2.3.4:443 to open the connection. The packet is routed to data center #1. The data center duly replies with a SYN/ACK packet, which my client answers with an ACK packet.
However, due to some bad luck, the ACK packet is routed to data center #2 which is also a destination for the anycast address.
Of course, data center #2 doesn't know anything about my connection, so it just drops the ACK or replies with a RST. In the best case, I can eventually resend my ACK and reach the right data center (with multi-second delay), in the worst case, the connection setup will fail.
Why does this not happen on ingress, but is a problem for egress?
Even if the handshake uses SYN cookies and got through on data center #2, what would keep subsequent packets that I send on that connection from being routed to random data centers that don't know anything about the connection?
But although I would be surprised if load were not also part of the route picker, I would also be surprised if the routers didn't have some association or state tracking to actively ensure related packets get the same route.
But I guess this is saying exactly that, that it's relying on luck and happenstance.
It may be doing the job well enough that not enough people complain, but I wouldn't be proud of it myself.
An attacker could choose to only compromise devices located near a particular data center, but that would really reduce the amount of traffic they could generate, and also other data centers would stay online and serve requests from users in other places.
Most routers with multiple viable paths pass was too much traffic to do state tracking of individual flows. Most typically, the default metric is BGP path length, for a given prefix, send packets through the route that has the most specific prefix, if there's a tie, use the route that transits the fewest networks to get there, if there's still a tie, use the route that has been up the longest (which maybe counts as state tracking). Routing like this doesn't take into account any sort of load metric, although people managing the routers might do traffic engineering to try to avoid overloaded routes (but it's difficult to see what's overloaded a few hops beyond your own router).
For the most part, an anycast operation is going to work best if all sites can handle all the forseable load, because it's easy to move all the traffic, but it's not easy to only move some. Everything you can do to try to move some traffic is likely to either not be effective or move too much.
* A new (usually closer) DC comes online. That will probably be your destination from now on.
* The prior DC (or a critical link on the path to it) goes down.
The crucial thing is that the client will typically be routed to the closest destination to it. In the egress case the current DC may not be the closest DC to the server it is trying to reach so the return traffic would go to the wrong place. This system of identifying a server with unique IP/port(s) means that CF's network can forward the return traffic to the correct place.
- See: https://news.ycombinator.com/item?id=10636547
- And: https://news.ycombinator.com/item?id=17904663
Besides, SCTP / QUIC aware load balancers (or proxies) are detached from IPs and should continue to hum along just fine regardless of which server IP the packet ends up at.
Also, there are a fair number of ASes that attempt to load balance traffic between multiple peering points, without hashing (or only using the src/dst address and not the port). This will also cause the problem you described.
In practice it’s possible to handle this by keeping track of where the connections for an IP address typically ingress and sending packets there instead of handling them locally. Again, since it’s a few ASes that cause problems for typical connections, is also possible to figure out which IP prefixes experience the most instability and only turn on this overlay for them.
I guess ingress is next, then? Two layers of Unimog to achieve stability before TCP/TLS termination maybe.
Originally IP was a way to allow discrete physical computers in different locations owned by different organizations to find each other and exchange information autonomously.
These days most compute actually doesn't look like that. All my compute is in AWS. Rather than being autonomous it is controlled by a single global control plane and uniquely identified within that control plane.
So when I want my services to connect to each-other within AWS why am I still dealing with these complex routing algorithms and obtuse numbering schemes?
AWS knows exactly which physical hosts my processes are running on and could at a control plane level connect them directly. And I, as someone running a business, could focus on the higher level problem of 'service X is allowed to connect to service Y' rather than figuring out how to send IP packets across subnets/TGWs and where to configure which ports in NACLs and security groups to allow the connection.
Similarly my ISP knows exactly where Amazon and CloudFlare's nearest front doors are so instead of 15 hops and DNS resolutions my laptop could just make a request to Service X on AWS. My ISP could drop the message in AWS' nearest front door and AWS could figure out how to drop the message on the right host however they want to.
I know there's a lot of legacy cruft and also that there are benefits of the autonomous/decentralized model vs central control for the internet as a whole but given the centralized reality we're in, especially within the enterprise, I think it's worth reevaluating how we approach networking and whether the continuing focus on IP is the best use of of our time.
By default, sure. You can easily bring your own IPs into AWS and use them instead, and I don't think it's hard to imagine the pertinent use cases and risk management this brings.
I was looking for the "just" that handwaves away the complexity and I was not disappointed.
How do you imagine your laptop expressing a request in a way that it makes it through to the right machine? Doing a traceroute to amazon.com, I count 26 devices between me and it. How will those devices know which physical connection to pass the request over? Remember that some of them will be handling absurd amounts of traffic, so your scheme will need to work with custom silicon for routing as well as doing ok on the $40 Linksys home unit. What are you imagining that would be so much more efficient that it's worth the enormous switching costs?
I also have questions about your notion of "centralization". Are you saying that Google, Microsoft, and other cloud vendors should just... give up and hand their business to AWS? Is that also true for anybody who does hosting, including me running a server at home? If so, I invite you to read up on the history of antitrust law, as there are good reasons to avoid a small number of people having total control over key economic sectors.
That's my whole point. You're thinking of it from an IP perspective where there are individual devices in some chain and they all need to autonomously figure out a path from my laptop to AWS. The reality is every device between me and AWS is owned by my ISP. They know exactly which physical path ahead of time will get a message from my laptop to AWS. So why waste all the time on the IP abstraction?
> I also have questions about your notion of "centralization". Are you saying that Google, Microsoft, and other cloud vendors should just... give up and hand their business to AWS?
AWS is just an example. Realistically a huge amount of traffic on the internet is going to 6 places and my ISP already has direct physical connections to those places. Maintaining this complex and byzantine abstraction to figure out how to get a message from my laptop to compute in those companies' infrastructure should not be necessary.
And in general the more important part is within AWS' (or Microsoft's or enterprise X's) network why waste time on IP when the network owner knows exactly which host every compute process is running on?
Instead of thinking of an enterprise network as a set of autonomous hosts that need to figure out a path between each other think of it as a set of processes running on the same OS (the virtual infrastructure). Linux doesn't need to do BGP to figure out how to connect two processes so why does your network?
Neither of these are true in general. And suppose AWS (or GCP, or Azure, or Cloudflare...) decides to add a new POP. How do they broadcast to your ISP and all the other ISPs in the world how exactly to send datagrams to it?
IP is concrete, not abstract. Whatever the form the network takes, when you make a request, your ISP is going to have a make a decision on how to route it over their physical assets to get it to the desired destination. Unless you are talking about your ISP provisioning a physical circuit directly between you and Amazon, with no multiplexing and no equipment on it, you are going to have those hops whether you use IP to choose the route or not. That is not really negotiable, or you're not describing a network (or something even remotely viable) at all. Maybe that path is invisible to you, but it exists.
And in fact, in many or even most carrier networks, this is abstracted in much the way that you describe within that particular network using MPLS. But this approach doesn't scale to the scope of the Internet, requires all edge devices have complete knowledge of every necessary path in the network, and makes inter-networking more difficult because every endpoint and its end-to-end path to every other needs to be shared and synchronized. This is actually more complex, and much more brittle, than the current implementation. And for what? I still have yet to understand what advantage you think would be gained here. Right now, if you want to talk to s3, you send a packet to s3, and your ISP does all the 'complex and byzantine' work. What do you care how they do it?
Ignoring the long tail is also silly. FAANG might represent a majority of traffic on the Internet, but the long tail is huge, and you can't just hand-wave it away like that. Enabling it is what makes the Internet what it is, and if your proposal doesn't account for it, it's dead in the water.
> And in general the more important part is within AWS' (or Microsoft's or enterprise X's) network why waste time on IP when the network owner knows exactly which host every compute process is running on?
Knowing that is the easy part. You still need to figure out a path and actually route the packets along it. You still need to deal with path selection, load balancing, fault tolerance, and synchronizing any necessary state. You still need devices along the path to know what to do with a packet they receive, somehow. It turns out that hop-by-hop routing is an efficient and viable way to accomplish this.
> Instead of thinking of an enterprise network as a set of autonomous hosts that need to figure out a path between each other think of it as a set of processes running on the same OS (the virtual infrastructure). Linux doesn't need to do BGP to figure out how to connect two processes so why does your network?
Because the 'network' you describe is not a network? It's processes running on the same machine? This is not analogous at all to a large distributed network like the Internet.
The ISP can internally use MPLS to do routing exactly the way you suggest: build a "circuit" between you and Amazon and then route the packets internally in their networks through this circuit, instead of using IP. This works because the ISP has a global view of their network and as such their routers don't need to work independently. MPLS was needed in a time where IP routing was too slow, but nowadays you can get full speed without MPLS.
But anyway, it doesn't matter how packets are routed internally within each network, IP is still required for routing between different networks. Which is actually super common!
> You're thinking of it from an IP perspective where there are individual devices in some chain
What do you plan to replace the individual devices with?
Looking at the traceroute in question, something I'd suggest you do for your own, I see one router owned by me, 6 by my ISP, 2 at an interchange, 5 for a backbone, a number of intermediate ones of mysterious ownership, and finally one owned by Amazon, presumably with a bunch of other Amazon hops that are hidden from me.
These are physical devices connected by physical links (wires, cables, fiber). Wires that people installed. Wires that break. Connected to devices that break. What in your proposed grand vision will happen there? Please start with the happy path and then give detail on the failure cases.
A similar problem applies in Amazon. The abstraction they hand you is pretty convenient. But that abstraction is made out of many millions of devices connected up in cunning ways.
Linux can connect two processes because the kernel has total control over a modest number of CPUs and RAM. That just doesn't compare to literal billions of interconnected devices with no central control. It's like saying that because your dad knows your mom's name and how to reach her, he should be able to do the same thing for everybody in his country.
At least from a security perspective though ip acl’s are falling out of favor to service based identities, which is a good thing.
You can see how AWS internally does networking here: https://m.youtube.com/watch?v=ii5XWpcYYnI
https://aws.amazon.com/blogs/aws/new-amazon-ecs-service-conn...
> AWS knows exactly which physical hosts my processes are running on and could at a control plane level connect them directly. And I, as someone running a business, could focus on the higher level problem of 'service X is allowed to connect to service Y' rather than figuring out how to send IP packets across subnets/TGWs and where to configure which ports in NACLs and security groups to allow the connection.
You shouldn't be? Doesn't AWS number your machines for you automatically and give you a unique ID you can use with DNS to reach it? And also provide a variety of 'ingress' services to abstract load balancing and security as well? I'm not a consumer of AWS services in my dayjob, but isn't this their entire raison d'etre? Otherwise you may as well just run much cheaper VMs elsewhere.
> Similarly my ISP knows exactly where Amazon and CloudFlare's nearest front doors are so instead of 15 hops and DNS resolutions my laptop could just make a request to Service X on AWS. My ISP could drop the message in AWS' nearest front door and AWS could figure out how to drop the message on the right host however they want to.
Uhm, aside from handwaving away how your ISP is going to give you a direct, no-hops connection to AWS, this is pretty much exactly what your ISP is doing. Hell, in some cases, your ISP has abstracted the underlying backbone hops too using something like MPLS, and this is completely invisible to you as an end user. You or your laptop don't have to think about the network part of things at all. You ask to connect to s3, your laptop looks up the service's IP address (unique ID) in DNS, sends some packets, and your ISP routes them to CloudFlare's nearest front doors.
There are some good arguments to be made for a message-passing focused rather than connection focused protocol model, but that doesn't seem to be what you're talking about. What you seem to be talking about is doing away with routing altogether, and even in a relatively centralized internet, that just makes zero sense. We will continue to need the aggregation layer, we will continue to have multiple routes to a resource through multiple hops that need to be resolved into a path, and we'll continue to need a way to uniquely identify a service endpoint.
TCP/IP is full of cruft and causes a shitload of unnecessary problems, but increasing the number of bits in the addressing space solves literally none of them.
It's useful that they use client side certificates. (They call this "authenticated origin pull", but it seems to be client side certs.
https://medium.com/@ss23/leveraging-cloudflares-authenticate...
Authenticated origin pulls should not be “useful”. They should be on and configured securely by default, and any insecure setting should get a loud warning.
Indeed, we do happy-eyeballs where we can and _strongly_ prefer IPv6 when origin host resolves AAAA. However, we still need IPv4, since there is still a big chunk of traffic that only works IPv4.
If the client gives us a chance, we'll strongly prefer IPv6.
or
2) We admit that v6 needs to be rethought and rethink it. I understand why v6 does not just increase IP address bits from 32 to 128, but at this point I think everyone has admitted that v6 is simply too difficult for most IT departments to implement. In particular, the complexity of the new assignment schemes like prefix delegation and SLAAC needs to be paired back. Offer a minimum set of features and spin off everything else.
Reduction in scope and complexity makes it easier to implement high quality IPv6 stacks. We also need more high-quality reference implementations available for newer and underserved platforms, but that's a different subject.
SLAAC seems to introduce insane churn in the IPv6 of end user devices. SLAAC (vs DHCPv6) seems to struggle in fully configuring an end user device (think DNS servers etc.
As someone who has beat my head on the IPv6 thing for a bit before giving up (I tried to go all IPv6), what is the "proper" way to setup a home IPv6 network with SLAAC and not DHCPv6?
If DHCPv6 for some reason is essentially always required, why not let subsume SLAAC for IP address assignments?
RFC 8106 IPv6 Router Advertisement Options for DNS Configuration
However, it only handles your address. Apparently there's an extension to make it also provide DNS addresses and so on. If you don't have that, then I guess you configure DNS manually. Or use 8.8.8.8.
It's not like your router is doing some magic DNS auto-discovery, by the way. It just tells devices the addresses that someone typed into its own configuration page.
I'm really tired of hearing it will "Just work" - that's proven to be a lie over and over. But I would love to be shown where these major players do their static blocks for ipv6 (having fought this fight for a while).
Are you using comcast? I think they only DYNAMICALLY assign you a 64 - so you can only create ONE subnet on your entire network. Again illustrating the shortage and difficulty in using IPv6. I'd thought /48 would be minimum PD, but that is not the case. Or even a /60? No go. There are workarounds I'm aware of, but this stuff absolutely DOES NOT "just work".
[1] - https://support.google.com/fiber/answer/6136162?hl=en&ref_to...
> Are you using comcast? I think they only DYNAMICALLY assign you a 64...
In my experience here in San Francisco with both IPv4 and IPv6 service (and IPv4 service elsewhere in the US years and years ago) Comcast will give you a v4 IP (or v6 subnet) for as long as the same edge device (or an edge device with the same MAC address (or DUID for v6)) continues to renew the lease. So, yeah, it's dynamic in theory, but static in practice.
> Or even a /60 [via DHCPv6-PD]?
Weird. Configuring my DHCPv6 server to do Prefix Delegation and request a /60 always worked just fine for me. I was disappointed that I couldn't get a /56 or /48, but was okay with the /60 that they gave me. Again, this was in San Francisco, so maybe other parts of the country are managed _way_ differently. Maybe.
Alternatively, are you _sure_ that the DHCP-PD request was refused and that that you didn't -say- fail to configure your system to actually assign slices of the /60 to your LAN?
...
Actually, now that I'm thinking about it, I seem to recall a problem like you're describing.
If I'm not misremembering that I had this sort of problem, then, maybe, try configuring your edge device to ask for a /60, change the DUID that it will use to make the request, and then bounce the WAN interface (or reboot the device, or do whatever is required to apply the changes). My memory is absolute shit (and likes to hallucinate things that never happened), so this might not do anything useful. But it's usually a pretty easy thing to try.
By "insane churn" do you mean "devices generate and allocate new IP addresses for themselves periodically (maybe daily, maybe more frequently)"? If you do, then that's not SLAAC, that's the head-assed thing sometimes known as "IPv6 Privacy Addresses". From what I've seen on Windows, OSX, and Linux, this makes it so that there's one IP that remains constant, and a parade of addresses that get assigned as time marches on. You can disable it on Windows, OSX, and Linux, and I would recommend doing so.
> SLAAC (vs DHCPv6) seems to struggle in fully configuring an end user device (think DNS servers etc.
Yeah, if you're interested in only using SLAAC, then the best you can do is set the `RDNSS` option [0] in your Router Advertisements and pray that the network configurator in the OS you're using has bothered to pay attention to it.
[0] <https://www.rfc-editor.org/rfc/rfc8106#section-5.1> (Do note that despite the date on this RFC, this option was first specified in 2007, and first specified in a non-experimental RFC in 2010... so, it's not like it's new.)
I think privacy extensions are unavoidable - they default on in many places. So I'm leaving them. Some devices actually rotate more often (ie, when connecting to different wifi points even if underlying network is the same, apple seems to generate another new IP). But compared to ipv4 (where you can almost immediately trace from an IP you have in a log to device) -> you need more support in your tooling to do that with IPv6 and privacy extensions.
Honestly, given that the vast majority of the sites that use v6 "privacy addresses" are going to be end-users at their home, and that most of those folks are going to be either using web browsers, and/or already logged into the servers that are servicing their requests, there are so very, _very_ many powerful ways that folks can be tracked that have absolutely nothing to do with their IP address.
"Privacy addresses" are just a nuisance.
> Some devices actually rotate more often (ie, when connecting to different wifi points even if underlying network is the same, apple seems to generate another new IP).
I'm not sure _exactly_ the setup you're talking about. If "connecting to different wifi points" means "disconnecting from one SSID and connecting to another SSID but still being on the same physical network", then I think that this is OSX randomizing your MAC address and/or OSX generating a new DUID when connecting to a different SSID.
Prefix delegation to customers is necessary if you want to avoid NAT. IPv4 stayed with delegating just one IP address to customer (household) because everyone used NAT in home routers due to IP address conservation.
SLAAC was introduced because people wanted a simpler alternative to more complex DHCP. That is why it is mandatory, while DHCP is optional.
> but at this point I think everyone has admitted that v6 is simply too difficult for most IT departments to implement.
I do not think that IPv6 is too difficult to implement, i rarely heard this argument. The main reason that there is no transition is that there is no economic nor political incentive to do so for individual organizations, so there is a coordination problem.
I am absolutely no expert, but could get my head around ipv4, but IPv6 - I always end up running into a fuss. I really wish they'd expanded address space to 64 bits, a few other tweaks, and called it good. Maybe call it IPv5? Is there any chance of doing something like this.
So many things that are so trivial or well known in IPv4 are a total nightmare pain with IPv6. Some quick examples:
Internet service providers will happily give you a block of static IPv4 addresses for a price. ATT goes up to a 64 ip address block easily, even on residential. Almost impossible to get a static block of IPv6 in the US.
Let's say you are SMB, you want WAN failover. With IPv4 this is simple. You can either get two blocks of static for your upstream, and route them directly as appropriate to your servers, or go behind a NAT and do a failover option. Whent the failover happens, your internal network is relatively unaffected.
Now try to do this with IPv6? You can't get your static IP's to do direct routing with, and NPT and anything else is a mess, and the latency in having your entire network renumber when the WAN side flaps is stupid and annoying.
In many SMB contexts folks are very used to DHCP, they use it to trigger boot scripts TFTP, zero touch phone provisioning and lots more, pass out time servers and other info and more. The set of end user devices (printers, phones, security cameras, intercoms, industrial iOT) that can be configured and supported with IPv6 is so poor and the complexity is so high.
Not all ISP's offer prefix delegation to end user sites. Because you have an insane minimum subnet size with ipv6 a lot of things that for example you just need two IPs (think separate network for customer premise equipment) now need a 18,446,744,073,709,551,616 addresses.
The GSE debacle means instead of a very large 64 bit address space we got an insane 128 bit address space. Seriously, how about 96 or anything else a bit more reasonable.
Even things like ICMPv6 - if you just let it through the firewall you could be asking for trouble, but blocking it also causes IPv6 problems. Ugh. Oh, it's simpler than IPv4 they say.
This causes issues on IPv4 as well, the only difference is that a lot of the dirty hacks and workarounds were removed for IPv6 so that people are forced to deploy it properly.
"forced to deploy it properly" = giant headache. I'm tired of IPv6 folks saying it's a pain in the neck because it's "proper".
With ICMPv4 if you really needed / wanted, you could basically drop ICMPv4 at the firwall edge (with TCP MSS clamping etc). And the attack space with ICMPv4 coming through I don't think was TOO bad.
When folks say they don't need to filter ICMPv6 for things like RS / RA / NS / NA traffic that seems SO SO sketchy to me.
NAT is that IPv6.
How does the ingress "router" load-balance incoming connections, which it must (even if the "router" is a default host or cluster)? CF isn't opening TCP, then HTTP just to send redirects to another IP for the same cluster.
I guess hashing, on IP and port, is already readily used in routing decisions, so eyeball-addr and -port of the inbound packets, 4-tuple (CF-addr [fixed, shared], CF-port [443], eyeball-addr, eyeball-port), provides a consistent.
I guess that this is good for 99.9 percent of connections, which are short-lived, and opportunistically kept open or reused. I suppose other long-lived connections might take place in the context of an application that tracks data above and outside of TCP-alone. I'm grasping for a missing middle, in size of use case, and can't quickly name things that people might proxy but need stable connections. CloudFlare's reverse proxying to web servers would count, if the web-fronting had to traverse someone else's proxy layer.
What are the rough edges here? What's are next challenges here to build around?
40% is caused by archaic protocols that don't allow enough addresses or local efficient P2P or ISP caching of public content (sort of what IPFS aimed to do), which would alleviate much of the need for CDNs in the first place. The remaining 10% is just solving hard problems.
I'm a little surprised that splitting by port number gives servers enough connections; maybe they are connection-pooling some of them between their egress and final destination. If there was truly a 1:1 mapping of all user requests to TCP connections then I'd expect there to be ~thousands of simultaneous connections to, say, Walmart today (black Friday) which is also probably on an anycast address, limiting the number of unique src:dst ip and port tuples to 65K per cloudflare egress IP address. Maybe that ends up being not so bad and DCs can scale by adding new IPs? https://blog.cloudflare.com/how-to-stop-running-out-of-ephem... covers a lot of the details in solving that and similar problems.
Each data center has a single IP for each country code (so that they can make outgoing requests that are geolocated in any country). In order to achieve that, they have a /24 or larger range for each country, and announce it from all their data centers, and then they route the traffic over their backbone to the appropriate data center for that IP.
Then in the data center, they share the single IP across all their servers by giving each server a range of TCP/UDP port space (instead of doing stateful NAT).
Edit: based on this part of blog post:
“With a port slice of say 2,048 ports, we can share one IP among 31 servers. However, there is always a possibility of running out of ports. To address this, we've worked hard to be able to reuse the egress ports efficiently.”
I don't see this behavior in my setup. I have a server in AMS and connect from IST, but `netstat` reveals a unicast IP address that is routed back to IST.
https://en.wikipedia.org/wiki/Realm-Specific_IP
With port ranges rather than being 'leased', being allocated on the the basis of per server within a locale.
So the IP goes to the locale, the port range is the the static RSIP to the server within that locale.
* pcp https://datatracker.ietf.org/doc/html/rfc6887
* ds-lite port allocation mode https://support.huawei.com/enterprise/en/doc/EDOC1100125799/...
Neat technical solution though!
Deploying IPv6-mostly access networks : https://news.ycombinator.com/item?id=33694293
(I can't for the life of me find the comment chain where this link was posted ~4 days ago ??)
RFC 8925 – IPv6-Only Preferred Option for DHCPv4 (2020) : https://news.ycombinator.com/item?id=33697978 (posted on HN by me)
Perhaps they should go v6 internally and implement one v4-v6 gateway to put all these tricks behind.
If you think you can do better than that, I look forward to hearing your plan. Personally, I think that's huge progress.
https://blog.cloudflare.com/introducing-cloudflares-automati...
The success of that convinced us we should do something to improve the Internet every year to celebrate our “birthday.” Over time we ended up with more than one product that met that criteria and timing, so it went from a day of celebration to a week. That became our Birthday Week. Then we saw how well bundling a set of announcements into a week was so we decided to do it other times of the year. And that’s how Cloudflare Innovation Weeks got started, explicitly with us delivering IPv6 support back in 2011.
> To avoid geofencing issues, we need to choose specific egress addresses tagged with an appropriate country, depending on WARP user location. (...) Instead of having one or two egress IP addresses for each server, now we require dozens, and IPv4 addresses aren't cheap.
> Instead of assigning one /32 IPv4 address for each server, we devised a method of assigning a /32 IP per data center, and then sharing it among physical servers (...) splitting an egress IP across servers by a port range.
Indeed, the starting point is sharing IP's across servers with port-ranges.
But there is more:
* awesome performance allowed by anycast.
* ability to route /32 instead of /24 per datacenter.
Generally, with this tech we can have much better IP usage density, without sacrificing reliability or performance. You can call it "global anycast-based stateless NAT" but that often implies some magic router configuration, which we don't have.
Here's one example of problems we run into - the lack of connectx() syscall on Linux - makes it hard to actually select port range to originate connections from:
https://blog.cloudflare.com/how-to-stop-running-out-of-ephem...
Of course not every destination is an IPv6 host, so IPv4 remains necessary, but at least IPv6 can avoid the need for port slicing, since you can encode the same bucketing information in the IP address itself.
I've seen this idea used as a cool trick [0] to implement a SOCKS proxy that randomizes outbound IPv6 address to be within a publicly routed prefix for the host (commonly a /64).
I guess as long as you need to support IPv4, then port slicing is a requirement and IPv6 won't confer much benefit. (Maybe it could help alleviate port exhaustion if IPv6 addresses can use dynamic ports from any slice?)
Either way, thanks for the blog post, I enjoyed it!
Probably they didn't need to do much work with IPv6, since half of the post is solving IPv4 exhaustion problems.
Warp uses its own set of egress IPs and their geolocation is close to your real location.
Maybe this ends up reducing cost on customers though, because the international transit happens in your backbone network rather than on the internet (customer-side).
If anyone from CF sees this, I can work with you and give you data on this. I’m dealing with this at one of the large social media companies.
Here's an example, this is NSFW - https://atragcara.ga
The only way for somebody to DDoS from Cloudflare would be using workers, however, this isn't practical as workers have a very limited IP Range.
I think you're probably right about the spoofing but it comes off a little dismissive when the possibility of a site that queries other sites, could be tricked into doing something it shouldn't, is always going to be in the realm of a possibility.
There was a proposal (BCP38) which said that networks should not allow outbound packets with source IPs which could not originate from that network, but it didn't really get a lot of traction -- mainly due to BGP multihoming, I think.
Some tier-1s do follow BCP38 though, so one day maybe? Still, there's plenty of abuse to be done without spoofing, so while it would be an improvement, it wouldn't usher in an era of no abuse.
Apart from perhaps a couple of sites like gob.gq, there's essentially nothing of any value on those TLDs. Allow-list the handful of good sites, if you must, and default block the rest.
Are there though, really? Can you give some examples?
To a first approximation, I contend that essentially everything on Freenom is bad. There are maybe a handful of good sites (the one I listed, https://koulouba.ml/, etc) but you can find those on Google in a few minutes with some site: searches.
I commend your efforts in blocking the scam sites, but also honestly believe that it would be better for you, your customers and the internet at large to default block Freenom. Freenom sites are junk, wherever they are hosted.
Freenom TLDs are just junk. Save yourself the hassle and default block :-).
I understand that they started using anycast for the egress IPs as well, but thats unrelated to the NAT problem.
As far as I understand the history of the IP protocol, initially an IP address pointed to a host. (/etc/hosts file seems that way)
Then it was realized a single entity might have multiple network interfaces, and an IP started to point to a network card on a host. (a host can have many IP's). Then all the VRF, dummy devices, tuntaps, VETH and containers. I guess an IP is now pointing to a container or VM. But there is more. For performance you can (almost should!) have an unique IP address per NUMA node. Or even logical CPU.
In modern internet a server IP: points to a single CPU on a container in a VM on a host.
Then consider Anycast, like 1.1.1.1 or 8.8.8.8. An IP means something else... it means a resource.
On the "client" side we have customer NAT's. CG NAT's and VPN's. An IP means similarly little.
The IP's are really expensive, so in some cases there is a strong advantage to save them. Take a look at https://blog.cloudflare.com/addressing-agility/
"So, test we did. From a /20 address set, to a /24 and then, from June 2021, to an address set of one /32, and equivalently a /128 (Ao1). It doesn’t just work. It really works"
We're able to serve "all cloudflare" from /32.
There is this whole trend of getting denser and denser IP usage. It's not avoidable. It's not "breaking the Internet" in any way more than "NAT's are breaking the Internet". The network evolves, because it has to. And for one, I don't think this is inherently bad.
I agree. NATs, particularly the Carrier NAT that smartphone users are behind, has broken the internet. It's made it so most people do not have ports and cannot participate in the internet. So now software developers cannot write software that uses the internet (without depending on third parties). This is bad. So is what you've done.
Someday ipv6 will save us.
Ao1 has super nice censorship resistance properties. With DoH + ECH, that essentially is game over for most firewalls. Can't wait to see just how Cloudflare rolls Ao1 out (I'd imagine it'd be opt-in, like Green Compute).
> Then it was realized a single entity might have multiple network interfaces, and an IP started to point to a network card on a host
Much to Nagle's chagrin: https://news.ycombinator.com/item?id=21088736 (a Berkeleyism as he calls it)
Another perspective is that the connection of an IP to specific content or individuals was a bug of the Internet’s original design and thankfully we’re finally finding ways to disassociate them.
I can totally see an argument against their CDN being too pervasive and problematic for TOR users, but this seems fine IMO.
There's a fourth way to resolve this, that works for the core use case, is less engineering, and was in production 20 years ago, but I can't fit it within the margins of this comment box.
// CF's approach has additional feature advantages though.
> PING cloudflare.com (104.16.133.229) 56(84) bytes of data.
> 64 bytes from 104.16.133.229 (104.16.133.229): icmp_seq=1 ttl=52 time=10.6 ms
With a ping like this, you know that I am not using Musk's Internet....