Cloudflare Gen 12 Server: Bigger, Better, Cooler in a 2U1N Form Factor
blog.cloudflare.com
blog.cloudflare.com
I knew fan power usage could be a substantial amount of a server's power draw, but I didn't realize they would need 320W to move enough air to cool 500W of dissapation.
10-20 years ago, 1U and blades made a lot of sense because space was a major constraint; now 2U or larger can make sense because 1U has a lot of tradeoffs for size, and power is the bigger constraint.
If space density is still needed, 2 nodes in 2u or similar narrow but tall configurations provide a more square profile that allows generating airflow with larger fans which is more efficient than the 40mm fans needed in 1U systems.
Oxide's rack uses 80mm fans that run at less than half the stock minimum speed to reduce this waste and run super quiet: https://oxide.computer/blog/the-cloud-computer
I recently picked up an R730xd. By default, a non-Dell-branded PCI-E device will force the fans to ~16000RPM, but there's an IPMI command to disable that and go back to regular, temperature-driven, fan control. Total server power came down by 70W after issuing that command.
Where the heck are you seeing 5-10kva?
One that I found nice too, Cloudflare is able to quickly install GPU's everywhere, because for over 6 years, they didn't utilize 1/6th of the slots for that purpose.
https://www.fool.com/investing/2023/11/22/cloudflare-started...
Would have loved to see that information from their blog, instead of their earning call :p
So, these newer CPUs while being a lot more powerful also draw a lot more power as well. In the end the squeezing 1000s of CFM through a 1.75" aperture wasn't worth the headache.
I'm glad to see cloud providers are still using datacenter standard 1U chassis. I assumed everyone had moved to Open Compute rack dimensions 537x44mm.
The neighboring cabinet to ours has a 3x Walmart box fans. I thought that was pretty innovative. Just use big silent, slow, fans instead of these tiny, noisy airhorns.
I thought fans consumed like 10 times less
https://www.delta-fan.com/newspost/ingress-protection-ip55-i...
40mm, 12.6W gets 21.56cfm (at zero pressure) or 3.416 inH2O (at zero flow). That's 1.7 cfm/W.
140mm, 46.8W gets 282.31cfm or 2.033 inH2O. That's 6.03 cfm/W, over 3x better.
Now, I'm being a bit unfair: I ignored pressure. I don't see fan curves, but a vaguely credible figure of merit if we're just trying to estimate efficiency is to multiple the flow and the pressure and divide by power. The little fan gets 0.68 and the big fan gets 1.44. That's still quite a lot better. (It does not mean that fan exceeds 100% efficiency. It's not delivering 282.31cfm at 2.033 in H2O.)
But, interestingly, these fans produce almost the same velocity. If the moderately lower available pressure from the large fan isn't a problem (which it shouldn't be -- we're talking about a regime where there is too much TDP per rack, so it's beneficial to spread things out, thus reducing resistance to airflow), these fans really can substitute for each other in the sense that covering the same area with big fans produces about the same airflow as covering the area with small fans.
Oh, and the big fan is louder, but not louder than 10 small fans!
So I stand by my claim: if you are cooling the contents of your rack by filling it with metal tubes and blowing air through the tubes using fans, use big fans, which you can achieve by using taller units. And 2U isn't enough for that 140mm fan.
I just don't see a rack full of 1U being anything but a giant liability in terms of trying to cool it and the following reliability issues.
Edit: Ah, I figured it out. It means "nodes", or the number of hot-swap separate computers held within the rack-unit chassis.
Is the power and performance / core count comparable, so it's more momentum preserving to stay with x86_64?
It's nearly time for another refresh to our x86 infrastructure, with deployment planned for 2024.
Example: https://www.supermicro.com/en/solutions/liquid-cooling#solut...
It's pretty safe to assume the situation will be the same with next generation EPYC 5 CPUs with maybe 192 or 256 cores, higher power draw, but more efficiency.
Also I noticed the SSDs must be changed by open top lid. Curious about anti-intrusion design.
Do they choose this design because they (may?) sharing datacenter or even sharing rack?
What I dont understand is why they haven't done 2U2N or 1N with Dual Socket design.
I'm really surprised by this, I always thought typical datacenters were kept at low temperatures, like below 20C.
40mm fan can push ~5,000 sqmm of air
80mm fan can push ~20,000 sqmm of air
Doubling the fan size results in 4x more airflow.
Like, this article is one of the top results for 2U1N, plus some dell pages and random jank.
https://chat.openai.com/share/37774e90-a271-4f6c-a5d3-3937e9...
I was curious too because I hadn't heard '1N' in a long time so I did the above and you might like the answer it gave me too!
I'm pretty sure this is wrong. N means "node", so a 2U4N is a 2U chassis that houses 4 servers. I'm pretty ambivalent for this reason for using chatgpt for research in things I cannot easily verify. I will admit this is a scarily convincing hallucination.
I'm definitely trying to balance it out! Apologies to all for the mistake here!
What CPU did they end up selecting?
Cores/RAM/Disk … even GPU since they mention the importance.
TL:DR - newer CPUs TDP is so high, they needed to change from 1U to 2U and that changes economics at a macro level per rack (since you’re density now halves).