The flip side of this is meta having a hack that keeps their GPUs busy so that the power draw is more stable during llm training (eg don't want a huge power drop when synchronizing batches)
I'm not au fait with network data centres though, how similar are they in terms of their demands?
I expect you're right that GPU data centers are a particularly extreme example
My current guess is that I heard it on a podcast (either a Dwarkesh interview or an episode of something else - maybe transistor radio? - featuring Dylan Patel).
I'll try to re listen to top candidates in the next two weeks (a little behind on current episodes because I'm near the end of an audiobook) and will try to ping back if I find it.
If too long has elapsed, update your profile so I can find out how to message you!