Nvidia Hopper Architecture In-Depth
developer.nvidia.com
developer.nvidia.com
Am I right in thinking the GPU-to-GPU communication is just shuttling chunks of data around for sharing inputs/outputs of computations? Or is there some other coordination going on between the GPUs directly with regards to the actual computations each is running? (Or is that still being managed wholly by the CPUs they're attached to?)
I have a sneaking suspicion that somewhere in an NSA datacenter there will be racks upon racks of these things running transformers over every text message and voice call coming from certain regions...
This could be used for DRM or nefarious purposes, but that’s not the point.
Obviously this is no more secure than the underlying hardware. All the attacks against SGX, for example, will still apply unless they are fixed.
I wonder what's different this time.
A group at IBM has been claiming for years that you can train with as little as FP4: https://papers.nips.cc/paper/2020/file/13b919438259814cd5be8...
There's been some speculation that if the Ethereum merge puts a lot of 3000-series cards on the used market then Nvidia might delay releasing this generation of GPU.