[1] https://www.jonpeddie.com/news/amd-to-integrate-cdna-and-rdn.... [2] https://centml.ai/hidet/ [3] https://centml.ai/platform/
[1] https://www.jonpeddie.com/news/amd-to-integrate-cdna-and-rdn.... [2] https://centml.ai/hidet/ [3] https://centml.ai/platform/
I have 7 NVidia 4090s under my desk happily chugging along on week long training runs. I once managed to get a Radeon VII to run for six hours without shitting itself.
I have 6 Radeon Pro VII under my desk (in a single system BTW), and they run hard for weeks until I choose to reboot e.g. for Linux kernel updates.
I bought them "new old stock" for $300 apiece. So that's $1800 for all six.
(I release it will be significantly lower, just try to get as much of a comparison as is possible).
Moreover, for workloads limited by the memory bandwidth, a Radeon Pro VII and a RTX 4090 will have about the same speed, regardless what kind of computations are performed. It is said that speed limitation by memory bandwidth happens frequently for ML/AI inferencing.
Because the previous poster had mentioned only single precision, where RTX 4090 is better, I had to complete the data with double precision, where RTX 4090 is worse, and memory bandwidth where RTX 4090 is the same, otherwise people may believe that progress in GPUs over 5 years has been much greater than it really is.
Moreover, memory bandwidth is very relevant for inference, much more relevant than FP32 throughput.
Titan V: 7.8 TFLOPs
AMD Radeon Pro VII: 6.5 TFLOPs
AMD Radeon VII: 3.52 TFLOPs
4090: 1.3 TFLOPs
The two are rather different and one market is worth trillions, the other isn't.
Beyond that, not much else.
There's a bunch of similar setups and there are a couple of dozen people that have done something similar on /r/localllama.