Amazon's Trainium2 AI Accelerator Features 96 GB of HBM, 4x Training Performance
anandtech.com
anandtech.com
We don't know what the tflops will be for the 300x yet, but it seems like the limitations here are more about memory than it is about tflops.
I'm genuinely curious because it seems like every other AI Accelerator/GPU thread on HN mentions MI300X, yet I haven't talked to/met anyone who trained anything with AMD.
As for training, it has been happening a bit on MI250's [1]. There is also A LOT that happens that doesn't make the larger news [2] and [3] and [4].
[0] https://techcommunity.microsoft.com/t5/azure-high-performanc...
[1] https://www.databricks.com/blog/training-llms-scale-amd-mi25...
[2] https://www.amd.com/en/resources/case-studies/kt-cloud.html
[3] https://www.lamini.ai/blog/lamini-llm-finetuning-on-amd-rocm...
With things like inferentia, although the performance numbers they have reported are really impressive-- in practice, model compilation is super finnicky. I'd love to see how seamless they can make usage of trainium.
Fact is that AMD was caught out of the initial AI race and is catching up. Lisa has said over and over again that she is committed to it.
X is pure gpu, A is more like an APU (cpu+gpu), C is pure cpu.
At $30/hour I’m literally spending a dollar or two just installing dependencies.
Whether AWS custom hardware is reasonably priced, and how much of a pain in the butt it will be to port my code, and how training speed compares to Nvidia GPUs will be my main deciding factors. I’m not locked in to AWS so Trainium will need to be faster and cheaper (cheaper per epoch) to be competitive for me.
Or whatever cloud provider you use does something funky and the setup on one machine is less portable to a GPU machine than you might expect.
If we were operating at 100x or 1000x the scale it might make sense to spend the extra engineering time, but we just pay for the NVIDIA chips (especially because a big chunk of our customers run at the edge where NVIDIA compatibility is even more important so we'd be doing double the engineering on an ongoing basis).
Overall the challenge is often more with making your code work on their accelerator, not porting it away from their accelerators so I wouldn't be too concerned about lock-in.