Training LLMs with AMD MI250 GPUs and MosaicML
mosaicml.com
mosaicml.com
AMD put two individual devices together, so you can tell people the combined memory is more than Nvidia has to offer.
> Overall we find that AMD MI250 achieves an average of ~80% of the per-GPU training throughput of A100-40GB and ~73% of A100-80GB
So their two-device bundle with more memory and theoretical compute just can't compete with Nvidia's prior gen single device.
Can MI300 compete with H100?