The devil is in the details. Training large LLMs requires a lot of custom infra (handling GPUs going down, efficiently pushing data to keep the accelators busy, deciding on which mechanism of parallelizing model training is better - data vs model parallelism or both, tuning hyperparams of optimizers which can be different for larger batch sizes, etc)
Mosaic is one of the better providers for this. AWS is nowhere near ready at this current point in time, it is pretty much a "dumb" infra provider in large LLM training at this point. (Of course they won't be standing still and will prob acquire that capability one way or another)