Long running jobs often use checkpoints which require high speed networking and storage, which I don't see an option for. Eg, I cam get EC2 instances with 100gbps networking.
Great job starting the service! But, I think you have ways to go before reaching Hyperscalers level reliability.