Long running jobs often use checkpoints which require high speed networking and storage, which I don't see an option for. Eg, I cam get EC2 instances with 100gbps networking.
Great job starting the service! But, I think you have ways to go before reaching Hyperscalers level reliability.
I have servers hosted in a similar quality datacenter with literally 9 years of uptime. It would be 10 but the customer shut down his services...
By the way, unless something has changed, you can't get more than 10gbps of throughput of a single tcp stream in those 100gbps setups...
See various speeds and feeds depending on instance size and type, and whether inside or outside the VPC at e.g. here:
https://d1.awsstatic.com/events/reinvent/2019/REPEAT_2_Deep-...