Thanks! AWS was costly compared to the free HPC cluster but it was sub $1000 per month and our bottlenecks were person-hours rather than money. We had a lot of work to do. As some have mentioned there wasn't a strong technical argument for switching to AWS, it was more about productivity with those tools. I'd built up a bit of credibility with the team at that point and was like "I can make this better" and my boss was said "ok cool". As I mentioned in the article the team was very open to try new things.
I can see an alternate timeline where I set up Buildkite in AWS to run Slurm jobs on the HPC cluster, rather than using EC2 spot instances. Using Buildkite + S3 were probably the more important infra changes.