SLURM does that wrapping for you, where you essentially just point to the file that you want to run, along with some high level GPU and CPU resource allocation tags, and it just schedules and runs it for you.
I have seen some people trying to run GCP (lol) with SLURM, and wouldn't be surprised if it is possible with AWS/Lambda or any of the other cluster service providers (Cluster-as-a-service, CLaaS?).
Just through one Google search, looks like its definitely possible with AWS: https://docs.aws.amazon.com/parallelcluster/latest/ug/slurm-...
Or maybe you're kind of wanting a serverless GPU cloud - check out Runpod, Modal, Baseten, and Replicate.
Links:
https://slurm.schedmd.com/documentation.html
https://www.run.ai/ml-workflow-management
https://github.com/skypilot-org/skypilot
https://www.runpod.io/serverless-gpu
Ray is another good candidate as well (and feels more modern imho).
EDIT: so the main thing is technologies like Ray have a way to do these things, but I honestly just want an easy way to do this. Maybe means I will have to set up something with Ray and AWS myself and make a wrapper for that?
https://docs.ray.io/en/latest/ray-core/key-concepts.html#tas...
https://docs.ray.io/en/latest/ray-core/examples/gentle_walkt...
That's what I would want something to do if I was building tooling like this myself. Although, I'd do it in golang instead of python so that the dependency chain was simpler. A single small binary is nicer imho.