We have jobs that users initiate that use 80+GB of memory and a few dozen cores. We run only one pod per node because the next size up EC2 costs a fortune and performance tops out on our current size.
These jobs are triggered via a button click that trigger a lambda that submits a job to the cluster. If it is a fresh node, user has to wait for the 1gb container to download from ECR. But it is the same container that the automated jobs that kick off every few minutes also uses, to rarely is there any waiting. But sometimes there is.
Should we be running some sort of clustering job scheduler that gets the job request and distributes work amongst long running pods in the cluster? My fear is that we just creat another layer of complexity and still end up waiting for the EC2, waiting for the pod to download, waiting for the agent now running on this pod to join the work distribution cluster.
However, we probably could be more proactive with this because we could spin up an extra pod+EC2 when the work cluster is 1:1 job:ec2.
Thoughts?
We're in the process of moving to Karpenter, so all this may be solved for us very soon with some clever configuration.