At least in the past, nvidia’s consumer grade GPUs would fail eventually if you used them for ML. Not always but I’ve seen it happen multiple times.
Namely, there are 2 big issues: - Virtualization: Getting vGPU functionality out of them (most of the time, there's no point in having a full 1080 Ti for only 1 job, we want to have it split on more than 1 VM). - Kubernetes workloads: I'm running a simple workload that uses the GPU (for testing for instance) and k8s wants it all for only one pod.
You'll need to make sure to not overload gpu mem, otherwise you'll get oom errors. You can use k8s patch functionality to add custom resources, essentially use that to represent gpu mem and include the amount of gpu the pod uses in its definition, that way you don't go over the limit.