How do you train new models if your GPUs are being used for inference? I guess the training happens significantly less frequently?
Forgive my ignorance.
Forgive my ignorance.
That isn't because we aren't training that often - we are almost always training many new models. It is just that inference is so computationally expensive!
Not that anyone should think any aspect (training nor inference) is cheap.