I have not found this to be true at all in my field (natural language generation).
We have a 7 figure GPU setup that is running 24/7 at 100% utilization just to handle inference.
We have a 7 figure GPU setup that is running 24/7 at 100% utilization just to handle inference.
Forgive my ignorance.
That isn't because we aren't training that often - we are almost always training many new models. It is just that inference is so computationally expensive!
Not that anyone should think any aspect (training nor inference) is cheap.