Yes, big time, but there continues to be lots of progress.
Most importantly, models are maturing, and this means less custom optimization is required.
Most importantly, models are maturing, and this means less custom optimization is required.
I suspect we will (or already are?) at a point where 95%+ of GPUs are used for inference, not training.