The incentives are rather in the opposite direction. Enthusiasts are trying to get models to run at all on the hardware they have, even if the resulting efficiency is poor, whereas a big company spending billions on hardware can afford to hire hundreds of performance engineers to tune their systems end-to-end for maximum efficiency, since the expense pays for itself even if they only manage to eke out a 1% improvement. I.e. why not make training and inference efficient when you can just throw money at the problem?