Many applications dont want to host inference on the cloud and would ideally run things locally. Hardware constraints is clearly important.
Id actually say its the most important metric for most open models now, since the price per performance of closed cloud models is so competitive with open cloud models, so edge inference that is competitive is a clear value add