Sufficiently large enterprises could get their own GPUs to run the models though.
One engineer's salary to accelerate a team of twelve is so cheap you can't afford not to.
Open models on-prem is the future, not a single doubt in my mind.
The biggest problem we had with on-prem was maintenance as it took a lot of staff and time to ensure decent reliability.