But we do live in this era, no? What is cloud compute other than a time share?
In fact, we're already further along than that in terms of local AI. I'm currently able to get usable results at 8-10 tokens/sec using open-weight models on my laptop's integrated GPU, running on battery power. A $4,000 DGX Spark (less than what an IBM PC cost at launch in inflation-adjusted dollars) can get 3-5 times the inferencing performance with models 3-5x larger.