I truly believe the cost per token for most people will eventually go approximately to zero even in the absence of subsidy unless there is a dramatic state intervention to curtail free computing and hardware R&D and manufacturing.
I truly believe the cost per token for most people will eventually go approximately to zero even in the absence of subsidy unless there is a dramatic state intervention to curtail free computing and hardware R&D and manufacturing.
Your premise of full token commoditization has a major wrinkle though. GPU-centric data centers with never allow tokens to go to zero since you have a 1.5 year technical deprecation and 5 year life on the GPUs. Maybe Cerebras or similar chip builders, after writing model weights to silicon, will improve those returns on capital and enabling the scaling you are envisioning.
From my perch, AI at the edge is severely underserved, and the GTM for the hyperscalers mostly ignore it. So if someone figures out how to scale AI at the edge (nvidia + HF perhaps) then I think your premise is on point.