Is there a world where we look back at how AI usage is charged today and we equate it with how we had minutes on AOL and how absurd it seems looking back?
However I think there is still a significant runway for these models to scale, so there will always be some sort of offering from providers. I can't imagine that our current use of the context window will be how that looks in a handful of years.
Maybe if we have some breakthrough in how to get equivalent ai capacity out of less compute it’ll seem crazy in retrospect.
But the magnitude of flops per token on SOTA models is mind boggling huge.
Any other outcome would be pretty terrible.