maybe we'll look back at token context windows like we look back at how much ram we have in a system.
If we have cheap billion token context windows, 99% of your use cases aren't going to hit anywhere close to that limit and as a result, your models will "just run"