I've been spending all my time recently thinking about LLM costs. From that, I am currently of the mind that the most interesting question is if they can get it to work in the first place. After that, like with "performance" in the past, I believe there are ways to optimize things. We see the code harnesses doing this out of necessity, for example.
Some of the tools we use: https://www.induction.ai/docs/context-management