I don't think this follows at all. Just like benchmarks get saturated, lots of tasks get saturated as well. Over time, you can accomplish a given task for much cheaper, and part of that is due to open weight models. That doesn't imply that they're competitive with frontier models for the most advanced tasks, which might represent a smaller fraction of overall work, and thus use a smaller portion of tokens.
That said, at the moment I'm finding that not much can compete with GPT-6 Luna on cost / performance (not using for coding, but for AI pipelines in my product).