(barring some breakthrough that reduces costs, which of course may happen, but for which recent model improvements are not strong evidence of)
If you meant 3.5 9B and you truly believe it's as good as 4o then I can only assume you have a very basic use case.
"Reasoning" and now "Agentic" AI systems are not some fundamental improvement on LLMs, they're just running roughly the same prior-gen LLMS, multiple times.
Hence the conclusion that LLM improvement has slowed down, if not stagnated entirely, and that we should not expect the improvements of switching to these "reasoning" systems to keep happening.
“ChatGPT came up with an idea which is original and clever. It is the sort of idea I would be very proud to come up with after a week or two of pondering, and it took ChatGPT less than an hour to find and prove”
I'm saying they're not an advancement in the tech in the way GPT 1 through 3 were. They're a different kind of improvement.
And as such the rate improvement cannot just be extrapolated into the future.
All interesting conceptual breakthroughs came after GPT3: RL and reasoning being the main ones.