I don't think that will make much difference in a year.
I think there's a pretty good chance that we've reached the point of diminishing returns, for our specific use case.
There are still like a billion other (more difficult) use cases to be tackled, but I think "generating code" has gotten really good to the point where the other bottlenecks will prevent further exponential progress on this specific task.
Sophisticated chain of reasoning LLMs like ChatGPT have baked in some natural language operations and they make it so i can create at a higher level of the language expression stack. But I'm still formulating my own expression. There's no conceivable path I can see where an improved model is going to be able to do what I do. I think that is clear from my ChatGPT threads at least.
With deterministic workflows, type-safe languages and test suites, agentic loops pretty much “can’t fail”. They will continue until the types resolve, the tests pass, and the project requirements are deterministically met.
By that point it’s literally just a case of typing a prompt in to a text field, and waiting.