In my experience modern models are better at all tasks than models from two years ago, especially complex multi-step tasks.
It wasn't "better" it was better at kissing your ass which matches what a lot of people want in a partner.
I suspect GPT 5.6 would be even better at it, if given the same sycophantic system prompt and lack of guardrails.
If you want creative writings, use the API and play with the sliders.
For example, every day people teach teenagers how to drive and with only dozens of hours of practice, they are on the road.
Another interesting task would be to take the AI in a robot body into a vegetable garden and teach it to pull weeds. This is another task that lots of children help out with.
I think most are actually worth, as agentic harnesses seem to optimize for solving poorly described problems rather than following complex procedures as written. In other words, instruction following maximizing models seem to make worse free-form agents, but they're really all that some domains need.
You can do many more things, when stuff is cheaper, even if the stuff were otherwise unchanged.
What? GPT-4.1 was not a small model! And why wouldn't you use reasoning?
You're of course going to see poor results when you restrict yourself to small non-reasoning models, but why would you?
In voice interactions, ttfat is actually relatively important. If you look at models with a <1s ttfat you eliminate almost every reasoning model, less some of the diffusion models and more obscure ddtree/dflash like speculative decoding implementations.
GPT-Live, which is coming to the API soon, responds instantly while reasoning in the background. So it can say "Hold on, I'll look that up for you" and continue to respond to the user conversationally while running an asynchronous reasoning task in the background.
It's not in the API yet, but it should be in the coming weeks. You'll see an enormous improvement compared to GPT-4.1.
That’s not a credible position, but there isn’t anything that I or anyone else can say to someone who simply doesn’t want to believe something.
I don't know if I agree with that but it doesn't seem like an irrational claim and does seem credible to me.
so far there is no end to this progress in sight so it's full steam ahead on this singular domain. once it plateaus you should expect to see the greatest disruptions in human endeavors ever as all the training flops will start flowing to other domains to disrupt and dominate.