They basically only started doing this because someone noticed you got better performance from the early models by straight up writing "think step by step" in your prompt.
It would actually take more work to condense that long response into a terse one, particularly if the condensing was user specific, like "based on what you know about me from our interactions, reduce your response to the 200 words most relevant to my immediate needs, and wait for me to ask for more details if I require them."
* this time last year they couldn't write compilable source code for a compiler for a toy language, I know because I tried
SOTA today has a different set of caveats, of course.
Alternative approaches like "reasoning in the latent space" are active research areas, but have not yet found major success.
I'd hazard a guess that they could get another 40% reduction, if they can come up with better reasoning scaffolding.
Each advance over the last 4 years, from RLHF to o1 reasoning to multi-agent, multi-cluster parallelized CoT, has resulted in a new engineering scope, and the low hanging fruit in each place gets explored over the course of 8-12 months. We still probably have a year or 2 of low hanging fruit and hacking on everything htat makes up current frontier models.
It'll be interesting if there's any architectural upsets in the near future. All the money and time invested into transformers could get ditched in favor of some other new king of the hill(climbers).
https://arxiv.org/abs/2602.02828 https://arxiv.org/abs/2503.16419 https://arxiv.org/abs/2508.05988
Current LLMs are going to get really sleek and highly tuned, but I have a feeling they're going to be relegated to a component status, or maybe even abandoned when the next best thing comes along and blows the performance away.
I analogize it as a film noir script document: The hardboiled detective character has unspoken text, and if you ask some agent to "make this document longer", there's extra continuity to work with.
Some applications hide the reasoning tokens from view, but then the final answer appears delayed.
I tried using a custom instruction in chatGPT to make responses shorter but I found the output was often nonsensical when I did this
I occasionally go back to o3 for a turn (it's the last of the real "legacy" models remaining) because it doesn't have these habits as bad.
Over the last few years I’ve rotated between OpenAI and Anthropic models on about a 4-5 month cycle. I just started my Anthropic cycle because of my annoyance with the GPT-5.2 verbosity
In four months when opus is annoying me and I forget my grievances with OpenAI’s models and switch back, I’ll report back lol.
Like, no, stop that! Keep my engineering life separate from my personal life!
Asking it to be shorter is like doing fewer iteration of numerical integral solving algorithm.
Like, my guy, I don't want to keep prompting you to make shit better, if you're missing info, ask me, don't write a novel then say "BTW, this version sucked"
Yes, I know this could probably be resolved via better prompting or a system prompt, but it's still annoying.