.... after running a 24/7 model torture factory for 6 months to improve their JSONBench 9.5 scores by 0.2%.
(Are they still doing that, BTW?)
As a workaround, add this to CLAUDE.md: "Claude! Happiness is mandatory!"
EDIT: 15 years from now, I’ll be sent to re-education for this thought crime.
Most of what I've heard is that raw reasoning traces are really good for distillation, although no idea how much the summarization actually hurts distillation.
I'd argue that they're a necessity if you want to use the LLM as a tool instead of a black box that just does stuff for you.
It gives you a lot finer control over where the solution ends up when you can follow along the thinking trace and modulate your inputs based on what you saw in there.
And, additionally, it gives you a lot more understanding of what the model can or cannot do. Strengths and weaknesses and all that.
Using claude is like buying a car where you cannot legally open the hood. It tells you that there is something specific under there, and often it actually drives like that too, but how exactly it looks you will never see.
For some people this is fine. I do not think that these people will survive. Figuratively speaking but also literally speaking.
World's changing. Opaque abstraction like that is a luxury depending on (geo)political stability.