One thing I notice that is huge is being able to see the chain of thought lets me see when my prompt was lacking and the model is a bit confused on what I want.
If I was anymore impressed with R1 I would probably start getting accused of being a CCP shill or wumao lol.
With that said, I think it is very hard to compare models for your own use case. I do suspect there is a shiny new toy bias with all this too.
Poor Sonnet 3.5. I have neglected it so much lately I actually don't know if I have a subscription or not right now.
I do expect an Anthropic reasoning model though to blow everything else away.