> correctly reason through one of my private benchmarks
Want to say more about these private benchmarks? :)
Want to say more about these private benchmarks? :)
Part of my suspicion is drawn from the thinking trace I got when I tested the car wash problem. That really does seem to have been post-trained; it's too good.
e.g. Gemma 4 26B solves this concisely without adding any filler about fuel economy or how long it will take, but it generally gets there by breaking down the problem in the thinking trace the way you'd expect.
Qwen 3.8 27B is just a little too certain right off the bat in low reasoning mode.