Aren’t all of these reasoning models?
Won’t the reasoning models of openAI benchmarked against these be a test of if Sam is losing?
Won’t the reasoning models of openAI benchmarked against these be a test of if Sam is losing?
With Gemini (current SOTA) and Sonnet (great potential, but tends to overengineer/overdo things) it is debatable, they are probably better than R1 (and all OpenAI models by extension).