DeepSeek V3.2 was tried without reasoning and it got 57% on ARC AGI 1. It's a 7 month model, so I'm pretty confident that base LLMs would be able to solve ARC AGI 1 without reasoning/CoT.
https://arxiv.org/abs/2607.06764
The base model got 15%. They built an elaborate looping harness that allows it to burn 100k tokens "thinking" about the problem, which got the 57%. This is just an alternative approach to reasoning.
So this paper doubled the price to get the same exact result at base Deepseek 3.2 at launch, and wasn't even tested on the verified set.
https://huggingface.co/datasets/arcprize/arc_agi_v1_public_e...
"kwargs": {
"max_tokens": 100000,
"stream": true,
"reasoning_effort": "high",
"rate_limit": {
"rate": 2,
"period": 60
}
}
The other paper ran Deepseek v3.2 without reasoning as a baseline and got 15.5%, which is much more in line with other base LLMs like GPT-5.2.