easy stuff happens by itself, but with a system large enough you need a scratchpad and a rubber duck.
easy stuff happens by itself, but with a system large enough you need a scratchpad and a rubber duck.
why we don't do GAN here, ie. second model verifying correctness/accuracy/etc. ?
Closed models probably do the same thing internally. What is shown externally is different though: you get a summary of the chain of thought, not the thoughts itself. This is done to prevent distillation.
The latest look we had at a frontier chain of thought is probably in the Huggingface incident report - I haven't actually read it yet but I saw the BlackHat talk, and it included some snippets. The thoughts look like they are approaching neuralese. The words are still understandable but the grammar is weird, simplified. In comparison, Qwen 3.8 27b thinks in valid English.