Basically it's a trick that recognizes that language transformers only perform computation to generate words so for complex tasks you can get better results by asking the model to explain its chain of thought and only give the answer at the end. This has the effect of giving the model "time to think". If it didn't generate those words it wouldn't have anything to hang that computation off since it is fundamentally a word-generation model.
Here's a simple example: keyterms from text are extracted with text detection from an image. Those keyterms will sometimes have bad reads where "aligned AI" might pop out as "aligned Al". A subsequent "internal thought" would be formed and ask, "What's wrong with the 'aligned Al' keyterm?". If an updated response is returned, we use it instead of the original output.
I think this is the paper that really kicked off this technique: https://arxiv.org/abs/2201.11903
You can google for more results / papers.