I'm not an expert, but I believe it is the act of the model itself being used to align itself with the intended outputs[0]
At a glance, this looks like a model pretrained to perform prompt-engineering. It should automatically use Chain-of-Thought in its responses in order to improve it's programming abilities, and, therefore be better aligned with users expectations.
It also has reflection. So they include code to execute the model output and return the response to the model for feedback.