There is no literal “human in the loop” for generations, of course, but the model is fine-tuned on examples written by human contractors of instructions being given followed by correct responses. I assume that training is essential to it being able to follow directions of this length, or really any directions at all. If you try using the pre-InstructGPT version of Davinci (model=“davinci”, not model=“text-davinci-002”), you’ll find it’s as cumbersome and annoying as you remember GPT-3 being in 2018.