> How would you even begin to assemble a list of prompts that would cover all of knowledge?
They don't need to, all that knowledge is already in public training sets from scraping the internet.
What is harder to get is the answer patterns to ensure good user experience. You want a lot of such answer patterns so the model knows what structure to use for different kinds of questions. That structure isn't to make the result formatted for humans, but contains reasoning paths the LLM takes in order to arrive at reasonable answers. Since an LLMs thinking is the words it writes the structure of how it responds thus corresponds to thinking patterns, and you want the LLM to learn a lot of those, internet data wont contain that but ChatGPT responses will.
TLDR: The words an LLM writes is also what it thinks, since it doesn't hide thoughts, so by training on what ChatGPT writes you also train on how it thinks. It isn't the facts you want but the thinking patterns.