I had similar issues when training personal models for https://meraGPT.com A meraGPT model is supposed to represent your personality so when you chat with it you need to do it as if someone else is talking to you. We train it based on the audio transcript of your daily conversations.
The short answer to how abilities like in-context learning and chain—of-thought prompting emerge is that we don’t really know. But for instruction-tuned models you can see that the dataset usually has a fixed set of tasks and the initial prompt of “You are so and so” helps model align it to follow instructions. I believe the datasets are this way because they were written by humans to help others answer instructions in this manner.
Others have also pointed out how RLHF may also be the reason why most prompts look like this.