Anyone care to comment on why the output of these models changes so dramatically given so little Q&A training? It's a 6 billion parameter model with only 50 thousand Q&A samples.
It's clear the model already "knows" the format of a Tweet (short length, attention-grabbing, contains hashtags). The model also knows stuff about language models (word2vec, tokenization), and can include entities from the question in its response (Dolly, Databricks). Yet, it just doesn't put these pieces together in the right way without the Q&A training.
Edit: For kicks, I asked GPT-4 this question: https://imgur.com/a/sM4uyBn