Not sure how you're prompting it but remember that it's not trained for chat or instruction following, it simply takes the text given to it and tries to continue it. Give it the right prompt structure, and it can (at least sometimes) output coherent completions, far more often than you'd see in a Markov-chain. Also, this version is more or less equivalent to the smallest version of GPT-2; the largest version was 1.5 billion parameters and was much more likely to generate impressive (at the time) output.
The assumption that LLMs would always need sophisticated inputs to generate useful outputs is where the term "prompt engineering" came from. Now that idea is basically dead. Absolutely wild how far these models have come in less than a decade!