I don't think such penalties were applied to GPT-2 or even GPT-3, yet they weren't repetitive like that.
https://platform.openai.com/docs/api-reference/completions/c...
See presence_penalty and frequency_penalty.
Sampling techniques is one of important arts of LLMs, you'll can find a lot of papers on them.
In general, smaller are more prone to repetition, but you can get caught in it even with larger models.