Or, if you're okay with a smaller LLM, go here, set temperature to zero and enjoy repetition: https://transformer.huggingface.co/doc/distil-gpt2
That post-processing infrastructure can use all kinds of mechanisms to prevent loops and induce more useful output. With the most basic simply systems simply refusing to select any output token that already in the input, to more complex stochastic process that explore the tree of possible outputs, to find branches whose overall result is improved by choosing less optimal immediate steps.
The vast majority of what make Chat GPT different to simpler GPT-3 models is this complex post-processing phase that allows designers to push and pull on the behaviour of the pre-baked static model underlying the chat interface.
My guess is that they have some kind of state-aware sampler, and they know if they are in the table, etc. Because then you can sample in a much better way. Just like with grammars, but grammar itself is probably not enough.
Sampling and tokenization are IMHO the biggest open problems. We have something which kind of works but it's nowhere close to be perfect.
There was an article/paper showing that GPTs (whole family) quickly get confident in looping if there is anything loop-like in the window. So basically, as soon as it loops once, it will never get out of it.
Note that sometimes loops are desired, like with docstring in the code, always starting before the function definition, and other structural things.
For example, here's a function from llama.cpp that applies repetition penalty: https://github.com/ggerganov/llama.cpp/blob/master/llama.cpp...
Here's the one from transformers: https://github.com/huggingface/transformers/blob/0a55d9f7376...
To summarize how they work: you keep some number of previously generated tokens, and once you get logits that you want to sample a new token from, you find the logits for existing tokens and multiply them by a penalty, thus lowering the probability of the corresponding tokens.
https://platform.openai.com/docs/api-reference/completions/c...
See presence_penalty and frequency_penalty.
Sampling techniques is one of important arts of LLMs, you'll can find a lot of papers on them.
In general, smaller are more prone to repetition, but you can get caught in it even with larger models.