Whisper (speech-to-text transcription model) sometimes gets stuck in an endless loop of repeating the same word. It's a known problem for lots of our modern transformer/attention-based models.
One of the hacky ways to avoid this is to ensure the last few output tokens aren't too similar. Some postprocessing filter watches the last few produced tokens. When it notices the model starts to repeat itself, the postprocessing usually perturbs the next token a little, e.g. taking the nth-top token instead of the most likely one.
Perhaps a model that genuinely wants to repeat the output instead needs a bigger "kick" to the representation which puts it somewhere completely different in the semantic space? Idk, just pulling this out of my hat.
(I have no idea how ChatGPT handles this or whether the raw implementation suffers from this problem, but Whisper-cpp has some manual entropy regularization postprocessing stuff to avoid getting stuck like that. It's super hacky and often doesn't help.)