> It is known that it produces verbatim copies of sections of code
This happens only rarely, like under 1% of the time. It happens mostly for well replicated code and not so much for code that only appears once. It can be filtered out with search and bloom filters of ngram hashes.
But the prompter can goad the model into copyright infringement by quoting the start of a copyrighted text verbatim, and asking for completion. The longer and more precise the prompt, the higher the chance of regurgitation. So, when it happens, we're often "asking for it".
Both regurgitation and hallucination seem to be LM problems we can tackle. They are complementary - in one we don't want the model to replicate the training data exactly (be creative), in the other we don't want the model to invent facts out of thin air (be factual). Both can be tackled by using search for reference testing.