The problem here ends up being that code, especially in popular languages, will always looks similar when you're doing something like finding the best implementation for an algorithm. So if you invoke CoPilot for a common problem, chances are it can pull the exact code it needs from is dataset, but it also could've generated that same code snipped had the solution not existed in its training dataset. And when you start out solving a problem then ask it to continue writing more code, it just assumes you're solving the exact same problem that the original source code was solving.
This could probably be remedied if CoPilot spit out a "this is % similar to <x> source code from the internet" so that you can know just how unique CoPilot is being. Legally, copyright is just a mess and was not ready for the scale of the internet nor the advancements in ML when there are machines that have a 50% chance of infringing on someone's copyright and 50% chance of creating something new.
Unfortunately this is frequently abused where researchers build a model under the exemptions, and then others use that model commercially, even if they wouldn’t be allowed to build that model directly themselves.
Anyways, the scientific progress would continue, but products would halt until product developers get some kind of agreements with content creators (eg maybe people start adopting a new kind of open-ish license).