Now let's say for any given length L, there exist programs P1..Pn that are generable from Copilot, given some description(s) D1..Dn.
But not even programs, because those have interfaces and implementations. Wasn't there something between Oracle and Google where copying the API was okay?
So let's say for some string of length L, there exist modules M1..Mn that are generable from Copilot, given some description(s) D1..Dn.
We say those modules are reachable, and modules are divisible so long as an alternate module Ma exists whose outputs are exactly the same, even if some portions of the string are substituted.
But Copilot could have generated those as well. Any alternate modules also exist in the generated space M1..Mn.
Now let's say instead of Copilot, we have a box with listings of modules M1..Mn of any length L. Assume retrieving all modules of a given length is instant.
Let's say we pull out a bunch of these modules and compare them to the copyrighted code. Let's say they are exactly equal.
Copyright has therefore been violated, right?
So we establish even a copyrighted work is an instance of a collection of characters that came from a box with infinite memory.
Then we may need a service that scans Copilot output and adds text like, "We recognize this could be copyrighted, which in this universe represents actions requiring legal compliance; yet our only duty is to inform, not take action."
Because is it not true that engineers must both 1) reduce stakeholder damage and 2) inform, not persuade?
So whether or not a copyrighted work is reused without permission, user assumes all responsibility, even if that law changes ten or a hundred years in the future.
Then it's not necessarily Copilot's requirement to eschew copyrighted output. It merely generates the most probable next token, given some existing collection generated. (Or something like that?)
Computation has increased to cover most, if not all, permutations ever possible.
The artist executes the definition of a `line()` method in a drawing package. The description D is opaque! It later takes the form of interviews, with curators and historians presenting views on some hidden data table.
We can also observe as well: suppose description D contains the exact copyrighted code to generate, and Copilot is asked to echo the output.
Since Copilot "generated" it, should we now accuse Copilot of violating copyright?
(But we assume the operator acted in good faith and did not knowingly include copyrighted code in the input.)
Okay, now let's turn to training. Suppose copyrighted works were used in Copilot's training. If we accept Copilot is an instance of our infinity box, no one owns the algorithm to the English language.
It follows that any string generator modeling probability space is merely pulling values out (in finite time).
Now we may say, "We can't ignore time. Our universe proceeds by time." So let's say copyright is important, by virtue of time.
Then only humans can violate copyright, not software. If the law extends to generated instances, it may overreach, attempting to control the babel box.
If the law defines "publishing to public spaces" a copyrighted work, that's another matter.
I say these things typing into a tiny text area. My source control is copy-paste. Of course, I reserve the right to change my mind, contradict myself, and basically be totally wrong.