I guess the likelihood decreases as the code length increases but the likelihood also increases the more constraints on parameters such as code style, code uniformity etc you pose.
I guess the likelihood decreases as the code length increases but the likelihood also increases the more constraints on parameters such as code style, code uniformity etc you pose.
That's just copying with extra steps.
The way to do it legally is to have 1 person read the code, and then write up a document that describes functionally what the code does. Then, a second person implements software just from the notes.
That's the method Compaq used to re-implement the original PC BIOS from IBM.
wasn't it to have one person run tests of what happened when different things were done, and then write up a document describing the functionality?
In other words I think one person reading the code is still in violation?
>by reverse engineering and then recreating it without infringing any of the copyrights associated with the original design.
reverse engineering is not 'reading the code'.
Except that with AI we can more easily (in principle) provide provable provenance of training set and (again in principle) reproduce the model and prove whether it could create the copyrighted work also without having had access to the work in its training set
Maybe, but that would still be copyright infringement. See My Sweet Lord.
If someone somehow managed to do that and then happened to have accidentally copied someone's code, how believable would their argument be?
No, and humans who have read copyrighted code are often prevented from working on clean room implementations of similar projects for this exact reason, so that those humans don't accidentally include something they learned from existing code.
Developers that worked on Windows internals are barred from working on WINE or ReactOS for this exact reason.