What I wonder is how far copyright protection goes, specifically when AI algorithms are involved. Suppose you "copy" a recipe + description (or a book for that matter) in the following way :
1) you create markov chains using a large number of cookbook texts
2) you create the markov chain for a specific recipe's description and go down alternate high-probability paths in an effort to creatively change the wording without changing the content (maybe this is comparable to the way a human would if he tried the recipe and then wrote down "how he did it")
3) you republish the result
Did you violate copyright or not ?
The reason I ask is that this is often used as an end-run around patents and copyrights. What most speech recognition programs do these days is preprocess the sound + feed into neural network + get output. (and while most "pattern recognition" algorithms for non-visual things seem to have a love affair with support-vector machines, temporal neural networks are certainly advancing there too).
Now if you analyse what those neural networks do there's 2 types of things
1) ~40% effectively is unrolled loops of (mostly) patented algorithms
2) 60% you effectively don't recognize
(to be fair it takes hours of seeing the network operate before you realize anything it's doing)
Clearly this is legal, in cases human implementations of the exact same algorithms wouldn't be, nicely sidestepping the problem of patented algorithms. And as a bonus you don't really have to know the subject matter (e.g. you can write a pretty good voice recognizer without 8 years experience as a linguist. Or you can write them for languages you don't actually know). And as a bonus, academics are miles ahead of the private sector where it comes to machine learning algorithms, so extremely useful thing are effectively free-for-all.
Of course this is also what humans do. If you look at neural networks, they can only do what they've "seen" happen before, or they can combine various things they've seen before. But they are utterly incapable of coming up with original work. So I don't think humans are any different to machines when it comes to producing original work based on combinations of previous works. This doesn't mean the thing that was copied was itself copyrighted, you can write about your own life, for example, or about nature, or ... but we'd call that original, when (in a strict mathematical sense) it's not.
So clearly we've de-facto accepted in our society that machine-processed works at some point start constituting original work.
Do we have any data what point that is ?