Yes, it does. Let's leave AI out for a moment.
If you lock yourself into your room with no Internet access, and you draw (or write etc) something independently that just happens to look exactly like some already existing copyright-ed work, you are still not infringing any copyrights.
If you produced the same work by making a copy, you would infringe copyright.
See https://en.wikipedia.org/wiki/Clean_room_design for how that's relevant in real life.
However you can infringe on patents and trademarks with independent work, ie even if you don't just make a copy.
How does an AI know what Obama looks like? It memorized thousands of images of him, most of which were by professional photographers and, hence, copyrighted.
I have never seen Obama in real life. If I can draw him because I know him from copyrighted works, is my draw copyright infringement?
In the Obama example it comes from copyrighted inputs. Without the photos in the magazines and newspapers I would not have idea how Obama looks. Any draw I do, it's going to be derivative of those copyrighted photos and, maybe, other inputs more generics that I saw in the past (that could be copyrighted too).
How is that different from what that software is doing? It's not, it's just that it's done in an industrial scale.
This is just the same problem an artisan had when factories started to appear.
Publishers license text to print in books. Music producers license samples to use in songs. Artists license photos and textures to use in illustrations and scenes. _License the source material used in training sets_. _Then_ go out and industrialize art, everyone wins!
Yes they almost definitely are, what else would they be hired for? You wouldn't watch the same news everyday (even if it feels that way)
> the number of training epochs has dropped dramatically
So? Lets imagine a documented nurodiverse person (lets use a photographic memory) "snapsohts" a famous "works" and then sits in a room, after 5 years training, to produce "some work" are they free of copyright claims? Have you not seen the nurodiverse person who draws the entire New York sky line from 1 helicopter ride?
Your approach seems non-sensical, 1 epoch - 10,000 epochs, it doesn't matter - its still a copy / derivative work (which may exempt it, fair enough).
Personally I think all current AI work is just copyright on steroids, DALL-E outputs Shutterstock logos given the right input and Githubs co-poilt is just a lawsuit waiting to happen (implying anyone has the funds to actually sue Microsoft these days, the T&C's of Github must be incredibly broad).
Craftmansship. They have subtle, but probably very good, intuitions about lighting, facial expressions etc. and also a lot of concrete knowledge about these things. All of which an algorithm can pick up on by examining thousands of images once each.
Once or many matters because if you just examine each data point one, there can be no overfitting in the traditional sense. And that's what shocks me about the "one epoch is all you need" fruits.
What makes you so sure? If my algorithm was literally just storing stuff in a hashtable to look up later, you'd get overfitting from a single exposure.
Think of it in terms of updating beliefs about the target distribution. With backpropagation, you predict based on the input, and update your beliefs according to how wrong you were. So in a sense it's unsound to re-use data - your beliefs already incorporate them! And traditional overfitting is all that - it's when you use up all the information in your training data. This was many people's objection to neural nets (and I thought it was a good objection at the time, and thought myself that the future lay with more "sound" methods, which performed better on most metrics anyway at the time, rather than with dodgy biomimicry which wasn't really even similar to biological brains at all).
But yes, there are other types of overfitting if you want to get philosophical about it. It's just that the one I and everyone used to worry about, from training too much on your data, just isn't important anymore. And most of those clever principled and less-principled regularization methods just don't matter anymore!
At what point is a neural network sophisticated enough that it can be compared to human learning for copyright purposes?
[1]: For most people at least. Some will have actually seen him in person, and for those this statement would of course not be true.
Humans aren't in a locked room either. Unless you put them there. Same for AI.
To say the same thing with less snark: in the future there might be a market for AIs trained on 'clean room' data.
That's not true.
Is taking a picture of a painting enough to violate copyright? Is a digital copy of a VHS tape a copyright violation? Is copy/pasting an image a copyright violation?
A computer is not creative because it cannot think. It can only copy what others do; it doesn't understand what it's doing, only how to do something. Equating a human with a computer is unreasonable at the current state of computing and I doubt we'll see a truly conscious AI in the future.
For an automated tool to replicate art into a state such that it would no longer be a mere reproduction of copyrighted material, one would assume that the tool would gain such sophistication that the tool itself should deserve copyright rather than its operators. After all, a commissioner of an art piece may provide the prompt but the art itself is under copyright by the author.
We'll have to see how the courts look at these reconstructions. Personally, I believe tools like DALL-E and Copilot are no more than fancy copy/paste systems and should only be trained on copyright-free materials for their use not to be subject to copyright issues.