So, counter-intuitively it may strengthen Stock Image Sites value
So, counter-intuitively it may strengthen Stock Image Sites value
[1] https://twitter.com/kevin2kelly/status/1551964984325812224
The same prompt with DALL-E gives me proper silhouettes with no watermark.
[1] https://github.com/CompVis/latent-diffusion [2] https://imgur.com/a/8tOI9QU
The minute Dall-E becomes a threat, these sites will enforce Dall-E to retrain from public / creative commons images. But Stock Image sites will and be one-up on Dall-E
It's almost you don't how to play chess or business strategy.
Apparently using copyrighted training data is okay in EU: Directive (EU) 2019/790 … on copyright and related rights Article 4 [1]
[1] https://eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=CEL...
Be sure to pass this along to Microsoft's GitHub Copilot.
So even if you are right, and courts rule that training is not fair use, it seems likely that all the big tech companies would lobby Congress to restore the prior state of affairs, as not having an ML industry is somewhat of a national security issue if the rest of the world is going full steam ahead on making Skynet.
I swear, world politics and economics isn’t much different from that, intellectually….
Since there is no legal precedent and the law itself isn't clear about this use case, it's basically a huge gamble on a legal gray area at this point. For VCs the risk doesn't really matter as ML startups only need to exist long enough to provide an exit with high ROI and for enterprise companies it doesn't matter as ML products are just one of many ventures for them.
It's worth noting that unlike Germany, where book and newspaper publishers have won rather unusual copyright claims against companies like Google, in the US the big publishing industries to worry about are movies and music, and most ML projects right now seem to focus on generating images or text rather than music or video. If "AI generated music" caught on like DALL-E 2 did, I think we'd see a lot more contention over how copyright law applies to ML training data.
So we could actually follow copyright laws and still have an ML industry. But I’m not sure big tech wants to ask the public for consent. They would rather have free reign to do whatever they want.
By the way I don’t like copyright law or the concept of IP. But I find it a little annoying that I’m supposed to respect IP law and ML stuff can just ignore it. Also if big tech was forced to encourage people to share stuff with an open license, this would be a huge net good for society! Instead nothing changes but big tech gets to take advantage of peoples copyrighted works and artists can’t do anything to stop it. That kinda sucks.
Today’s ML output throws everything into a mixer and then blanket-calls it ML-generated output, because the original training content has been separated from the social and legal frameworks that govern it.
A NN producing a copyrighted work without respecting the license is not fair use, we know that much.
The funniest ones will be where ML independently reproduces a picture from unrelated materials with humans making a futile effort trying to figure out how it obtained this result.