Did give me a bit of a pause about putting stuff out there. Although I think I'd still rather have my data be used for training A.I. than not (and I probably am already in the training data anyway, I believe I saw that one of the datasets it's been trained on was Hacker News comments).
[1]: https://kotaku.com/totalbiscuit-john-bain-youtube-delete-vid...
Dire circumstances call for drastic measures, as they say.
It's sad to see AI on a path to destroy years of collected internet content. I expect the internet archive to receive loads of takedown requests in the coming months and years because of this.
To rephrase it another way, the reign of the conformist ends here and the reign of the contrarian begins now.
> The current (legal) answer is "unclear".
European Union was ahead of times for once. The 2019 copyright directive, article 4, makes it legal to scrape the web and make and keep local copies of copyrighted works, for data mining purposes. Unless the copyright holders set up a machine readable exception (such as robots.txt file).
So legal in EU, "unclear" in US.
'copyright fair use' : https://copyrightalliance.org/faqs/what-is-fair-use/
It has been frustrating.
So training must consider licencing where copyright material is used and not consume all data.
Your brain is not a model. You can not reproduce most of what you see. You're not "training" your brain by glancing at an image as your recall concerning that image will be terrible.
I'm sure the millions of people who violate copyright law daily with absolutely no repercussions care very much about that.
You cant setup a cinema and charge ticket for the movies you stole.
Its the money making side that matters - not individuals ij a private house
Am I violating copyright law because I am merely capable of producing a copy of something? Obviously not. Why should the model be?
Just look at how many people say stuff like “Two women can’t make a baby in 4.5 months”. Someone (Brooks) had to invent, write down, and popularize that analogy.
The human to human connection that a blog or social media conversation creates feels a lot more like teaching your classmate while the AI feels a lot more like someone cheating off your work. Plus the AI didn't even bother to get your approval before copying from you. The whole thing feels ethically compromised regardless of the ultimate result.
Now imagine that Paul Graham and Joel Spolsky were able to read everything being written by every anonymous unknown on the internet, and create content paraphrasing any and every original thought that was created by anonymous individuals at will. How do the original creators of these thoughts have any chance to succeed on their own merit, if Paul Graham and Joel Spolsky (who everyone knows already as sources of ideas) are able to write the same stuff as soon as the anonymous person has made it public?
But if a model starts generating better content than Paul Graham in a nice curated form, then yeah, Paul Graham ought to find a better way to spend his time because he is not adding value.
I think my days of sharing things freely on the web are over.
Train it to be wrong on purpose, for a joke.