I realize I am posting this on a public forum that is almost certainly being used as a training corpus, but I’ve mostly withdrawn from public social media at this point and I wouldn’t be surprised to see that happen with more people.
I realize I am posting this on a public forum that is almost certainly being used as a training corpus, but I’ve mostly withdrawn from public social media at this point and I wouldn’t be surprised to see that happen with more people.
And now we're unleashing ML training on all that digital, online data. Which industries will discover that this is the thing that means putting your data online, digitally, was a mistake? Certainly artists are feeling it now... maybe programmers, too, a little.
So how do you put the genie back in the bottle? Live performances, with recording devices banned? Distribute written material only on physically printed media - but how to prevent scanning? Or just escalate the DRM war - material is available online, but only through proprietary apps on locked down platforms?
Or is this going to take regulation - new laws to protect copyrights in the face of ML training?
It wasn't always the case, that you could assume that if some information exists, it should show up in a single search. That's an expectation we invented only about 25 years ago. It's possible that the result of all this is that we figure out that we can't actually sustain the free sharing of information that makes that possible.
The problem is, to borrow a phrase: information wants to be free...
Ah! The HN echo chamber. The world at large vastly does not care at all. Govs are banning TikTok apps on gov employees phones all over; does anyone care and use less TikTok? Only HN and probably many there are lying about it. Privacy and this type of content abuse prevention is valued by a handful of people unfortunately.
I don’t think people are going to continue investing time and effort into giving away high quality knowledge on the internet just so Sam Altman can train an AI to repeat it, put it behind a paywall, and charge you for it.
Yes, people will still post memes. How useful is that as training data?
It is also not only Sam Altman; many good people value this information, not limited to open source AI scientists. Let’s not build our lives around a few grifters, otherwise you might as well just quit and do up and flip old houses for money; that’s beyond the reach of AI for the foreseeable future and at least then you don’t have to deal with scammers and people who make society worse at scale.
Virtually nobody's going to care enough to change their behavior.
What are you on about? The Detroit auto industry failed due to Japan producing a superior product.
> By the end of the 1970s, the Japanese automakers dominated the domestic producers in product quality ratings for every auto market segment, representing a formidable competitive advantage (National Academy of Engineering and National Research Council, 1982, p. 99). The quality gap between U.S.-produced cars and foreign cars was beyond dispute (Kwoka, 1984, p. 518) [1]
Because it is a prisoners dilemma.
If you personally stop posting, it has very little effect on the job market.
Because other people will still be posting this content.
This, you are at a disadvantage, even if in agregate it hurts you for this content to be posted.
More likely some sort of physical reaction, hopefully not a violent one.
Why would people suddenly just turn to AI for all their entertainment? Part of the reason people love art is because they relate to it.
Like, in 2023 with the internet, blog posts, youtube, reddit, movies, etc, there's still enough demand for books that people buy them. Just because a new medium is created doesn't mean it suddenly just becomes the only medium.
Internet eliminated the need to go to library to find information.
Hell, even books/writing changed the way people acquired and stored information and people no longer needed to be present in the university. Plato was against writing in a time when writing was just becoming accessible in a widespread way, and in just 2000 years we literally can't imagine a world without writing.
> Throughout human history, anyone can profit off of hearing something you said in public.
I think the big difference here is that this will now be automated. This comment I'm writing right now is being "donated" to any company that wants Hacker News on their dataset, but they didn't even need to go through the work of reading my comment to use it, they just feed along with hundreds of terabytes of random text they get from the internet and apply heuristics/other models to filter the dataset.
I feel somewhat uneasy about it, specially when writing code now. I don't think my code is good in any way shape or form, but there's a reason I license them as GPL, I don't want it used without being contributed back in some way, I write that for the greater good. License infringement was always a touchy subject since it's quite hard to find license infringement in proprietary software, but now that an opaque black box is involved in that process the people write the software might not even know they violated a license.Just throwing ideas, I'm not really in favor or against LLMs using public datasets. At least on one side, it levels the playing field among all participants. On the other side, however, big tech will always have an edge in training and testing these large language models. My current stance is to just wait and see what happens, and then react accordingly.
Don't get tricked by the past decade or so's focus on the idea of "content"! This was all mostly a ploy by YouTube et al to get people to make more channels so they could sell more ads.
There is no real value to your sense of individuality in an economic sense, you shouldn't worry about it being exploited. Don't confuse the (beautiful) experience of your own interiority with a commodity.
But regardless of all that, I feel like this is an incredibly weird way to intepret what he is saying. Its not about the rising value of human generated datasets, its about the fact that there is still necessary well of labor that all this stuff must pull from in order to eliminate other forms of labor. That is why he says "not everyone can use AI." Not "proprietary data will become more valuable."
If it will be valuable, who will own this value? Probably not the same people generating and curating it, and those are precisely the people who can't use AI.
(And there are at least 1000 people making my very point this very moment..)
If there’s one thing I refuse to believe people will ever stop doing, it’s posting.
I’m genuinely curious how much time you spend with non-techie people to come to the conclusion that your opinion comes close to representing the masses. Try to make this argument to a 16yo posting dances on TikTok or a boomer reposting poems about their grandkids.
“Don’t you understand?! The big tech companies are vacuuming up your posts as TRAINING DATA! It’s going to become part of the CORPUS! You have to be like me and WITHDRAW from public social media!”
“Yeah I don’t care lol”