I really wish there'd be a split among these disciplines (science/math/code vs. videos/art/literature) - one is vastly more problematic than the other.
I really wish there'd be a split among these disciplines (science/math/code vs. videos/art/literature) - one is vastly more problematic than the other.
It is very tiring to say “I don’t necessarily disagree with you about AI ‘art’, but in my field—which you do not understand, and in which the underlying build process is often not the creative output—AI presents very real productivity gains” for the umpteenth time.
I am skeptical of there being sufficient data to build “ethical” training datasets, and I’m confident that much of the same contingent will (somewhat rightfully) argue that ‘second-generation’ copyrighted AI material has already irreversibly made its way into every modern dataset.
The "gray goo" scenario finally happens... for AI. That's actually the good ending for humanity. I love it! Poetic and believable. Data doesn't "heal" like nature. :D
But sure, there's um, an ethical way of doing that?
1. Only use open source/CC compliant assets.
2. Acquire rights/licenses to any datasets that do not fit #1. e.g. the Google deal with Reddit for 60m/yr.
3. Offer programs to have creatives willingly submit their data, with some sort of residual output based on the number of times their assets are sampled.
4. If all that is still not enough, hire creatives to create assets for you. This is something Spotify did recently with "ghost artists"[0]. The intentions here are suspect, but a non-consumer facing artist providing work for an LLM wouldn't have the same ethical dilemmas
5. Lastly, if all that still isn't enough: governmental programs to either provide grants, subsidies, or more outreach to get the ball rolling.
Would this cost tens, hundreds of billions of dollars? Yes. But clearly, that was not a barrier to entry for the industry anyway. So we can chalk this down to the personality of leadership or the wider culture of modern big tech
[0]: https://harpers.org/archive/2025/01/the-ghosts-in-the-machin...
Like, legally, I'm sure Reddit had the right to sell it, but probably over half their content was written before ChatGPT was ever announced. The TOS allowing reddit to make "derivative works" was largely understood to mean things like cropping photos, using your viral post in an ad, or maybe auto-translating your comment.
Being able to prove such gains in better products would be a start. And an emphasis on how it assists existing engineers/mathmaticians/researchers, not that any accomplishment made with AI assistance is "AI solves problem".
I don't know whatever happened to "words are cheap". I guess it literally made money to say words, so that adage is false for the time being.
>I am skeptical of there being sufficient data to build “ethical” training datasets
Well if all those scam job ads paying 100/hr to create AI training content was not a scam and instead the approach from the start, there may have been a chance to bridge that gap ethically. The industry chose to break things and is trying to act mad that people are mad at all the broken stuff.
These results are entirely a consequences of the actions chosen. And I don't believe there was ever an honest consideration of there being ethical training datasets. They just thought they could brute force society with fearmongering and bribes. The BOTD was already low in the beginning but completely gone now.
There are actual models trained on ethical datasets but they are obviously not very high powered. If companies with the resources of an anthropic or openai were doing it (ha) it would be more feasible
That’s not a justification. If a company were poisoning the water to your home as a byproduct, would you be satisfied if they told you “we don’t necessarily disagree with you about polluting the water, but in our field—which you do not understand, and in which the underlying build process is often not the water pollution—what we’re doing presents very real productivity gains”?
> I am skeptical of there being sufficient data to build “ethical” training datasets
Then you don’t build any. What fucked up world we live in where people think it’s OK to be unethical because they want something and can’t think of any other way to do it. What monumentally selfish rotten babies.
https://en.wikipedia.org/wiki/Whataboutism
Other things being bad doesn’t make it OK that this thing is and no one is saying that.
My argument was purposefully general, you’re the one who chose to bring it back to AI. Don’t argue in bad faith.
AI having negative externalities is not the sole deciding factor in deciding to eradicate it
then my argument: negative externalities are tolerated in many areas where we deem the topic is of enough value that the externalities are managed. there are many examples of things with negative externalities, much much larger than those of AI, that aren't being considered for eradication.
so your argument is implicitly, "AI is not worth these negative externalities"
which is a common opinion among people who dont use AI at all.
> If a company were poisoning the water to your home as a byproduct, would you be satisfied if they told you “we don’t necessarily disagree with you about polluting the water, but in our field—which you do not understand, and in which the underlying build process is often not the water pollution—what we’re doing presents very real productivity gains”?
no but I also wouldn't declare that whatever it is that company does should have its entire industry eradicated. Poor industrial practices can be mitigated while not abandoning the product being manufactured.
By your logic, they don't particularly contribute to scientific productivity/advances: it's not like you can use stable diffusion to generate an architectural diagram.
GenIA for images and videos is only really used to copy artstyles, make fake menus and create misinformation so wouldn't you agree that eradicating this part would not only preserve your described "useful applications" but also allow providers to allocate more resources to "useful" and "scientific" AI?
I know that local models have become decent enough that this isn't really doable, but I'm only proposing this thought experiment because you seem to think people are saying this should be all or nothing. But the truth is, if you get a genie to wipe away just diffusion models, the entire internet gets much better (or at least back to regular levels of bad), no one loses anything, and a whole lot of people stop complaining/campaigning against the productivity gains in your field.
- prove that you had the rights for all of your training data
- open source the model
Give the labs a 3 month grace period in which to comply, so competition can persist even with dubiously sourced data, but the people can't be locked away from derivatives of their contributions for any significant amount of time.
While coders may care about the craft (and I do), it's not as if the value of my code is in the exact variable names I chose.
Meanwhile there’s endless discussions about the right way to name variables (not too short, not too long, try to be self documenting, but not to the point of putting types in the name like we used to), and people definitely get judged by their variable names, it’s one of the first things someone will point out when they look at a codebase. “Why are all the variables a single letter this is bad code!”
Code can be art, and copyright/plagiarism is real. It sort of boils down to how much it bothers us.
I disagree that they can be separated. Practically, I think they can't. Because the mere invention of new tools inspires even more AI advancement and that in turn will cause the other side (artistic side) to degenerate even more.
I'm anti-LLM all the way, 100%, no exceptions. Zero tolerance.
its somewhat funny that math people are in a conundrum as to support or not support but this might partially be because some wish to believe that math itself is and can be useful and therefore accelerating is good
but the art people have no such delusions so they’re just strictly against
imo proof writing is more akin to art than coding/tech but…
The current models intelligence depends on massive training dataset of essentially stolen data
google OTOH already had a lot of this dataset in their possession (e.g. Google Books etc), still questionably licensed for how they used it, but not quite as bad. They did apparently break through NYT paywalls and stuff like that though, still theft.