(text commentary below the comic, in case your OS has decided to conceal the scroll bar from you and you didn't notice the page is longer)
(text commentary below the comic, in case your OS has decided to conceal the scroll bar from you and you didn't notice the page is longer)
Boy does that ring true.
> Meta: We stole so much stuff that, actually, we didn't steal any stuff
> Google: If things were different, things would be totally different
> Andreeson Horowitz: But we already spent so much money
> Microsoft: Think about how this would hurt the little guys, like us
> Anthropic: Fucking shut up about it
> Hugging Face: It's super legal but honestly it might not be idk
> StabilityAI: It's legal in places that aren't here
It's worth noting that, back when it suddenly became technically feasible to copy all songs for free, the response from the rightsholders of music was, essentially, to demand that we all collectively pretend it isn't. It might make sense for the rightsholders of training data to try to do the same, if they could speak with one voice and if they had lobbyists, but I guess the fact that they can't and don't settles that.
To me, training a model is less like redistribution and more like reading something and having it influence your thoughts. You may be able to reproduce certain memorable sections of things you've read verbatim (common with poetry, for example) and you may well be able to reproduce a 'style', but you don't have the entirety of the source material memorized in a way that it could be redistributed in full (but if you did, that may be well an infringement at such a time).
If Steve the megarapist is writing articles about the latest popular injustice, does that make Steve virtuous?
If he applied his approach to Python coding, he'd be one of those folks who isn't happy until their entire program is a single `do()` method that executes by using decorators to synthesize classes and metaclasses out of thin air, that themselves generate the Python code as strings, that itself gets `eval()`'d to execute the intended program.
It's not for you.
(this is a reference to an ancient Penny Arcade bit: https://www.penny-arcade.com/comic/2004/03/24/the-adventures...)
> (this is a reference to an ancient Penny Arcade bit: https://www.penny-arcade.com/comic/2004/03/24/the-adventures...)
Thank you for posting this. I remember it, had forgotten it, and re-reading it after completely forgetting about it is almost as good as reading it for the first time. It still holds up 20 years later!
I for one really appreciate seeing language being used so artistically. (Though if I was reading a technical spec, product description, legal contract, etc... absolutely not.)
Perhaps if you read Gabe's post you could have saved yourself the trouble of making this comment.
Seems lazy to call a genAI model "intelligence" without providing any substance as to why one should believe that, but OK.
The model isn't an "intelligence" because it's not making a choice about which data you train on and be "inspired by", as people here say.
The humans that operate those models do. It's those humans that are exploiting the artists.
The AI, as many have said here, is just a tool.
Funny how it's "just a tool" when a human uses it to create art, but it's "intelligence" that gets "inspired" by art when a human uses it to create a software product.
There’s no evidence any of us are making choices either. I don’t think LLMs are intelligences in the way most people mean the word (I think that would be vastly overestimating their abilities) but appealing to choice when it’s so poorly understood — as poorly understood as intelligence, even — is not useful.
There is very clear evidence that the humans that operate the computers make the choice on which data the model gets trained on.
The discussion is about those choices.
We don’t understand what it means to be intelligent in the way people generally mean, and we don’t understand (scientifically) what it means to choose. So we can’t use choice to usefully define intelligence.
"I'm sorry, your honor, there is no scientific way to demonstrate that I have made a choice to shoot the victim dead. Therefore, I am not guilty!"
The way that an AI gets a data set to be trained on is via a choice made by humans. You can't reason your way out of the fact that a choice is made, that it's not made by the AI, and that it has consequences.
But sure, let's pretend the conversation is about intelligence of an algorithm that has no agency and no say in the matter.
Being obtuse on purpose is fun!
This is a very poor definition of intelligence and is not convincing of deserving training data
Models aren't capable of existing without humans in the loop making the choice of which data they get trained on.
Consequently, those humans and their choices need to be regulated.
That doesn't mean they arent stolen and that doesn't mean they are fine cuisine.