I don't get any attribution for the code or messages I wrote that GPT4 was trained on, and there are multiple examples of LLMs regurgitating people's code (even copyrighted code) without attribution.
I don't get any attribution for the code or messages I wrote that GPT4 was trained on, and there are multiple examples of LLMs regurgitating people's code (even copyrighted code) without attribution.
Point being, there’s a legit conversation to be had about AI training data, but it’s not the same conversation as humans intentionally creating a derivative work. I think it’s too reductive to insist those are the same thing.
The training dataset is just upstream of the fine-tuning step but both are supplying vital data to the whole pipeline.
Still rude though. Rude on rude.
I fail to see how anyone with your viewpoint isn’t flat out being intellectually dishonest. It is understood that it is both (usually) legal and ethical to not attribute someone when they teach someone an abstract skill set which they use to develop something “new”. LLMs inarguably bring us closer to reaching that point without implicating human effort in the creation process. “You did it with a computer, so copyright!”is such a naive and frankly wasteful take as it threatens to forego massive societal advances in favour of intellectual property rights. What sort of society is that!? An overly American one, that’s for sure.
This isn't a student learning from a teacher, its a student copy-and-pasting an answer they found. It isn't learning an abstract skill set which it is using to develop something new, its literally just copying the code.
It is utterly dishonest to make the comparison you do in your comment.