You mean the knowledge that Claude has stolen from all of us and regurgitated into your projects without any copyright attributions?
> But I see a lot of my fellow developers burying their heads in the sand
That feeling is mutual.
You mean the knowledge that Claude has stolen from all of us and regurgitated into your projects without any copyright attributions?
> But I see a lot of my fellow developers burying their heads in the sand
That feeling is mutual.
You can't, and shouldn't be able to, copyright and hoard "knowledge".
You can twist this around as much as you like but there are several studies showing that LLMs and and will happily reproduce content from their training data.
Correct. But if read your code, produce a detailed specification of that code, and then give that code to another team (that has never seen your code) and they create a similar product then they haven't broken the law.
LLMs reproducing exact content from their training data is symptom of overfitting and is an error that needs correcting. Memorizing specific training data means that it is not generalizing enough.
That costs significantly more and involves the creation of jobs. I see this as a great outcome. There seems to be a group of people who share the opposite of my views on this matter.
> and is an error that needs correcting
It's been known for years. They don't seem interested in doing that or they simply aren't capable. I presume because most of the value in their service _is_ the copyright whitewashing.
> Memorizing specific training data means that it is not generalizing enough.
Is that like a knob they can turn or is it something much more fundamental to the technology they've staked trillions on?
I don't see it that way. If whatever you're doing can now be automated then it's become a bullshit job. It no longer a benefit to humanity to have a human sit on their ass, stand on their feet, or break their back to do a job that can be automated. As a software developer, it's my job to take the dumb repetitive stuff that humans do and make it so that humans never have to do that job again.
If that's a problem for society, it's because society is messed up.
> It's been known for years. They don't seem interested in doing that or they simply aren't capable.
I don't find that to be particularly big problem. Fundamentally an AI isn't just compressing all human knowledge and decompressing it on demand; it's tweaking parameters in a giant matrix. I can reproduce the lyrics of songs that I've heard but that doesn't mean there is a literal copy of that song in my brain that you could extract out with a well placed scalpel. It just means I've heard it a bunch of times and the giant matrix in my brain is tuned to be able to spit it out.
> Is that like a knob they can turn or is it something much more fundamental to the technology they've staked trillions on?
In a sense, it a knob. It's not fundamental to the technology; if it's reproducing something exactly that likely means it's over-trained on that data. It's actually bad for the models (makes them more incorrect, more rigid, and more repetitive) so that is a knob they will turn.
Also: Over the past 20 years, I could count the number of times on one hand that I was been able to get away with out-right copy/paste from SO.
And comes with a price tag paid to people who neither own nor generated that content. You don't think that shifts the ethical boundaries _significantly_?
I would very much like someone to give me the magic reproduction triple: a model trained on your code, a prompt you gave it to produce a program, and its output showing copyright infringement on the training material used. Specific examples are useful; my hypothesis is that this won't be possible using a "normal" prompt that's in general use, but rather a prompt containing a lot of directly quoted content from the training material, that then asks for more of the same. This was a problem for the NYT when they claimed OpenAI reproduced the content of their articles...they achieved this by prompting with large, unmodified sections of the article and then the LLM would spit out a handful of sentences. In their briefing to the court, they neglected to include their prompts for this reason. I think this is significant because it relates to what is really happening, rather than what people imagine is happening.
But I guess we'll get to see from the NYT trial, since OpenAI is retaining all user prompts and outputs and providing them to the NYT to sift through. So the ground-truth exists, I'm sure they'll be excited to cite all the cases where people were circumventing their paywall with OpenAI.
Then you have been mislead:
https://arstechnica.com/features/2025/06/study-metas-llama-3...
> I would very much like someone to give me the magic reproduction triple
Here's how I saw it directly. Searched for "node http server example." Google's AI spit out an "answer." The first link was a Digital Ocean article with an example. Google's AI completely reproduced the DO example down to the content of the comments themselves.
So.. don't know what to tell you. How hard have you been looking yourself? Or are you just trying to maintain distance with the "show me" rubrick? If you rely on these tools for commercial purposes then the onus was always on you.
> So the ground-truth exists
And you expect a civil trial to be the most reliable oracle of it? I think you know what I know but would rather _not_ know it.
As to your Ars article, I'm familiar because I read Ars.
> The chart shows how easy it is to get a model to generate 50-token excerpts from various parts of Harry Potter and the Sorcerer’s Stone. The darker a line is, the easier it is to reproduce that portion of the book.
50-token excerpts are not my concern, that's 40 words. The argument that the plantiffs need to make is that people are not paying for the NYT because ChatGPT (part of the four fair use pillars, I could expand, but won't). That's gonna be tough. Let's revisit this after the ruling and/or settlement.