Of course they're not going to stop at just code. They need all the rest of it as well.
Of course they're not going to stop at just code. They need all the rest of it as well.
It's trivially easy to get claude to scrape that and regurgitate it under any requested licence (some variable names changes, but exactly the same structure - though it got one of the lookup tables wrong, which is one of the few things you could argue aren't copyrighted there).
It'll even cheerfully tell you it's fetching the repository while "thinking". And it's clearly already in the training data - you can get it to detail specifics even disallowing that.
If I referenced copywritten code we didn't have the license for (as is the case for copyleft licenses if you don't follow the restrictions) while employed as a software engineer I'd be fired pretty quick from any corporation. And rightfully so.
People seem to have a strange idea with AI that "copyleft" code is free game to unilaterally re-license. Try doing that with leaked Microsoft code - you're breaking copyright just as much there, but a lot of people seem to perceive it very differently - and not just because of risk of enforcement but in moralizing about it too.
Source: find literally anything on GitHub using dependencies that are MIT licensed and being distributed without following the terms that state you must also redistribute the licence for each
I think that is one of the main reasons there is so much pushback against this, a lot of people are now addicted to their stream of washed code and want to claim ownership over what is essentially a derived work. The key then becomes 'if a work could not have been written by the author that claims it does that claim survive'. I think it should not but there is plenty of disagreement on this.
Your idea of how humans use content under copyright is mistaken.
I want to make an ajax request using jQuery. I look up an example in StackOverflow. I use a very similar code to the example given in the post and by not giving any attribution I just claim ownership.
Same with Spring in action books or looking up Java class references. Many times I look something up and use it as reference just tweaking the examples given.
Millions of programmers have done this.
LLMS in principle use the training data to generate an answer to the prompt, similar to the process I described.
Is there even any evidence that "crypto bros" and "AI bros" are even the same set of people other than being vaguely "tech" and hated by HN? At best you have someone like Altman who founded openai and had a crypto project (worldcoin), but the latter was approximately used by nobody. What about everyone else? Did Ilya Sutskever have a shitcoin a few years ago? Maybe Changpeng Zhao has an AI lab?
That was a biometric surveillance project disguised as a crypto project.
> Is there even any evidence that "crypto bros" and "AI bros" are even the same set of people
No, the "AI" people are far worse. I always had a choice to /not/ use crypto. The "AI" people want to hamfistedly shove their flawed investment into every product under the sun.
Has it been adjudicated that AI use actually allows that? That's definitely what the AI bros want (and will loudly assert), but that doesn't mean it's true.