Exactly this. I've been very confused by the free software advocates that seemed to hate AI until I realized their reasons for releasing software under an open source license were very different than what I assumed they were.
Exactly this. I've been very confused by the free software advocates that seemed to hate AI until I realized their reasons for releasing software under an open source license were very different than what I assumed they were.
1. The credits for my code are protected (with the GPL you have to tell where your code originates from). That's the ego part. It's important to me because I'm not paid for my software. So credits are an important reward.
2. I want people to think twice about reusing my code. I use the GPL license because I think sharing software is the ultimate goal. So I force people to share my software by using the GPL. That may sound "extreme" (that's the whole open source vs free software debate) but, not being a full time politician, I can't change laws to push society in the direction I want. At my level, that "push" is the GPL choice. Maybe it's not noble enough, maybe it's cowardice, but it's my way (compare that with those who simply don't care).
AI severely weakens both of these. And for people like me, this forces us to reconsider our position. For my part, I accept the legal point of view that A.I. doesn't steal code, and just reproduces the ideas in the code. So, as far as ideas can flow in society, I'm OK with that (that's the principle behind copyright laws).
If AI has its way, one day one will not need to write software, we'll just ask the AI. In that case, software will be dead and free software will die with it too. By then I'll do my "local politics" another way and follow the next RMS.
Then they don't need to train on github, no? Why not release a new model trained from Knuth's Art of Programming, Cormen's Introduction to Algorithms and the C specification.
Feel free to throw in any other published literature related to STEM, but stick to the code samples from the books.
I'm certain it'll be able to change the color of a CSS button, right?
LLM can be trained on a code and at the same time reproduce the core ideas. That's what LLMs do after all - they convert the training data into their own internal models and representations, and then reproduce the ideas.
Sure, some things/patterns, that were repeated multiple times, LLMs will tend to repeat verbatim as well, but that's not that big of a problem.
As a person who invented a few algorithms on my own I absolutely love LLMs and I don't mind them being trained on my work, but yeah - I've been way less likely to publish open source over the last year. In the past, if some of my stuff got traction, the credit was close to automatic (early adopters credited or at least knew where they got it from). Nowadays, LLMs will train on these ideas, rewrite them, and give no credit.
Still, I prefer this to having no LLMs at all.
> but stick to the code samples from the books. > I'm certain it'll be able to change the color of a CSS button, right?
A good enough LLM will just decompile a browser, figure out CSS spec from it, and yes - figure out how to change the color of a CSS button from first principles. There is no point to do this with CSS, but with other things it's now easier to just dig through sorces or direct bytecode than to bother checking docs.
Can it decompile a browser using a specification of x86 and the source code of the compiler?
e.g. without training on the source code and binaries of all software it was able to rip from the internet?
Also, it would be relatively easy to build synthetic datasets for training.
I, for one, don’t mind models being trained on stuff I produced and shared publicly over the last 20 years. I did it for common good, including commercial uses, and this is one of them.
Plenty of people who never produced any open source trying to argue as if if they did.
Just skip the source code. Consider it an easier challenge than reinventing relativity from 19th century physics.
This would be like teaching cooking without looking at any recipes, just from physics and first principles. Or learning music without looking at the sheet music / listening to any existing songs, just generic musical theory and chords. I don't think humans can do it "zero-shot" either...
And this implies that training on the "theoretical" side of things doesn't give you much insight on the "practical" side. Stuff like cyclomatic complexity, UML diagrams and all that stuff might be well-represented in literature but way less so in real software, so training on the literature will produce completely different software than training on production software code.
https://patentimages.storage.googleapis.com/db/8f/cb/dad63e9...
I'm going to stop arguing with you now.
Because they're really really stupid and only make up for this by being really really stupid really really fast.
This has been ruled, by actual courts, to not be "stealing" (not even in the "you wouldn't steal a car, piracy is theft" sense that film and music studios campaigned on).
The last I heard was the "Chinchilla" scaling law was ~20 training tokens per parameter. Humans are, if you'll excuse a very hand-waving Fermi estimate, 100,000 times more data-efficient at learning stuff (it's really hard to tell given we're visual creatures that happen to speak, while LLMs are text-based things that happen to see).
What makes you think that wouldn't work? I think a lot of the hype around AI is vastly overblown but that seems to be well within the scope of what they can be expanded to do in the not too distant future. AlphaGo was trained through self-play reinforcement learning IIRC and I don't really see a reason that some sort of equivalent couldn't be done for generating code starting with textbooks and access to a Linux CLI as a reference. It would be an interesting experiment at least.
I also don't think that just because you have to ask AI to write the software, free software will die. Because people don't understand their own requirements, I don't think we're going to get to a place where AI can one-shot, even moderately complex software, and so creating software will continue to be some effort. I fully expect that the norm will become that we give away free software and expect other people to pick it up and tune it to their own needs with their AI. But that doesn't mean that free software is dead. It means it evolves.
Otoh I’d say the current models are better at predicting expectations than average programmers. Average programmers don’t know ux or business, LLMs do.
But that's literally anathema to the spirit of GPL. Copyleft exists only as a reaction to copyright which is sadly ingrained in legal systems, but the original thought about free software, at the time of GPL inception, is that in an ideal world, copyright shouldn't exist for software ; it's leveraged by GPL only to protect against abuse of copyright holders that could close open code, which is thus made impossible "legally" with the GPL. AI makes this distinction fal into "practically" as pretty much anything is "open" for individual use now (e.g the only use that matters).
If anything, the remark that I have to make to the parent is that in principle attribution is also required by non-copyleft licenses. However, I doubt it's respected for the hundreds of crates or npm modules in a typical Rust or JavaScript project...
Given the supposed commoditization of intelligence and recent memory prices, I'd say LLMs are literally anathema to the spirit of the GPL.
Spirits don't get much legal protection.
Any collective group with open contribution will end up hating it. Why is this a surprise?
Nothing up my sleeve.
This of course requires audit and that requires that the code is written for human consumption, otherwise nobody will bother. Sure you can vibe a printer driver but how sure are you that your LLM didn't include a backdoor in the millions line of slop?
If not that, the ability to trust the reputation of the author which creates an incentive not to willingly insert a backdoor.
Yes, `npm install` was always bullshit because most people didn't bother to check, which is exactly why it was exploited multiple times, which created a conversation about "supply chain security".
If you want to argue that "npm changed the world" then you are correct. It did not change the world for the better though.
So LLMs might decide to add backdoors without prompting?
I look at the whole chain of influence behind them, and it's quite a lot more upsetting than the smiling face I interact with.
The whole "AI is a Lovecraftian tentacle monster wearing a smiley face" thing applies to simple bureaucracies (replace Lovecraft with Kafka), to corporations (replace Lovecraft with IDK, most anarchists?), and governments (Orwell?)
Tools made by tools made by tools, along more steps than most people know even when their job is one of them. Somewhere there's a kid working a dangerous mine without the right safety equipment, elsewhere there's a sweatshop, another place a "reeducation camp".
But who do I see? A cashier, mostly. Someone whose job involves smiling to customers even when we're idiots.
Say I download an app. Who made it? The programmer? Their PM? Apple's store requirements? The US government, for whom there was a special tickbox I had to agree to last time I uploaded an app?
Trust is for the mechanic, driver, builder; but for 90% of my interactions I have to trust my government set good rules and other people followed them.
I can't do this with AI, neither good rules nor them being followed, but I also can't do it when the OS company and app devs are foreign, as they generally are to me now.
That being said there is absolutely nothing wrong with building a portfolio. It is 100% valid motivation. The only issue with that theory is that no one cares about your open source contributions. Not employers, not peers, no one.
What I do find fascinating tho is that fans of the most selfish egoistical companies and tech groups ... insinuate artists, open source programmers, writers or anyone else is selfish or otherwise disapponting if any part of their work involve actual human motivation.
And I'm not deriding those that build open source to boost their own reputation. I'm just much less interested in their work than the work of folks that are motivated by improving the commons (even if they get no credit).
Because, they are a lot of work. They are not hobbies or kitchen soup two hours a week effort. These require a lot of sustained effort and time, that would either takes time from your paid job, or would make you effectively work 60 hours a week which means you have no other life.
End result is that these are paid positions, so.
If a modern-day Beethoven were told that his work will be slurped in by a machine and he will remain unknown to the world, would it be acceptable to him? To you?
It's a perfectly good reason, but I lament the loss of people who are doing open source just to get credit much less than I would lament the loss of people doing open source because they wanted to contribute to a shared commons.
Fame only has utility so long as it gets you paid. If you can get paid without being famous, that's probably a win.
I'm not very concerned about unrecognized genius: they are all around us, and the usual rule is that we've never heard of them. I think the that's just fine.
Fame can have many good reasons. Just feeling the warm glow of adulation is plenty enough for some.
What they do want, however, is to not have to spend time on the thousands of AI generated pull requests and bug reports.
The current point you're making is (I think) that OSS devs don't want credit, they want to be free of spam. I, for example, host my free software on fossil and don't accept contributions at all. It's not that I don't want them, it's that I don't want the noise, so I feel like that's a valid way to distribute open source.
But when I raise that, you're saying I'm moving the goalpost. I think the goalpost is in the same exact location as it has been for the entire discussion: open source authors that don't care about getting credit and want to distribute their free software can do so without any worry about AI.
If your underlying point is that AI has lowered the barrier to entry for code submissions, making previously-viable approaches to collaborative development infeasible, that's completely fair, I just wasn't treating that as fundamental to open source development in the same way you were.