NetBSD bans all commits of AI-generated code
mastodon.sdf.org
mastodon.sdf.org
So more of a provenance issue w.r.t licensing than an ideological one.
Will be interesting to see.
Supply chain attack is knocking on the door. Maybe less people should work on OS kernels not more?
If there's a chance the providence of the code is tainted, the only thing that makes sense that doesn't cause the possibility of immense problems later is to reject it.
We also need lawsuits painful enough that companies stop trying to get away with the automated plagiarism, because that's how it's being used, and they know it.
Now, should we [care about getting sued]? Take the number of [copilot users] in the field, A, multiply by the probable rate of [a user getting sued], B, multiply by the average out-of-court settlement, C. A times B times C equals X. If X is less than the [opportunity cost of missing out on user training data], we don't [care].
I'm only aware of these particular projects, but I imagine there are others. Anyone know of any FOSS projects that have explicitly said they are okay with AI-generated code?
Like others have pointed out, it's nearly impossible to police, and to what degree. Boilerplate code is boilerplate code for example, and whether you were assisted in writing it, or you wrote every bytes through your fingers is irrelevant, as it's an obvious continuation.
At this stage, you just have to assume that every publicly available code, regardless of the license, has been gobbled up by LLM training data, maybe multiple times over, and the cat is out of the bag. You might just want to protect your IP by being clear about keeping your hands clear of it all.
https://www.gnu.org/prep/maintain/html_node/Legally-Signific...
In the same light, I feel that LLMs that have been solely and verifiably trained on M̶I̶T̶ MIT-0 or 0BSD should be fine, and this may just be a temporary decision until those become more commonplace.
In general, though, this will be hard to enforce, but I do understand why they took this stance considering the current flock of LLMs and their training data. Drawing a line in the sand on licensing, even if they cannot truly ensure this with absolute certainty, is understandable and feels appropriate for them, especially considering LLMs have now been in common use for coding tasks for multiple years, so this does not seem to be a knee-jerk reaction.
Why did you say number of lines appeared to be part of their concern?
The MIT license requires attribution. You meant MIT-0 maybe?
The old requirements relied on people being honest also.
> You meant MIT-0 maybe?
Absolutely, my mistake.
[0] https://www.gnu.org/prep/maintain/html_node/Legally-Signific...
> code generated by a large language model or similar technology (e.g. ChatGPT, GitHub Copilot)
so I suspect it's actually pretty unambiguous; I've never seen an LLM that was plausibly mistakable for any other technology. (Though if this is merely my ignorance, please suggest an example)
There will be a time in not-so-distant future when people look back at this kind of policy as silly.
With the amount of people getting killed by careless drivers, we should bring those back, at least in cities.
this is fundamentally different. this is a machine, a mathematical equation, that converts enormous amounts of history and facts into a facsimile of creation, thought, expression. it does it convincingly in very small doses, but it doesn't do it all that well. unless your task is to regurgitate something relatively mechanical.
a car was something new. a mechanical horse. it wasn't some distillation and averaging of every extant horse into some "perfect" mutant horse. it was genuinely new. it also wasn't a mirage or a hallucination.
AI is not a creative force, rather reductive. it's at best a tool. so if you want to copyright something or you want to contribute to NetBSD, sure, that's fine. but copy-paste of someone else? no matter what the underlying process, it's still wrong.
it's not Luddite behavior for saying "no, we don't want to do that"
Also HN: AI is evil because it's stealing copyrighted works!
also me: AI is evil because it tricks people into giving up their creativity, which isn't the boring part, it's the real part
also also me: wtf? why bring that up?
The factory owners said, "lol, nah" and fired them all to hire cheaper labor to run the autolooms so they could keep the increase in profitability to themselves.
The Luddites were rightfully pissed and wrecked shit. The factory owners, had they been less greedy, could have kept the Luddites on board.
Only if copyright laws are abolished, or stop being enforced completely.
[0] https://urbanaccessregulations.eu/countries-mainmenu-147/swi...
[You could argue that it is bad because the pattern may change, but not having to type it out allows me to think about that instead of the typing -- and a super common bug for me to write is where I copypasta something X times and forget to fix a detail in one of them]
The stated reason is licensing, but after seeing some of the crap code that AI-crutch-coders have tried to pass as worthy in some other projects I've seen, it's good to see more pushback against drowning in mediocrity.
I can understand why they have done this, though.
I mean, any question you ask is going to be logged as well as the output/reply. I do wonder if there are some legalities to this decision not just the quality of code being (simply) copied and pasted into projects.
For example, lets say I create a game and its 90% code from chatgpt, and I sell the game on steam for some pennies. Could the company behind chatgpt claim some kind of ownership... even if I modified the generated code most of the time? Afterall, they can see the questions I have asked, and the results generated. I am sure they could figure out the projects its for and investigate, right?
Obviously free/opensource projects is a little different... but there could be other legal factors unless they specify "we do not accept ai generated code" or similar.
It is certainly something I would put in place for my own public projects.
I wouldn't want to ever work without Copilot again. So frustrating to not have my magic autocomplete when I have to make some changes in a basic editor.
You should get committer flag in NetBSD because of how LITTLE new work you bring to the table, and because of how MUCH code you REMOVE.
Most of the time it just feels like an autocomplete that is actually useful, it's scary good at writing the rest of the line exactly the way I was going to.
Anyway, your argument seems much more reasonable than the original post - they're just faffing about copyright which is completely pointless since Microsoft already declared they'll take care of any copyright lawsuits related to Copilot.
And a lot of us have been saying these LLM's are nothing but advanced autocompletes... and maybe that isn't a bad thing?
How very kind of them!
edit: it's almost like you, as a user, get a taste of having so much money youre legally untouchable!
Maybe we just get more of the same errors in more code because they all copy from the same faulty AI.
Yes, copilots are really helpful, even if you do know what you're doing.
Or just wait to get left behind, while you sit on your ivory pedestal of superiority. Your choice.
I’m not tethering myself to a distant datacenter just to be able to write some code.
I'm not saying copilots aren't useful. I'm saying that this "you'll be left behind" rhetoric smacks of crypto esque hype.
"while you sit on your ivory pedestal of superiority. Your choice."
ha! if only! i actually sit in a dim room with remnants of caffeine and pizza and wonderful half-broken computers and brand new ideas and code and interesting challenges and discoveries. and friends.
i did choose.
correctly.
It's like all the kids who use it to write their papers. Great, you didn't get caught, your instructor agrees that you are the worlds most average and uninspired writer.
(and for boilerplate there isn't any fundamental difference between copy-pasting it from ChatGPT or anywhere else copyright wise since it's practically identical and unverifiable)
- statistically, a below average programmer, or - a really, really, slow typer.
Writing code is the easy part. Copilot, and transformers in general, are only able to generate "statistically average ish" output. A statistically average ish coder is also pretty bad at coding (this stands to reason, or else programmers wouldn't complain about how bad every other programmer's code is). So if Copilot makes you more productive, that's a bit of a red flag.
If, on the other hand, you're a really slow typer, then unfortunately you're out of luck, because nobody has figured out a superior input method to pressing buttons. I think in this case, Copilot is an OK solution, but you'll find yourself correcting most of the code it writes anyways.