RFC: Banning "AI"-backed (LLM/GPT/whatever) contributions to Gentoo
mail-archive.com
mail-archive.com
Genuinely in disbelief what I was reading in that thread. How do people think that an autogenerated nonsense description is better than not describing the software at all?
[still gets $30M in VC funding because of course it does]
What's worse is that they have the condensed information and decide to make it harder to consume.
Banning developers from internally using Copilot, a personal tool that is essentially a glorified autocomplete so you don't have to type as much, all because of "copyright infringement," "ethics" and "energy waste" concerns, is dumb. It is unrelated to the problem at hand, unenforceable, bizarrely overreaching, unnecessarily divisive and also dumb and I hate it.
> Haha, I'm glad you're finding this interaction amusing!
A central reason was "reputation". It's a PR move because ordinary people are coming to associate AI with really bad things. If they see us using AI for icons, why would they not assume the rest of the content is AI generated too? It doesn't matter whether the general public are right or wrong.
I think what we're hearing in this thread is mostly a developer viewpoint. Try to see it like the customers see it.
Of course using AI makes many tasks easier. Changing tack is a schlep because I have to go back an replace the old instances. Honestly I'm almost regretting the decision and shudder to think what this accumulated debt would be like with a couple million lines of code and docs instead of a dozen pictures.
Seems very courageous of them to make a policy choice early.
This was raised in that linked thread and is a completely valid concern. Who could look at that package manager and feel confident about its other design decisions even if it wasn't largely AI generated itself?
As an open source developer, I have complex feelings about users who act like they're "customers," entitled to make demands of me. (Unless they are, in fact, paying customers. Paying customers are great, but they're often not the goal of most of my projects.)
The stuff I write on my own time, I write first and foremost for myself. I'm happy to share it with other people, but that's a gift. Sometimes sharing will get a little scene going, and we can all have fun together. But I certainly don't spend $400+/year to "code sign" my binaries for MacOS and Windows. 90% of time, if I hear from a Windows user, it's because they're upset with me and they want something. And I will often deliberately avoid advertising certain tools, because I have zero desire to spend my life providing free tech support.
I'd rather have 100 users who were into what I was trying to do, than 10,000 unpaying "customers" where I had to worry about things like "It doesn't matter whether the general public are right or wrong." I do that at work, not for free on the weekend.
So if I want to enable CoPilot's auto-complete when designing a new file format, then eh, I really don't care if this makes some user suddenly upset and unwilling to use my project. I'm happy to label things clearly so that they know what they're getting and they can make an informed decision. But they're not my customers unless they paid me money.
The idea that intelligences - whether they be human, artificial or alien - should be forbidden from learning from code freely shared on the internet goes against everything I like about open source.
I think it's fair that no one should be able to use reproduced copyright code verbatim, whether that by by a human memorizing something or a computer copying it.
But I take the complete opposite view on the ethics of letting machine learn from work. I think this should be encouraged.
Love it or lump it, this has been the battleground for most free software distribution. Just having a licensed copy of code does not permit you to do anything with it, even the BSD license has restrictive terms for the user. I'm an enormous Copyleft advocate, but arguably Open Source is only enforced by stopping people from using it illegally. If ChatGPT's fate is to turn into an GPL-licensed code-launderer, then projects have a great basis for banning it among their contributors.
> I think it's fair that no one should be able to use reproduced copyright code verbatim
Then I don't see how you can be upset at Open Source projects for adopting basic standards. They also want to protect their own license and community, with the main difference being that they're not in it for the money. Again - there was never any point where "Open Source" was synonymous with "do whatever you want with the code unconditionally".
There exist copyrighted algorithms for certain applications, especially for AI application these days, but a reengineering should be possible and similarities should be handled as liberal as possible.
Therein lies the rub. Most of the discourse and debate around LLMs and copyright revolve around the central question of what it means to learn.
Virtually everyone agrees that a human learning from reading code doesn't violate copyright (by somehow copying the knowledge into one's brain), because the human brain is some kind of copyright laundering machine, maybe? I don't really know whether there is any argument for that apart that can't be reduced to an appeal to common sense.
On the other hand, a DL algorithm learning from processing tokens of code scraped from public sources such as GitHub doesn't present the same kind of obviousness. My personal belief is that it's also learning and shouldn't be forbidden, but I can't deny the negative consequences of that. We're already seeing a lot of bad things come out of the democratization of GPTs.
I don't think this is so universally agreed upon as you think, or at least the implications of learning from copyrighted material. This is why projects like ReactOS and WINE strongly prohibit contributors from reading leaked Windows source code, incase they learn a little bit too much and accidentally reproduce copyrighted material.
> because the human brain is some kind of copyright laundering machine, maybe
Absolutely not. Music is full of legal cases where someone learned and copy a bit too much and went to court over it.
https://en.wikipedia.org/wiki/Pharrell_Williams_v._Bridgepor...
https://abcnews.go.com/Entertainment/jury-reaches-verdict-ed...
https://en.wikipedia.org/wiki/List_of_songs_subject_to_plagi...
This isn't really agreed on at all. See https://en.wikipedia.org/wiki/Clean_room_design which wouldn't exist if it was agreed on.
However, note the case law example of the NEC V20 which found in NEC's favor:
> While NEC themselves did not follow a strict clean room approach in the development of their clone's microcode, during the trial, they hired an independent contractor who was only given access to specifications but ended up writing code that had certain similarities to both NEC's and Intel's code. From this evidence, the judge concluded that similarity in certain routines was a matter of functional constraints resulting from the compatibility requirements, and thus were likely free of a creative element.
I once came up with the idea of a physical copyright laundering machine. It had three CD-R drives (this shows how long ago I had the idea). You’d insert a CD-ROM to launder, and two blank CD-R discs. To one of the CD-Rs, it would write a one-time pad, to the other it would write the input CD-ROM XOR the one-time pad. A hardware RNG (I wanted to use a quantum process such as radioactive decay for more emphatic indeterminism) generates the one-time pad. It also generates a single random bit which determines which output CD-R gets the key and which one gets the ciphertext. That bit is never revealed to the user (or recorded in any way). The end result is two CD-Rs, one containing random data, the other a copyrighted work encrypted with random data-but it is impossible to know which is which.
I never actually built one of these machines. I wanted to patent it, but gave up when I realised how much patents cost. I also eventually realised that my machine would never work, because it was approaching the law with the mind of a developer not the mind of a judge - I doubt any judge would actually be convinced by my copyright laundering machine, they’d find a way to rule against it, whatever exact way that might be. The law and computing are both systems of rules, but the rules in the former involve far more discretion and flexible interpretation.
Because it's an end not a means. The concept is central in philosophy, law, ethics, education.
Maybe all we can hope for is consolation. Like, if the tech elite is successful in this campaign of soft-dehumanization ("we are all LLMs anyway") it will open up some new cultural pathways to greater respect for animals and the environment.
What is a cow but some kind of copyright laundering machine anyway?
Subjectivity as we know it is quite contemporary, articulated in part by things like Kant's kingdom. I think things will start to change again. The world of love, art, particularly human passions might fade away to something different. You can almost feel people hungry for this. It doesn't have to be good or bad, its not for our subjectivity to understand after all. But I will say it does feel like we are leaving summertime now, towards a colder future.
I think it's inevitably bad because dehumanisation ends in war. Dehumanisation, for me, is synonymous with violence. It isn't just technology but the terrifying post-modern subjectivity that allows 8 billion people to co-exist in relative peace. "Accelerationism" looks set to build still more and better weapons, fewer ways of resisting using them, and less capacity to care, so I fear the colder future you speak of will be a nuclear winter. Of course, as you say, it's like some people want that.
Or humans as pets. If you want to see humanity’s future, consult the nearest dog or cat.
- 7 instances of the word "shit". I don't mind swearing, but it is indicative of the author being maybe a bit too emotional for a technical proposal.
- It is unnecessarily broad. Not using AI to create bug reports? What if you use AI tech to find a bug? Are you not allowed to report it? The stated issue here seems to be mostly about code completion, but it is stretched to everything AI-related, everywhere.
The point raised are copyright, quality and ethics, which are valid points, but not specific to AI.
Copyright: You have the same problem when copy-pasting code, and people do that, you can't really single out AI. Instead of banning AI, a more sensible guideline would be to just be aware of copyright when importing code from elsewhere, including AI generated code, but also copy/pasting from online sources (ex: StackOverflow) and using external libraries. There are tools to check for copyright compliance.
Quality: AI-generated code is often lower quality, but so is code written by bad coders, judge by quality of contribution, not by how it is done. As for the "we can't really rely on all our contributors being aware of the risks", maybe start by picking contributors you can rely on. And if you think they may not be aware of the risks, tell them about the risks rather than saying "you can't do that".
Ethics: I don't know what Gentoo stands for, but I'm guessing it is mostly about making a good source-based Linux distribution. Don't hijack the project for your own goals. Now, I have no problem with a Linux distribution that has "no AI" as one of its core values, but it doesn't have to be Gentoo.
I mean, the hot new trend seems to be people reporting completely imaginary bugs that some AI tool thought it saw to projects, so, eh, I can see where they're coming from there.
"The goal of Gentoo is to design tools and systems that allow a user to do that work as pleasantly and efficiently as possible, as they see fit." https://www.gentoo.org/get-started/philosophy/
And this is what Torvalds had to say about LLM-enhanced submissions to the kernel. https://www.youtube.com/watch?v=w7-gJicosyA
AI can be used poorly, AI can be used well.
Gentoo is right here, and until we pass these hurdles, I don't use any of these systems, even with a 100 feet pole.
2) The specific issue you're talking about is because they don't see letters, they see tokens, which are groups of letters / subwords. It can't count those because it can't actually "see" what it's counting. This is being worked on as well.
"Banning LLM content" is in my opinion an effort spent on the wrong thing. If you want to ensure the quality of the code, you should focus on ensuring the code review and merge process is more thorough in filtering out subpar contributions effectively, instead of wasting time on trying to enforce unenforceable policies. They only give a false sense of trust and security. Would "[x] I solemnly swear I didn't use AI" checkbox give anything more than a false sense of security? Cheaters gonna cheat, and trusting them would be naive, politely said...
Spam... yeah, that is a valid concern, but it's also something that should be solved on organizational level.
I get the impression he agrees this road to LLM content is inevitable, but also kind of emphasises the role of the reviewer who takes the final decision.
Llm are absolutely helping with catching buts and code quality already.
Progress to where? One should not use "progress" as an unqualified noun to denote a scalar. Progress is a vector, with both magnitude and direction. The direction part is really important.
Even when I wrote very single line of code myself, I use AI to ask it about questions regarding the programming language or the library that I use. Banning that is just handicapping yourself.
I do like the sentiment. You absolutely do not want people to commit code they don't understand themselves, but the solution isn't to outright ban AI. The solution is to have trusted, knowledgeable developers who are aware of the limits of AI and use it appropriately.
Heck, even when I come up with a solution myself without needing to research it, I'm still likely to get it wrong or make mistakes. That's why we rely on tests, specs, code review, etc, because humands still make mistakes.
it's kind of annoying to ferret out the bugs it has carefully concealed in the code it wrote for you, though
Shouldn't that be enough in the first place? RTFM is a real thing.
but it's pretty common for gpt-4 to be able to instantly find the thing you need out of 150 pages or 1500 pages or 15000 pages. or the five things you need, and how to combine them. if you want to know what libraries can probably solve your problem, asking gpt-4 is a much better idea than reading the manuals of all possible libraries
The negative view: write plausible-seeming explanations justifying the code as correct.
additionally, similar to how large PRs are more likely to just be skimmed and replied with a "LGTM!", an LLM missing some bad stuff but still producing a seemingly thorough review would increase the chance of the bad stuff making its way in.
allowing LLMs to write code would be fine if its truly verified by a human, but let another LLM hallucinate and cloud a persons judgement and you've got a problem
"Please review the above code. How does it work? Is it well designed? Is it efficient? What are its good points and its bad points? How should it be improved? Is it readable and maintainable?"
i feel like gpt-4's code review (included below) was mostly correct and useful. however, the efficiency concerns in particular are unfounded, and the python approach to handling errors like those cited is to just let the exception propagate, suboptimal though that is for usability. also, i'm not sure i agree with its high opinion of the modularity, usability, and readability
simply pasting gpt-4's partly incorrect analysis into a textarea when someone asks you for a code review would be malpractice, but it could be a useful source of ideas for possible improvements to suggest and possible problems to investigate. also, i've found it to be especially good at coming up with plausible ways that someone could misunderstand my code, which is to say, ways that my code is unclear and therefore could have its clarity improved
overall i'd say that gpt-4 is a lot better at reviewing code than at writing code. i've gotten significant improvements out of its code by asking it to review its own code looking for improvements, then asking it what the code would look like with those improvements. also, having it write test suites, then showing it the results of running the test suites
— ⁂ —
The provided Python script is designed to generate and print arithmetic tables for balanced nonary (base 9) numbers using a unique set of Unicode glyphs. The script covers operations such as multiplication, maximum, addition, and subtraction. Here's a breakdown of how the code works and an evaluation of its design:
### How It Works:
1. `gen_table` Function:
- This function generates a table for a given binary operation (`op`) like multiplication or addition.
- It takes an optional `name` parameter for labeling the operation in the table header.
- The `values` parameter specifies the range of values to be used in generating the table (default is from -4 to 4).
- It constructs each row by applying the operation between all pairs of values.
2. `nonary` Function:
- Converts an integer into its balanced nonary representation using specified glyphs.
- Handles negative numbers and numbers greater than 8 by recursive calls.
3. `layout` Function:
- Formats the table into aligned columns for better readability.
4. `paste` Function:
- Combines multiple tables horizontally with spaces between them.
5. Main Execution Block (`if __name__ == '__main__'`):
- Generates specific tables for different operations and prints them formatted properly.
### Evaluation:
Good Points:
- Modularity: Functions like `gen_table`, `nonary`, and `layout` are modular, making specific tasks well encapsulated.
- Reusability: The functions can be reused or extended easily; e.g., adding new operations or changing glyph sets.
- Readability: Despite some complexity in logic, especially in recursion and formatting, variable names and function names are descriptive which aids understanding.
Areas for Improvement:
1. Efficiency Concerns:
- Recursion in `nonary`: For large numbers, this could lead to deep recursion levels.
- Repeated calculation of widths in `layout`: This could be optimized if performance becomes an issue with larger data sets.
2. Error Handling:
- There’s no error handling if invalid inputs are passed to functions like non-existent glyphs or unsupported operations.
3. Documentation & Comments:
- While there's a detailed module-level docstring explaining high-level concepts, inline comments explaining complex parts (like recursion in `nonary`) would improve maintainability.
- More detailed docstrings on what each function expects as input and what it outputs would help other developers understand their interfaces quickly without diving into code details.
4. Testing & Validation:
- Adding unit tests would ensure that changes do not break existing functionality and help validate edge cases like very large/small numbers or unusual sequences of operations.
5. Scalability & Generalization:
– Currently tailored specifically to balanced nonary system; generalizing this to handle any base might make it more useful as a tool for exploring other numeral systems.
6. User Interface Improvements:
– Providing command-line arguments or interactive prompts could allow users to specify ranges or select specific tables without modifying source code directly.
### Conclusion:
The script is well-designed concerning modularity and separation of concerns but lacks robustness due to missing error handling and testing components. Enhancements towards documentation, user interaction, scalability/generalization could significantly increase its utility while maintaining ease-of-use for educational purposes or numerical system explorations.
I am grievously concerned about the sea temperature being off the charts for a year now and we do not know why. Wasting tremendous amounts of energy to generate bullshit doesn't seem like a prudent action. Yes it's not yet a large chunk of total energy usage but we do need to stop the hype before going further when it'll be.
Because as https://hachyderm.io/@inthehands/112006855076082650 explains it's good for naught else.
> You might be surprised to learn that I actually think LLMs have the potential to be not only fun but genuinely useful. “Show me some bullshit that would be typical in this context” can be a genuinely helpful question to have answered, in code and in natural language — for brainstorming, for seeing common conventions in an unfamiliar context, for having something crappy to react to.
> Alas, that does not remotely resemble how people are pitching this technology.
And then of course there are all the ethical concerns.
The demand for technology leads to advancements that meet our needs. As we continue to innovate, we must focus on consuming more energy rather than less.
You are eager to decide what is useful and what is not. Can you predict the future? Can you predict the full impact of technologies? Can you see second, third and forth order effect? Likely not. For instance, many may not have anticipated the significant role smartphones play today.
It concerns me when some individuals attempt to control others' resource usage, potentially leading to authoritarian rule driven by fear. Such actions might result in adverse effects before any noticeable climate changes occur in the near future.
Also, if you need to convince people to live in such circumstances then a little convenience goes rather far so that also needs to be considered.
Modern washing machines are certainly more water efficient than hand washing and I wouldn't be surprised, again, if they would be more energy efficient too. Once again: humans consume energy too. Edit: and as someone else noted, we look should look at the societal effect. Well, it's quite clear the washing machine is an extremely big plus as it automates a time consuming, hard physical task.
So far every order effect of LLMs are terrible as they are built on the backs of exploited workers and are used to further disenfranchise workers and also artists.
> As we continue to innovate, we must focus on consuming more energy rather than less.
I think in your fervor to put me down with a flippant comment you went too far. You know this is patently untrue, aren't you? I mean, since you mentioned washing machines surely you are aware both the United States and the EU are pushing hard for more efficient washing machines? https://energy-efficient-products.ec.europa.eu/ecodesign-and... https://environmentamerica.org/center/media-center/biden-adm...
As for lights, LED lights consume less energy and are safer than the old incandescent bulbs. That's once again progressing towards less energy.
Sure we consume more energy than we did pre-industrial revolution but that doesn't mean we must continue upwards.
The things you listed are "wants." Perhaps we could say that washing machines has turned into a need, in much the same way crude oil has. What would the world have been had we tamed nuclear power before oil was commercialized?
> You are eager to decide what is useful and what is not
I think GP is eager to decide what is net beneficial, which is a tradeoff between usefulness and cost (monetary, social, environmental.)
I don't personally care that much about Earth. It's a rock in space, and it will continue existing with our without a functioning ecosystem, but I try to be conservative with my actions, so that the people who do care, may continue enjoying it.
At this moment, AI is a "want," not a "need."
Or does the desire for technology leads to advancements that meet our wants? Wants versus needs, and desires versus demands get confusing sometimes.
> It concerns me when some individuals attempt to control others' resource usage,
From a psychological point of view that's understandable. And it does portend an ugly authoritarianism. From a realist standing, it's inevitable if resources are limited. Right now we seem to be limited by the heat capacity of the planetary ecosphere. To avoid that becoming an open conflict I think we need to enrich the debate to talk about appetites rather than needs.
A huge majority of tech advancements are driven by supply rather than demand. Capitalism and modern economics pushes companies to build whatever they can market and sell, they aren't designing new products and tech because consumers already asked for it.
The hypothesis I've seen that makes the most sense is that enforcement of ultra low sulfur diesel for marine shipping caused the sea temperature rise [1]. I don't remember which podcast it was now or I would link it, but I first heard this a couple months ago from a researcher who is concerned that this quick rise in sea levels combined with the start of an El Niño this year will lead to a very hot summer with more unpredictable weather.
Anecdotally, weather predictions in my area (south eastern US) have been absolute trash. My best guess is that prediction models are way off now based on changes in underlying conditions the models never had to account for.
[1] https://www.carbonbrief.org/analysis-how-low-sulphur-shippin...
Personally, I think LLM's have a place but need supervision for the foreseeable future. In case there are diminishing returns on the horizon, we might need a new AI breakthrough to go further than what we have. But I am absolutely helped in my job by AI assistance. Key is however to understand the system well enough to realize when something is off. Engineers won't be wholly replaced by AI anytime soon, but many can already be helped by them.
Let's take for example the point about code laundering by LLMs and their un-traceability. Clean room reimplementation is widely considered ethical in the software world (for example Wine is built this way), but thinking of it in this context, how is it not laundering to avoid lawsuits, technically legal but ethically dubious? Where's the line? How can writing the boilerplate and generic scaffolding be considered laundering? Certainly there are gradations, and a blanket ban is myopic? This all has been discussed in that thread.
What about writing Autowiki-style documentation, filling the gaps nobody wants to work with? Machine-assisted translation? Indirect use of the tools?
Another point, quality. The person committing the code should be responsible for it, LLM-assisted or not. And the reviewer is responsible for verifying that it meets the quality bar before accepting it. It's as simple as that. (although it's addressed down the thread by pointing out the concerns are redundant).
If you wanted a copilot style thing that doesn't print out GPL'd software, you wouldn't feed it a load of GPL'd software for it to remember in the first place. What we have is like doing a clean room reimplementation by loading the original in your text editor before you start and deleting parts of it at random.
Pretty soon anyone looking to add "open source contributor" to their GH profile can take a Gentoo issue and ask an AI to cook up a solution, put that on a PR, and send it in.
This will be a nightmare for maintainers. I'm not sure if there is a solution, since AI usage will spread regardless of how good/accurate it is and there's no way for us to differentiate between plausible bullshit and actual contributions, without reading it carefully. Reputation of contributors is probably the best proxy for genuine contributions, but that's a catch 22, so it can't be the only way.
I believe I've noticed only one 'LLM spam' comment on an issue needlessly comparing different Javascript package mangers.
Pretty ironic coming from a distribution that requires every user to compile everything from source.
No. There are not practically enough cycles saved to overcome the often minutes of CPU time to build every update of a large software package.
> When every program is compiled for a specific architecture
Every binary distribution already "compiles" for a specific architecture.
The inane variant "tuning" generally renders miniscule speedups. Even if there is performance to be gained it can be done on a limited case by case basis.
I’ve seen folks use ChatGPT to generate code and review it for security flaws. It does often solve the tasks. And it leaves behind many kinds of vulnerabilities: injections, overruns, etc.
Based on the little empirical evidence we have about informal code review [0], it seems that we ought to limit or outright ban generated code. A Human can only read so much code before their impact on catching errors significantly drops. OSS project maintainers have enough on their plate and we don’t need to exhaust them with trying to maintain AI generated code.
[0] https://sail.cs.queensu.ca/data/pdfs/EMSE_AnEmpiricalStudyOf...
Update spelling
What mattered was eliminating the time I had that morning, to do meaningful work. Responding to that (and since, several) PRs was not the desired outcome. It also kicked off internal, processes - which in turn added more time. If you think about it, it's an interesting hack at least!
For a large highly visible open source project, there's also probably intellectual property concerns. LLMs are being trained wholesale on code that's under IP protections (be that copyright or otherwise) and having that end up in your code base could land you in trouble. 'The AI did it' will probably not be seen as a viable defense.
Ask the LLM to generate some unit tests for the code it just generated. Then if the test fails, ask the LLM to fix the bug in the code.
I find that this catches and fixes a surprising number of LLM bugs in generated code.
If an LLM is intended to generate code, perhaps it's standing orders should include "Generate a comprehensive test suite, run it, fix the bugs, then iterate. Send me an email when you've finished."
Why does it need three companies working on it? Isn't this just a matter of prompt engineering?
I've never played with an LLM, so I really don't know. Perhaps it requires some ordinary linear code to produce a "comprehensive test suite".
Ultimately, the problem I can't see 3 or 300 companies solving, is the correct interpretation of technical instructions in English. Native English speakers have trouble with that, and I doubt that a machine can out-perform a native speaker. Maybe I'm wrong, and it can; but it needs to also be able to convince me that my doubt is misplaced.
Writing technical specs is like writing laws; you're using vague words to describe something precise. We don't use machines to interpret laws, and I don't trust machines to interpret specs.
And if you don't, and therefore cannot quantify, then for future reference simply state "it feels like the LLM makes me more efficient"
And they will thank me for it! I've just saved them 3h45m.
This ratio stopped maintainers getting buried in review requests etc.
With the rise of LLMs there's a risk the ratio will get flipped: someone can make a patch in 5 minutes which needs 15 minutes of review.
That could make being a maintainer much more time-consuming - and mean the job is much less about maintaining code, and much more about dealing with timewasters politely, which ain't many people's idea of a fun hobby.
That should be the way, at least, until AI is proven to better at writing code than humans.
With the XZ utils backdoor, people are currently removing any contribution done by that attacker. Even if the work done is of limited quantity, evaluating every character and every byte is a huge endeavor requiring large amount of man hours. At some point it often is easier to discard code then doing the necessary in-depth analyses and testing in order to determine with a 100% certainty that the code do what the code should be doing with no side effect, security issue or legal implication.
If you think I’m an AI, you have to review it so closely you right as well just have written it yourself.
At scale it doesn’t work to simply judge the content on its merit when the content can be generated at no cost. The content becomes indistinguishable from spam, no matter the intent of the contributor.
...until you count how many fingers the humans have.
Same is true for AI generated code.
- Have been written with CoPilot enabled in my editor, and
- which optionally use GPT 3.5 as a translation API, and
- Which use OpenAI's text-to-speech model to generate spoken dialog files for testing.
I suppose I can try to mark my projects in a such a way as to inform Gentoo that it's against their policy to package them.
Overall, I would guess that my CoPilot-assisted code is slightly worse than code I hand-craft. The biggest difference seems to be that with CoPilot I write fewer tiny functions, and I tend to keep more related code in one place. On the other hand, CoPilot makes writing test code extremely quick. And I'm not talking about generic boilerplate here: CoPilot can write non-trivial parser or type inference code that relies heavily on internal project APIs that do not exist outside my project.
Overall, I'd guess that CoPilot allows me to produce twice as much code at 90-95% of the quality. Which since we're talking about open source projects that I maintain in my spare time (and that were painfully over-engineered to begin with), is probably a decent tradeoff.
> We can't do much about upstream projects using it.
As a potential upstream, I could potentially help them by adding some kind of metadata to my package, indicating, "Some portions of this code were written with CoPilot active." And this could allow them to automatically filter out and reject packaging requests from users.
(As an open source author, I'm deeply ambivalent about distro packaging anyways. I release my software as pre-built, standalone binaries specifically to avoid the tarpit of distro packaging politics. If my software is packaged for a distro, it will almost always need to go through someone else, who may or may not do a good job, or keep the software up to date, or break the software in a way that creates more support requests for me.)
This is tangential, but as a user I vastly prefer distributions, because I can rely on stuff working in context, on observing the distro conventions, and on automatically receiving security updates. It is a much more pleasant and convenient mode of software sourcing.
Maybe a naive question but, how will they know?
There's no implication in what I wrote that this is the case. The fact is that weeding out the crap is a tedious and difficult job, and maintainers (especially volunteer ones) don't deserve the extra burden brought by automated crap generation.
> It will push out good contributors who use AI responsibly...
My entire comment, which is just one sentence long, explains why this will not be the case.
>... and presumably make zero effect on people just wanting to abuse AI to get a patch into Gentoo.
That's not everybody (or if it is, Gentoo is screwed anyway, which does not mean it's not worth trying to stop it happening.)
These points have been made by multiple people in these comments over the last few hours.
When it is discovered that your submission is AI generated, it is enough reason to discard the submission without having to review it any further.
Many open communities do the same.
If not depending on AI tools then depending on a... hunch? So like a modern-era witch hunting?
There's a bit of a "intro paragraph, five numbered list items, concluding paragraph" style format that you start to notice pretty quickly.
Then a bot responded to the discussion poster's concerns and it was humorous but also it offered no way to resolve the issue.
So there are one or two cases where a maintainer might notice something off and this policy would offer a clear-cut way to reject whatever submitted the inaccurate decisions or to take the AI out of the discussion forum.
But for the cases of copilot-authored code I don't think there's any reliable way to detect or reject it. This probably falls under their "but not upstream changes" caveat.
I don't personally agree this particular line in the sand will help in all cases -- it is a difficult standard determining whether something is AI-created, this will likely increase the burden on the humans in the loop. But as policies go, it makes sense to have a line drawn in the sand for outright rejecting it on source not content, especially in the context of a package manager and Linux distribution. The burden on said humans in the loop will be even greater if they don't have a rule in place granting blanket dismissal on this characteristic, especially if they're correct in seeing an increase in AI-produced packaging of unknown binaries.
Of course, it'll get harder and harder to spot the problems, but that just brings the "bug" closer to the human-generated level.
Banning AI doesn't fix the problem, especially since the type of person that would have AI generate a description and then not even read it also isn't going to follow the rules.
> The type of person that would have AI generate a description and then not even read it also isn't going to follow the rules.
This is not a sound argument against this rule; it is an argument for the proposition that current LLMs present a threat to the open-source model, despite this rule.
Sure they can, but the descriptions are completely irrelevant to what the package or project at hand does.
What’s the fundamental difference between that and a human doing it without an AI’s help at all? It could be essentially the same outcome, just more time and human effort
The code is usually equally braindead. Not that it matters of course, most code that humans write today isn't much better, but if you value quality like many foundational open source projects do, the difference becomes obvious.
I can understand the sentiment behind this proposal, but it is way too nuanced and complicated to just solve it with a few basic rules.
The RFC specifically mentions tools like Bard, ChatGPT and Co-Pilot as what it considers to be ban-worthy.
If I’m going to be extremely pedantic about it: with Co-Pilot, the a actual contents of the contribution are AI generated. With a translator like DeepL, the contents are authored by the contributor, the translation tool just translations what’s already there
What are you referring to?
I agree that “AI” needs to be more closely defined, as for example using A* search is also AI.
On the one hand, I’m on the fence about this heavy-handed approach. Tons of people, myself included, use AI assistants to create high quality work in less time. Of course I’m also aware of tons of low quality garbage.
On the other hand, I’m all for banning automated submissions which have been on the rise for the past couple of years, which are often thinly veiled (if at all) ads for AI startups. GitHub in particular should allow owners to ban all unsanctioned bots, and report unlabeled bots.
The inference energy cost is likely on the same order of magnitude as the computer + screen used to read the answer (higher wattage, but much shorter time to generate the response than to formulate the request and read it).
The training energy cost is significant only if we ignore that it is used by many people. For GPT-3, I've seen plausible estimates of ~1 GWh, which would equal to about 400 tons of CO2, about as much as a single long-distance plane (total, not per passenger, fuel consumption only) round trip. Estimates for newer models usually ignore the existence and likely use of more efficient accelerators.
Water consumption skyrocketed while training up the current generation of models [0].
The energy (and water) usage to service requests is not staggering but it is concerning that is uses as much as it does for the output it generates [1]
[0] https://futurism.com/critics-microsoft-water-train-ai-drough...
[1] https://www.forbes.com/sites/cindygordon/2024/03/12/chatgpt-....
If you acknowledge up front that AI is unfit for purpose and is very likely to introduce some serious security problems then it seems wise. When it turns out the LLM models have all been compromised to insert backdoors into the compiler toolchain, you win by being the last distro left standing. You could look at it as a very high risk strategy for the same reason, if you think you'll be "left behind". Either way who dares wins (or dies). Dare to go against the mob, or dare to bet the farm on a principle.
AFAIK Gentoo is one of the more conservative communities. But I'd also expect to see this policy being considered in BSD circles too.
I've used AI to learn about massive codebases. It's a bit stupid but still extremely helpful. The free ChatGPT was capable of explaining the concepts in the code and the file system structure of the project, allowing me to get started much faster. It sure as hell beats being a help vampire on some IRC channel or mailing list.
This technology is literally too good to be banned. We should be working on taking it as far as humanly possible by getting it running locally and completely uncensored.
It is impossible to use most modern devtools completely without AI. I think it is better to regulate the usage and enforce transparency.
There are multiple repliers that very clearly don’t understand how LLMs work. “It’s computers, so I can intuit it!” is typical techie hubris.
Like - it suggests the next 10 characters or so it's either correct or correct enough that it saves me some typing.
So in the same way that people complain about CGI in movies (if the CGI is good CGI then they probably haven't even noticed it) - the only AI you notice will be the bad stuff.
People are producing plausible sounding bullshit because of ease of just cranking out and iterating code quickly. And we have made it way too easy to incorporate potentially copyrighted other people’s code.
As Donald Knuth would know, back when you had to generate punch cards and stand in line to load them on the mainframe, you spent a lot more time carefully designing and logically working through your code, instead of producing massive amounts of plausible bullshit. So yes, I agree.
You are talking about banning interactive editors with copy/paste and interactive compilers and debuggers right?