The rise of ChatGPT-enabled GitHub spam
mastodon.social
mastodon.social
Whoever said that generative AI was going to raise the internets noise floor was right on the money.
If you want to really do everyone a favor, consider that TDD has never been easier. Don't write a prompt, write a unit test-- and instruct GPT to write the code that makes it pass.
(Don't ask AI to do your work...ask it to finish your work.)
If you think about it, it does make sense why: the prompt tends to have no where near the context of how the module should be delineated. Normally the programmer knows how the abstraction within his program is separated, but we have no good way to tell it to chatGPT. That means unless your program structure is still following one of the common structure and hasn't morphed into your business domain structure, ChatGPT will mock the wrong layer/ abstraction in the test suite.
The real benefit to my system is that it weeds out the zero-effort clowns by forcing them to validate their own work before I waste my time trying to do anything with it.
For example, I'm in some aviation-related discussion groups, and people will often ask technical questions which are best answered from a manual. In the past, eventually someone with access to the relevant manual would post an excerpt, answering the question.
Nowadays, commonly someone will paste something from ChatGPT, without disclosing that it's just language model output. It will have the form of a correct response, but with invented details, and a few pages of confusion will result while people try to understand the implications of whatever GPT dreamed up, until someone with access to the real manual comes along, after which the original poster will usually admit that they copy-pasted from ChatGPT.
I'm not sure if these people genuinely believe that a language model is somehow a font of knowledge, or if, like GPT, they don't care about whether something is accurate or not, but just care about whether other people believe them.
Whenever a post gets popular enough, odds are good that somebody will ask ChatGPT and post its answer as their own. Usually, it will either choose a popular but wrong book in approximately the correct genre, or just flat-out invent a book by a real author. And then it will spit out a mixture of real plot details, distortions, and falsehoods in order to try to be as convincing as possible.
These answers are very obviously AI-generated (if you're familiar with the book or author in question) and stand in stark contrast to the human-generated answers. It's not uncommon that real people will make guesses that turn out not to be correct, but they almost never drastically misremember books they've read.
I find it incredibly frustrating, because even though ChatGPT could in theory be useful for this kind of thing, it only works if you know to apply the appropriate level of skepticism. When people start posting AI garbage without distinguishing it, they're drastically lowering the signal-to-noise ratio of what is otherwise a very useful resource, all in the name of gaining a few meaningless karma points.
And to forestall the inevitable objection: yes, I know it's possible that there are also lots of correct ChatGPT-generated answers, and I'm just not counting them because they don't draw attention to themselves. But I doubt that's the case, because I've experimented with it myself, and its success rate on all but the easiest questions is extremely low.
Of course the Microsoft issue is raised here. Probably this means if Microsoft ignores it, open source projects are going to have to move off of GitHub on to a platform they control and is explicitly designed to deal with spam contributions.
Everyone's been so scared of intentionally malicious groups, we all forgot about the average lazy fraud who wanted to boost his github commit graph, make his subreddit seem alive, or whatever the dumb case may be.
There will be a learning and adaptation period for everyone
This is also an opportunity for the maintainers to automate their responses and figure out how to chatgptize themselves (at least in terms of answering previously answered issues/questions)
Presumably there will always be a need for a person to answer the novel questions
Semi related news, but Leiningen recently added an anti-plagarism agreement for PRs in attempt to disqualify AI genetated PRs from being accepted. [1] I wonder if this was in response to any GPT spam.
[1] Pentultimate paragraph in the "Contributing" section of https://codeberg.org/leiningen/leiningen/src/branch/main/CON...
Direct-to-post ChatGPT chat is extremely easy to spot right now. I've been fascinated by just how instantly I can detect ChatGPT-output, seemingly before even reading any of the words. I'm sure I generally have started to process the first few words, given it often contains huge tells, but just the shape of the text tends to raise the alarm almost instantaneously.
(Though of course, if the output is high quality and valuable, does it really matter?)
This will work until non-RLHF models become good.
[0] claude-v1.3 — claude-instant-v1 is absolute garbage and I don’t understand why it’s offered at all
And every step of the way companies get their cut of the generation costs.
It does good work but also enables awful practices because of efficiency.
Please don't take what I wrote seriously.
tl;dr they might just do it for the lulz
A more primitive version of the technique was YouTube comment bots that simply re-posted other users comments verbatim - now with LLMs they can easily generate unique white noise posts instead.
I don’t require it, but I definitely look at applicants profiles. Of course, if I see lots of spurious and stupid commits that’s a negative and worse than nothing at all.