AI has definitely made it easier to generate more plausible-looking spam, and removed one way to easily identify bad accounts, but it hasn’t changed the underlying behaviour differences that put a lower bound on how bad the problem can become (as long as GitHub has the team and tools in place to use them in defence).
"Oh boy, another brilliant idea for generating crappy code! Let's just automate our way to technical debt and maintenance hell. Who needs readable, maintainable code anyways? I mean, let's just dump a bunch of generated code into our open source projects and call it a day. Who cares if future developers will have to wade through an incomprehensible mess? I'm sure they'll thank us for our foresight and ingenuity. Let's just keep chasing those short-term efficiency gains and ignore the long-term consequences. What could go wrong?"
;)
https://blog.cloudflare.com/introducing-cryptographic-attest...
Also, ultimately some human behind the scenes is getting prompts into the LLM to produce the content, so any test of personhood simply attests that a human was present to click the button, but not that they wrote what is submitted or even know what it is. Take a look at what happens when CAPTCHAs become commonplace and the cost to solve them with AI becomes too high for spammers [1].
I believe preventing computers from using the internet is a losing game, and improvement of quality resides in:
- In the case of forums, better, manual, human moderation
- In the case of source code, better LLMs that actually produce usable code and/or which could flag which PRs are interesting for review.
- In the realm of news, a cultural shift towards verifying information instead of relying on authority (or the lack of it). This will not happen, but the reality there has been bad much before LLMs.
[0]: https://news.ycombinator.com/item?id=35843566
[1]: https://www.nytimes.com/2010/04/26/technology/26captcha.html
Archived [1]: https://archive.ph/NbkLu
Commit signatures, signoff hierarchies.
Github could solve it with required signatures and automated credibility score.
When using you'd provide "entry points" that you trust (ie. orgs like google, fb, accounts of people you know etc) and "entry points" you blacklist.
You take advantage of directed acyclig graph contributions create (spammer trusting reputable entrypoint is meaningless).
Only single, possibly nested identity has to blacklist something to contribute to downscoring whole spammy contribution graph.
Negative score can be reinforced by others explicitly blacklisting it as well.
I think we'll not get away without anything like that - cryptographic identities, trust graphs etc.
That is an example of a key that wasn't stored in secure hardware. I also believe that hardware security is going to get better over time. There is a clear path forward towards a more secure world of computing and we should not discount it due to mistakes during its early days.
>Take a look at what happens when CAPTCHAs become commonplace and the cost to solve them with AI becomes too high for spammers [1].
Spam mitigations can always be bypassed at a cost. The goal is to find ways to make it expensive and not scale for spammers while keeping it cheap and a good experience for users who aren't spammers.
Since it appears to be a person doing this, any tests of personhood would fail immediately.
Allow sites to do hardware bans. This increases the price to create a spam account. Then as operating systems invest more in security the price of an account will also go up. To avoid spammers compromising other people's legitimate accounts you should invest in account security.
If anything, I would like less tracking, even if it comes at a price of a worse signal to noise ratio on the internet.
Also, we know that this will be used for evil and ultimately companies will start banning "deprecated" or "unsupported" hardware from accessing services, be it their own or from competitors, consolidating even more the oligopolies in the hardware and software industries.
A site could chosen to only ban a computer for day or just give it a bad reputation (require a credit card number or phone to make an account). Or real customers could learn not to buy from shady sources and instead by new, safe devices.
>uniquely-identifying tracking of the user and its hardware, which I don't believe is what's missing in the internet right.
It is missing from the internet and it is part of the reason why spam is a big problem.
>which would make reality the fears of the FSF and co. regarding secure boot and the trusted platform module being used for user control.
User control? It allows services like GitHub to provide a better experience to users by having much less spam. This makes the lives of average users better.
>start banning "deprecated" or "unsupported" hardware from accessing services
It makes sense if you want to have your service be more secure. You could also provide a degraded experience to people with insecure hardware. For example you may not be able to make a PR with insecure hardware, but you could still bookmark a project.
That's already happening with IP reputation and captcha essentially. It's just per-connection rather than per-device.
> Or real customers could learn not to buy from shady sources and instead by new, safe devices.
New safe devices can be compromised. (Paying off someone at bestbuy to install malware would become a very cheap investment) "Shady source" is not a clear category. (See all the scams "sold by Amazon")