Proof-of-work to protect lore.kernel.org and git.kernel.org against AI crawlers
social.kernel.org
social.kernel.org
isn't linux afraid of retaliatory tariffs? should I stock up on linuxes just in case? I've already beefed up toilet paper reserves.
Of course, they're not allowed to work as mascots while touristing. wink wink
- Server: here's a bit of a cancer protein
- Client: okay, here's some compute
- Verifier: the compute checks out
- Server: okay, you are authorized to access cat.gif
wherein a hosting company sees AI bot scans that appear to be coming from millions of unique addresses, thousands of ASNs, many residential and often with a single connection from an IP. The AI bots are proxying through either hacked IoT devices or apps that pay people pennies to let their phone be used as a proxy.
Likely your proof of work will be distributed to the proxies. It'll just make millions of webcams and phones run a little hotter without slowing down the AI bots at all.
Meaning for the same request rate they would need more hosts, which costs them more money.
> Weighs the soul of incoming HTTP requests using proof-of-work to stop AI crawlers
Like picking up the problem from the http headers and returning it as an follow up query.
Or is running a browser part of the PoW essentially?
I've not been able to read your blog with my personal news reader so I was hoping to implement that.
DeepMind's MassiveText dataset was sourced from ~2.35B documents. A difficulty of 4 leading zeros requires an expected 16^4 SHA-256 hashes per site. Benchmarks [1] show an H100 at ~12k MH/s, meaning it would take just ~3.5 hours to solve for all 2.35B pages.
[1] https://gist.github.com/Chick3nman/e1417339accfbb0b040bcd0a0...
I'm curious, what other approaches are you currently considering? In my mind, all roads lead to rate-limiting identifiers with privacy through zk-proofs.
Tor has similarly been using Proof of Work as part of the defense for onion services for around a year and a half now.
I’ve also seen some clear net web sites that use PoW to slow down account creation. Some websites will even adjust the difficulty for individual visitors depending on the recent number of sign-ups coming from their IP block. More signups from an IP block -> higher PoW difficulty for anyone from that IP block -> fewer accounts created by anyone in that IP block over a span of time.
And even in cases where Cloudflare forces a captcha, this POW ran much more quickly than I could solve one by hand
Makes me appretiate the Linux kernel mailing list based contribution method. Very open, very simple.
At this point I guess CF will never fix compatibility bugs in their interstitial pages, and in captcha, with non-default setup of Firefox.
No, not to the point of making it bearable, but at least it becomes rarer for it to take minutes.
Other sites with Cloudflare only take some nice twenty seconds, others just never ever let you go through.
Those checks are a serious contender for worse thing ever happened to the web.
I couldn't agree more. And not only Cloudflare but also Google is increasingly imposing extremely heavy CAPTCHAs, which destroy experiences for many users. They should really think more about false positives too. (And the irony of course is that Google is the biggest scraper in itself.) Thus, please come up with something else than CAPTCHAs or heavy JavaScripts blobs in general.
And if your argument is that it helps DDoS by being a frontend proxy, well, you still need more bandwidth than the DDoS uses, in which case you could do this with a simple "click here" page just as easily.
But please prove me wrong if I've misunderstood something.
> You can see it in action on this recently decommissioned system I'm using for testing purposes: https://ams.source.kernel.org/
Something seriously wrong with it. When I run it with my normal German/EU home connection, it does ~17k iterations. When I run it with a US Atlanta VPN, it only takes ~6k iterations.
I may have to just give up on that though :(
Paying fractions of a penny to view websites has minimal impact on average users but is punishing to spammers.
And no, "anonymous mixer" services don't work either. They're yet another layer of useless profiteering middlemen, which the web already has more than enough of.
I wouldn't trust anything to be safe from the dragnet surveillance apparatus of FVEY.
The problem is, if you want some "proof of money"/"proof of stake", site operators set that up on their own which is a ton of work and people will not want to set up payments for their favourite porn site AND the government can trace back people from payments to site visits, or site operators contract a major vendor (similar to Stripe, Paypal, ...) who handles it and can then trivially be subpoena-ed for records.
That is, do the proof-of-work before visiting the website, then present the currency token that proves the work has been done. The blockchain is there to prevent double spending of the currency token.
I do feel like this is a kind of "those who don't understand it are doomed to re-invent it" type of technologies.
You don’t need to worry about double spending or any of the privacy issues that come with a blockchain when the work is done real time. I know exactly how much work is being done on the server per request with this solution - one hash. There is no need to crawl through a block chain, or submit a transaction, or wait for an external system to catch up.
I want a solution to these crawling bots, not a distributed database. So why would I reach for a distributed database when this simple proof of work system works better, is more understandable, and doesn’t have external dependencies?
I'll mention that, in an extreme case, one could provide a proof-of-work challenge that's directly tied to some cryptocurrency mining effort, so that the work/energy expenditure is captured by the one providing the challenge. This has the effect of the challenger giving money to the one providing the challenge, just in a real-time proof-of-work way. Since money is effectively being exchanged anyway, this punishes systems that don't allow up-front payments to do away with the page load delay.
Although true in an ideal financial sense, it's demonstrably false because having a pay wall of any kind will severely limit usage.
My view is that when the payment threshold is so low, the issue is inconvenience or user friction, not the amount of money involved. I suspect people would be fine with small payments if the user experience was better. For example, Amazon, Netflix, iTunes or Spotify.
The core problem is that alot of crawlers aren't spending their money. They are part of a botnet so they are just spending the victim's money.
But hopefully most of the crawlers aren't botnets or funded by free VC money so they have an economic incentive to avoid crawling systems requiring proof-of-work.
(numbers for presentation purposes only)
If every request goes from 10ms to 100ms - a human won't care as they'll click like 5-6 times total, while reading the output between page load, but a bot will crawl the site 10 times slower.
The monetary payment for removal of the logo is just for companies that don't posses the "know how" to do that and can instead pay a consulting fee.
also where is the anubis avatar, that's so disappointing not to see it
We want AI automation everywhere and crawling is important.
Who is we? I definitely don’t want AI automation everywhere