Interesting project, thank you for sharing! From Botpoison's website[0] under FAQ:
> Botpoison combines: > - Hashcash , a cryptographic hash-based proof-of-work algorithm. > - IP reputation checks, cross-referencing proprietary and 3rd party data sets. > - IP rate-limits. > - Session and request analysis.
Seems like it is PoW + IP rate-limits. IP rate-limits. though very effective at immediately identifying spam, it hurts folks using Tor and those behind CG-NAT[1].
And as for invisibility, CAPTCHA solves in mCaptcha have a lifetime, beyond which they are invalid. So generating PoW when the checkbox is ticked gives optimum results. But should the webmaster choose to hide, the widget, they can always choose to hook the widget to a form submit event.
[0]: https://botpoison.com/ [1]: https://en.wikipedia.org/wiki/Carrier-grade_NAT
full disclosure: I'm the author of mCaptcha
I still think PoW alone is not enough as it can be automated, albeit at a slower rate. Most of the time I worry more about low-volume automated submissions than high-frequency garbage. The real value is in the combination of factors, especially what BP call the "session and request analysis" and other fingerprinting solutions.
Very true! I chose to use “captcha” because it's easier to convey what it does than, say, calling it a PoW-powered rate-limter.
> The real value is in the combination of factors, especially what BP call the "session and request analysis" and other fingerprinting solutions.
Also true. I'm not sure if it is possible to implement fingerprinting without tracking activity across the internet --- something that a privacy-focused software can't do.
I have been investigating privacy-focused, hash-based spam detection that uses peer reputation[0] but the hash-based mechanism can be broken with a slight modification to the spam text.
I would love to implement spam detection but it shouldn't compromise the visitor's privacy :)
[0]: please see "kavasam" under "Projects that I'm currently working on". I should set up a website for the project soon. https://batsense.net/about
Disclosure: author of mCaptcha.
I'm in two minds about how I feel inconveniencing those behind CG-NAT. I don't want to punish the innocent, but the ISPs aren't going to move towards better solutions (IPv6) without a push from their paying customers, and they'll never get that push with sufficient strength if we work tirelessly to make the problem affect us and not affect those subscribers.
This is definitely not true today, and i'm not sure it was ever true. I remember the first (or what i remember to be the first) <form>-submitted CAPTCHAs asking me to copy some text, answer a riddle or perform some simple math... all of which could be bypassed by a script. The goal was to avoid a random web-scanning bot to fill your DB with garbage, not protect against targeted attacks.
Nowadays there's large portions of the web i can't browse because of bad IP reputation (i browse through tor) and because the most widely deployed CAPTCHA systems insist that i'm not human.
Google RECAPTCHA in particular has an obsession with fire hydrants (that's not even a thing here in France), traffic lights, bicycles and buses. However, it's never clear if i should click the square where just a tiny bit of the object appears, and some images are so unobvious that what i know to be a bicycle part i'm not sure the algorithm will pick up as such. So i'm trapped in endless CAPTCHA loops in which the robot claims i'm not human.
I realize a GPT3-powered bot would probably say this, but i'm 100% human. Meanwhile, malicious bot authors buy CAPTCHAs for a few cents a dozen from sketchy online marketplaces. CAPTCHAs only prevents legitimate use, as malicious actors have way more considerable means and resources at their disposal than we ordinary users do.
PS: It's not just Google though. As much as they claim to be the good guys, CloudFlare claims an enormous part of their malicious traffic comes from the tor network. However, given that they block 80-90% of my `GET /` requests as if they were malicious, i wouldn't be surprised if these stats were complete bullshit that did not stand scrutiny. But well, if the algorithm says so... (post PS: even HN a few months back started blocking many GET requests from Tor except on the homepage).