I will change the wording on the website and remove the percentage.
864 karma · joined November 10, 2014
Co-founder and CTO of https://friendlycaptcha.com.
Github: https://github.com/gzuidhof
I will change the wording on the website and remove the percentage.
The solving can be a few times slower on a low end device which you should keep in mind. To aid with this when setting the difficulty for your website it shows you an estimate for various device types. This is indeed a downside of PoW approaches.
There is one factor that helps: you can start solving as soon as the form loads, so as the user enters their details/comment it can start solving - I have a hunch that people on mobile devices are inherently slower at entering their data which should help a bit..
Anyway - if you set the difficulty quite high and the solving takes 30 seconds, it takes the user 15 seconds to enter the form - the user would still have to wait 15 seconds. That's not very different from the time to solve image captchas (it's actually lower and doesn't come with a 2MB payload download which isn't great on phones either, and they can keep their privacy + sanity). You could give the user something to do that makes sense for your website (ask them for feedback?).
I've seen people report that they have reduced spam to near nothing already with just a honeypot, but of course I can't verify those claims.
Also with multiple requests from the same IP in a short timespan, the difficulty increases.
There are downsides to to any captcha, but in my opinion make a much better tradeoff. Accessibility and privacy are respected, and there are no annoying tasks.
* Private notebooks make a lot of sense, and perhaps this is where it can be a viable product as well instead of just a useful open source tool? I have plenty of runway, but still it would be more sustainable if it could eventuallu pay my rent.. Similar to the (old) github model I mean (pay for team features / private notebooks you can share with some) with some reasonable limits? My main goal right now is to raise awareness for it. I think many many web developers especially would find the notebook paradigm really useful, but how do I show that?
* A lot of stuff.. auto save, I store every revision but currently you only can retrieve the latest, social features (e.g. recent notebooks, avatars), actual documentation (in notebooks of course).
If you have ideas or thoughts on where I should take this, do reach out! Email in my profile.
Secondly notebooks have to run sandboxed. Most notebooks run the output/code in an iframe but the editor is in the main window (e.g. JSFiddle). The moment you do that you can no longer support notebook style output (cell-output-cell-output) without some serialization step as you can only embed the iframe in one place (so there would be a single output pane). This is also how Project Iodide worked [1].
Starboard's approach is to put the entire editor/runtime in the sandbox which I think is a superior approach for a notebook. It only talks to the outside frame over a thin API (for saving notebooks, refreshing, syncing the content).
[0]: https://caniuse.com/es6-module-dynamic-import [1]: https://alpha.iodide.io/
For sandboxing user content they need to be on their own origin, otherwise by opening the wrong notebook your account could be taken over. Ideally you also have a unique origin per user as well, as otherwise they would share LocalStorage and cookies. So here every user gets <username>.notebook.host for their notebooks (in the future I might completely randomly generate it instead). This is also the domain used for static embedding.
starboard.gg is used as the main entrypoint: it is just a static website (SPA-like). And the APIs are hosted on starboardapi.com. Those two could technically be combined.
And finally there's another one.. starboardproxy.com which acts as a CORS proxy for use in notebooks.
Python works for stuff that was written for it purposefully. Some python libraries work great: numpy, matplotlib, pandas. But many others are not supported directly and can be installed through micropip, but that's quite confusing! These are the issues with Pyodide currently:
* All python code that is executed is synchronous, which means it can not make requests (or call sleep). You can actually make a request using pyodide.open_url('path'), but that makes a synchronous request which isn't really a good idea for anything but small files. I believe asynchronous Python is possible, recent versions of emscripten support it, but it needs someone to put the pieces together (which is not easy!)
* Some packages are huge without being split up. Scipy is actually the only one that's really problematic, I believe it's around 80MB? It should be possible to split it up (into scipy.interpolate, scipy.stats, etc)
* Micropip is asynchronous, so you get a promise when you use micropip.install('my-package'), but you can't "await" it.
* Loading Python initially freezes the browser for a second or two.. not a great user experience.
* Libraries which are not pure python currently need to be manually made compatible with patches - including torch.
Python in the browser (not just in Starboard) needs more love. This is powered by Pyodide[0] which has been making steady progress, but the project is without corporate backing since Mozilla's change of direction. Perhaps Observable can allocate some of their funding towards supporting Python in their notebooks too through this project? Also consider this a call to action for other contributors who want to see Python in the browser become a reality :)
I'm building something similar called Starboard Notebook[0] that has a different set of trade-offs. It ends up being something in between Jupyter and Observable:
* One of the goals is to build Jupyter how it would have been if it was designed for the web (only).
* It's open source [1], plays nice with git (the format is plaintext), and supports local viewing & editing [2]
* Because of that you can host it yourself, put it on your blog / github pages, anywhere.
* There is little magic, it actually is just Javascript at it's base. This means you can use standard browser APIs and HTML, and when you are ready to "graduate" the notebook implementation that should be straightforward. In my eyes notebooks are only for the first 20% of the work that does 80% of the job for small applications/experimentation (which is often where it ends anyway).
* You can "build the ship as you sail": you can load new cell types dynamically at runtime. This is also how Python is supported (through WebAssembly).
* You can have interop with Python and Javascript which is really powerful. Example: Create a drag an drop form using HTML+JS, then process the dropped CSV file using Pandas and visualize using matplotlib.
[0]: https://starboard.gg [1]: https://github.com/gzuidhof/starboard-notebook [2]: https://github.com/gzuidhof/starboard-cli
If you type while(true){} in a notebook only the iframe will break (and usually your browser will prompt you after a while to kill it). When you do only the iframe is no longer functional.
I don't think there's an elegant way to solve it any differently in the browser.
[0]: https://starboard.gg
I'm also building a web-based literate programming environment called Starboard[1] that's probably a hundred times the amount of code (but then it has additional features such as top-level await, plugin, and Python support).
Consider supporting lit-html[2] literals instead of strings for HTML output, it works really well in notebook environments.
One more thing you can consider: a esm tagged literal that uses import(<data url of the code>), you can then use ES Module imports to dynamically load code. Here's a blog post about that approach [3]
[1]: https://starboard.gg [2]: https://lit-html.polymer-project.org/ [3]: https://2ality.com/2019/10/eval-via-import.html
Right now we use standard blake2b as nobody has repurposed a miner to solve hashes for spamming yet.
The thing is, a determined spammer will be able to attack any CAPTCHA - even in labeling tasks there is always the fallback to human-in-the-loop which is cheap at scale (or even free if these are MITM'd users..).
Any (new) CAPTCHA system will have flaws and break in some way at scale, we're open to ideas and of course will try to address any (future) concerns. We are trying to provide a viable alternative to ReCAPTCHA that respects the user - and we will iterate on these problems as we go. Without some new thinking and openness to new approaches we'll be stuck with ReCAPTCHA.
Small nit: the difficulty is set to require around 2.5 million hashes, not 115 thousand. Your point still stands though.
The difficulty can be scaled in a predictable way - it's similar to rate limiting but less all or nothing. We're about to release automatic difficulty scaling per IP, so if many CAPTCHAs are requested/submitted from a single IP the difficulty increases exponentially. Also being able to set the initial difficulty for your usecase and audience is something that should help.
Aside from that there's some more measures on the roadmap: using lists of known-to-be-datacenter IPs, and reputation lists such as [2], as hints to increase the difficulty.
But you're right - it will still be affordable to attack any CAPTCHA, FriendlyCaptcha is no exception. Proof of work approaches have downsides too.
The main ideas behind FriendlyCaptcha vs ReCAPTCHA:
* The user experience is superior. It can happen in the background while the user is doing something else. There is no labeling task.
* We don't have any incentive to collect user data or track users (GDPR compliant, no tracking cookies etc)
* It's as easy to add as ReCAPTCHA to your website. The API is a near copy of ReCAPTCHA's API. You can host the JS code yourself, or even bundle it. With recaptcha it must be third party.
* It works in any browser less than 8 years old (IE>=11), although of course it's much slower in old browsers that don't support WebAssembly.
* It doesn't have inherent accessibility problems (poor eyesight/hearing doesn't matter).
* Open source at its core [3], the SaaS wrapper is not open source.
[1]: https://news.ycombinator.com/item?id=24921288 [2]: https://www.stopforumspam.com/ [3]: https://github.com/friendlycaptcha/
While it's not the perfect captcha either (which I think is impossible), it makes a better tradeoff in terms of UX, price and privacy.
I really only use it as a view counter - seeing that there are actually a few readers motivates me.
[1]: https://plausible.io
The alternative: loading a third party script and multiple images (~2MB) to label for ReCAPTCHA and spending time performing the task also takes some battery (and mental) power.
On using blake2b: I chose blake2b as I was looking to use a hash function that is small in implementation, readily available and already optimized. With WebAssembly the solver can achieve (close to native) speeds and be least be an order of magnitude or two closer to optimized GPU algorithms.
Using specialized hardware, image tasks (and even more so audio tasks which must be present for accessibility reasons) have the same issue that they can be solved by GPU algorithms (i.e. machine learning, in which even a low percentage success rate would already be enough). If you search on GitHub you will find there are more ML captcha cracking repos than captcha implementations - they are probably even easier to get started with than adapting GPU miner code.
Image/Audio Captcha vs ML is an arms race that can be beat for split seconds of compute (even on CPU) or cheap human labeling: it's just as broken. FriendlyCaptcha optimizes for the end user (privacy + effort + accessibility) by not engaging in the arms race - I think it makes a better trade-off. Like the sibling comment pointed out the captcha solving can happen entirely in the background so that hopefully it doesn't even make the user wait.
As for rate limiting/difficulty adjustment: it's not perfect and it could lead to problems if you share the IP with a spammer (and let's be realistic: even with a million users on one IP there won't be tens of users signing up to some forum per minute). Also normal captchas have problems here though: users from these locales already get presented with much more difficult+frequent recaptcha tasks (I also doubt they are localized: American sidewalks are harder to label if you've never seen one in real life). Setting a reasonable upper limit to difficulty may be good enough here.
On not using blake2b: I have considered mutating the hashing algorithm every day randomly to make writing an optimized solver for it all that more difficult - but that would mean one could no longer self-serve the JS+WASM and be done with it. I won't rule it out for FriendlyCaptcha v2 if this does ever become a real problem.
Swapping out the hash function should be easy (the puzzles are versioned to allow for this). If you have a different function in mind and someone implements it in Assemblyscript (so we also have a JS fallback) then we can definitely consider it.
The solver runs in WebAssembly and is really really fast (~4M hashes per second) - but not every browser supports WASM yet (around 0.3% empirically). The JS fallback is around 10 times slower (more in 5+ year old browsers) - for those users you want at least a decent solve time too.
For Gitlab's audience the difficulty can probably be increased a lot - it all depends on the website and usecase. I'm sure the JS fallback's performance can be improved (it involves a lot of operations on 64bit ints that need to be represented as two numbers in JS), happy to accept PRs [1] :)
[1]: https://github.com/FriendlyCaptcha/friendly-pow/blob/master/...
With that the cost (and complexity) of an attack can hopefully be in the same ballpark (or higher) than ReCaptcha - without your end user having to label cars or send data to Google.
But in the end a determined spammer will get through any captcha cheaply (for reference: ReCaptcha solves are sold by the thousands for $1) - we just hope we can do better than ReCAPTCHA, especially UX-wise.
I built a privacy friendly alternative to ReCaptcha called FriendlyCaptcha [1], is there a possibility to see this integrated as a more user friendly alternative?
Happy to chat (e-mail in profile)
Without lit-html Starboard would be nowhere near as flexible or elegant. Starboard has to be completely customizable at runtime, that would probably be difficult to achieve with anything that has a build step, with lit-html it's "just Javascript" all the way down.
If you usually use a framework like React, Vue or Svelte it is still worth trying out lit-html, they are not mutually exclusive, it's simple to learn and deceptively powerful. It's great to have it on your toolbelt.
I am not sure what makes the most sense in terms of monetization in the longer term. I think it's a difficult for many open source projects.
I don't have to tailor ads to the visitor or something, I just want accurate enough stats in aggregate.
* Observable has quite some magic, it's not just Javascript (there is a compile step to make it possible, see the other comment). With Starboard the plan is to make it trivial to export as a static website, in the end Starboard is just away to make a simple website without too much magic that you can instantly share.
* The reactive system they have there is powerful, sometimes probably a bit too powerful. Starboard is just plain javascript (for better or for worse). You can run cells out of order and have complete freedom. Any restrictions on the order that cells should be executed and re-evaluated will have to come from external plugins.
* Observable is a walled garden, you can't export the notebooks into say a text file and use it independently (as far as I know). They own the ecosystem and you can't really move away from it (and fair enough, it's a commercial product!). Starboard notebook is fully open source and it is built with customization, vanilla-ness and portability as main features.