HNHacker News
TopNewBestAskShowJobs

protoduction

864 karma · joined November 10, 2014

https://guido.io, e-mail me at me@guido.io.

Co-founder and CTO of https://friendlycaptcha.com.

Github: https://github.com/gzuidhof

submissionscomments
protoduction··on hCaptcha now runs on fifteen percent of the internet
Judging by the downvotes (despite answering the question truthfully), I see it's not a good way to present ourselves, and frankly we don't have to make that claim. It's hard to estimate the real percentage, our customers are happy but measuring what is no longer there is tricky in the real world.

I will change the wording on the website and remove the percentage.

protoduction··on hCaptcha now runs on fifteen percent of the internet
Everybody gets the same difficulty initially which you determine as a site admin, so one should base this on their audience (e.g. Gitlab would have a different device profile from a government website).

The solving can be a few times slower on a low end device which you should keep in mind. To aid with this when setting the difficulty for your website it shows you an estimate for various device types. This is indeed a downside of PoW approaches.

There is one factor that helps: you can start solving as soon as the form loads, so as the user enters their details/comment it can start solving - I have a hunch that people on mobile devices are inherently slower at entering their data which should help a bit..

Anyway - if you set the difficulty quite high and the solving takes 30 seconds, it takes the user 15 seconds to enter the form - the user would still have to wait 15 seconds. That's not very different from the time to solve image captchas (it's actually lower and doesn't come with a 2MB payload download which isn't great on phones either, and they can keep their privacy + sanity). You could give the user something to do that makes sense for your website (ask them for feedback?).

protoduction··on hCaptcha now runs on fifteen percent of the internet
Admittedly it's not calculated so it may be a stretch, it's based on the assumption that the vast majority of spam out there just looks for forms to submit without smarts (which is also why honeypots can be pretty effective, especially if you have a small website that nobody will take the effort to work around it.)

I've seen people report that they have reduced spam to near nothing already with just a honeypot, but of course I can't verify those claims.

protoduction··on hCaptcha now runs on fifteen percent of the internet
I built an alternative[0] that takes a proof of work approach. As a site owner you set the difficulty that makes sense for you: so perhaps you would want 20 seconds of computation before you can submit. The nice thing is that this can happen entirely in the background while the user fills in the form.

Also with multiple requests from the same IP in a short timespan, the difficulty increases.

There are downsides to to any captcha, but in my opinion make a much better tradeoff. Accessibility and privacy are respected, and there are no annoying tasks.

[0]: https://friendlycaptcha.com

protoduction··on Show HN: Jupystar – Run any Jupyter notebook in the browser
* I'm not planning on shutting it down at any point, the cost to operate it is very low (it's just a bunch of static files at its core). But I am personally also always weary to put stuff in online services, that's why I made it all open source: other than observable actually all the pieces are there (including editor and offline support), not just runtime/parser. You can download the files and check them into git, and edit them, and host them trivially. There's no "download all" button yet but if there's a demand for that why not :)

* Private notebooks make a lot of sense, and perhaps this is where it can be a viable product as well instead of just a useful open source tool? I have plenty of runway, but still it would be more sustainable if it could eventuallu pay my rent.. Similar to the (old) github model I mean (pay for team features / private notebooks you can share with some) with some reasonable limits? My main goal right now is to raise awareness for it. I think many many web developers especially would find the notebook paradigm really useful, but how do I show that?

* A lot of stuff.. auto save, I store every revision but currently you only can retrieve the latest, social features (e.g. recent notebooks, avatars), actual documentation (in notebooks of course).

If you have ideas or thoughts on where I should take this, do reach out! Email in my profile.

protoduction··on Show HN: Jupystar – Run any Jupyter notebook in the browser
I think one of the main reasons was the lack of dynamic import support until fairly recently [0], being able to import code dynamically without crazy <script> hacks is very useful in a notebook. The tools you mentioned all pre-date that I think.

Secondly notebooks have to run sandboxed. Most notebooks run the output/code in an iframe but the editor is in the main window (e.g. JSFiddle). The moment you do that you can no longer support notebook style output (cell-output-cell-output) without some serialization step as you can only embed the iframe in one place (so there would be a single output pane). This is also how Project Iodide worked [1].

Starboard's approach is to put the entire editor/runtime in the sandbox which I think is a superior approach for a notebook. It only talks to the outside frame over a thin API (for saving notebooks, refreshing, syncing the content).

[0]: https://caniuse.com/es6-module-dynamic-import [1]: https://alpha.iodide.io/

protoduction··on Show HN: Jupystar – Run any Jupyter notebook in the browser
Well, you need at least two for something like this.

For sandboxing user content they need to be on their own origin, otherwise by opening the wrong notebook your account could be taken over. Ideally you also have a unique origin per user as well, as otherwise they would share LocalStorage and cookies. So here every user gets <username>.notebook.host for their notebooks (in the future I might completely randomly generate it instead). This is also the domain used for static embedding.

starboard.gg is used as the main entrypoint: it is just a static website (SPA-like). And the APIs are hosted on starboardapi.com. Those two could technically be combined.

And finally there's another one.. starboardproxy.com which acts as a CORS proxy for use in notebooks.

protoduction··on Why We're Building Observable
I think the support for HTML,CSS,JS is unlike any other notebook, as well as the possibility to change the runtime at runtime (importing by URL new language plugins / other functionality). This makes it great for documentation (with examples you can actually execute), as an output format for automated reports, as a scriptable Tensorboard, as a platform for interactive articles, and for educational purposes (tutorials, homework).

Python works for stuff that was written for it purposefully. Some python libraries work great: numpy, matplotlib, pandas. But many others are not supported directly and can be installed through micropip, but that's quite confusing! These are the issues with Pyodide currently:

* All python code that is executed is synchronous, which means it can not make requests (or call sleep). You can actually make a request using pyodide.open_url('path'), but that makes a synchronous request which isn't really a good idea for anything but small files. I believe asynchronous Python is possible, recent versions of emscripten support it, but it needs someone to put the pieces together (which is not easy!)

* Some packages are huge without being split up. Scipy is actually the only one that's really problematic, I believe it's around 80MB? It should be possible to split it up (into scipy.interpolate, scipy.stats, etc)

* Micropip is asynchronous, so you get a promise when you use micropip.install('my-package'), but you can't "await" it.

* Loading Python initially freezes the browser for a second or two.. not a great user experience.

* Libraries which are not pure python currently need to be manually made compatible with patches - including torch.

Python in the browser (not just in Starboard) needs more love. This is powered by Pyodide[0] which has been making steady progress, but the project is without corporate backing since Mozilla's change of direction. Perhaps Observable can allocate some of their funding towards supporting Python in their notebooks too through this project? Also consider this a call to action for other contributors who want to see Python in the browser become a reality :)

[0]: https://github.com/iodide-project/pyodide

protoduction··on Why We're Building Observable
I agree that Observable is fantastic, but also wish it was more open.

I'm building something similar called Starboard Notebook[0] that has a different set of trade-offs. It ends up being something in between Jupyter and Observable:

* One of the goals is to build Jupyter how it would have been if it was designed for the web (only).

* It's open source [1], plays nice with git (the format is plaintext), and supports local viewing & editing [2]

* Because of that you can host it yourself, put it on your blog / github pages, anywhere.

* There is little magic, it actually is just Javascript at it's base. This means you can use standard browser APIs and HTML, and when you are ready to "graduate" the notebook implementation that should be straightforward. In my eyes notebooks are only for the first 20% of the work that does 80% of the job for small applications/experimentation (which is often where it ends anyway).

* You can "build the ship as you sail": you can load new cell types dynamically at runtime. This is also how Python is supported (through WebAssembly).

* You can have interop with Python and Javascript which is really powerful. Example: Create a drag an drop form using HTML+JS, then process the dropped CSV file using Pandas and visualize using matplotlib.

[0]: https://starboard.gg [1]: https://github.com/gzuidhof/starboard-notebook [2]: https://github.com/gzuidhof/starboard-cli

protoduction··on Show HN: Scratch.js – Interactive JavaScript Scratchpad
A determined user will still be able to figure out something that blocks forever, for instance run a WASM program that has an infinite loop in it.
protoduction··on Show HN: Scratch.js – Interactive JavaScript Scratchpad
I think so, but a worker won't have access to the DOM and a bunch of other APIs, so the code would be fairly limited in what it can do. Which may be fine for some usecases!
protoduction··on Show HN: Scratch.js – Interactive JavaScript Scratchpad
On Starboard[0] I approached this by sandboxing the notebook code in an iframe on a different origin. This sandboxing has to be done anyway to prevent XSS.

If you type while(true){} in a notebook only the iframe will break (and usually your browser will prompt you after a while to kill it). When you do only the iframe is no longer functional.

I don't think there's an elegant way to solve it any differently in the browser.

[0]: https://starboard.gg

protoduction··on Show HN: Scratch.js – Interactive JavaScript Scratchpad
I love how simple and effective this is.

I'm also building a web-based literate programming environment called Starboard[1] that's probably a hundred times the amount of code (but then it has additional features such as top-level await, plugin, and Python support).

Consider supporting lit-html[2] literals instead of strings for HTML output, it works really well in notebook environments.

One more thing you can consider: a esm tagged literal that uses import(<data url of the code>), you can then use ES Module imports to dynamically load code. Here's a blog post about that approach [3]

[1]: https://starboard.gg [2]: https://lit-html.polymer-project.org/ [3]: https://2ality.com/2019/10/eval-via-import.html

protoduction··on ReCAPTCHA and the Anonymous Experience
We can change the hashing algorithm at will which is different from cryptocurrencies (potentially even on a timer). By changing here I don't even mean swapping out entirely, but even randomly changing the operations inside the hashing function - which will make it a moving target for any ASIC or even GPU implementations.

Right now we use standard blake2b as nobody has repurposed a miner to solve hashes for spamming yet.

The thing is, a determined spammer will be able to attack any CAPTCHA - even in labeling tasks there is always the fallback to human-in-the-loop which is cheap at scale (or even free if these are MITM'd users..).

Any (new) CAPTCHA system will have flaws and break in some way at scale, we're open to ideas and of course will try to address any (future) concerns. We are trying to provide a viable alternative to ReCAPTCHA that respects the user - and we will iterate on these problems as we go. Without some new thinking and openness to new approaches we'll be stuck with ReCAPTCHA.

Small nit: the difficulty is set to require around 2.5 million hashes, not 115 thousand. Your point still stands though.

protoduction··on ReCAPTCHA and the Anonymous Experience
You're right that it won't stop determined attackers, there was some prior discussion here [1]. The idea is that it's good enough - while not punishing your users as much.

The difficulty can be scaled in a predictable way - it's similar to rate limiting but less all or nothing. We're about to release automatic difficulty scaling per IP, so if many CAPTCHAs are requested/submitted from a single IP the difficulty increases exponentially. Also being able to set the initial difficulty for your usecase and audience is something that should help.

Aside from that there's some more measures on the roadmap: using lists of known-to-be-datacenter IPs, and reputation lists such as [2], as hints to increase the difficulty.

But you're right - it will still be affordable to attack any CAPTCHA, FriendlyCaptcha is no exception. Proof of work approaches have downsides too.

The main ideas behind FriendlyCaptcha vs ReCAPTCHA:

* The user experience is superior. It can happen in the background while the user is doing something else. There is no labeling task.

* We don't have any incentive to collect user data or track users (GDPR compliant, no tracking cookies etc)

* It's as easy to add as ReCAPTCHA to your website. The API is a near copy of ReCAPTCHA's API. You can host the JS code yourself, or even bundle it. With recaptcha it must be third party.

* It works in any browser less than 8 years old (IE>=11), although of course it's much slower in old browsers that don't support WebAssembly.

* It doesn't have inherent accessibility problems (poor eyesight/hearing doesn't matter).

* Open source at its core [3], the SaaS wrapper is not open source.

[1]: https://news.ycombinator.com/item?id=24921288 [2]: https://www.stopforumspam.com/ [3]: https://github.com/friendlycaptcha/

protoduction··on ReCAPTCHA and the Anonymous Experience
I built FriendlyCaptcha [1], it's a proof of work based alternative to reCaptcha that is accessible.

While it's not the perfect captcha either (which I think is impossible), it makes a better tradeoff in terms of UX, price and privacy.

[1]: https://friendlycaptcha.com

protoduction··on Ask HN: Do you use analytics on your personal blog/site? Why?
My personal site is HTML/CSS only, but I decided to add a privacy-ok simple analytics solution to it [1]. So now it does have some JS on it, but I can live with its <1KB size. By deferring it it's just as fast as before.

I really only use it as a view counter - seeing that there are actually a few readers motivates me.

[1]: https://plausible.io

protoduction··on Wikimedia is moving to Gitlab
It's not perfect, but maxing a single core for 20 seconds on an older smartphone is a necessary evil for this kind of captcha.

The alternative: loading a third party script and multiple images (~2MB) to label for ReCAPTCHA and spending time performing the task also takes some battery (and mental) power.

protoduction··on Wikimedia is moving to Gitlab
It is also important to note that the 6-12 seconds and 7-14 seconds reported in the paper is for the garbled text CAPTCHAs, not for image labeling tasks (fire hydrants, cars, etc).
protoduction··on Wikimedia is moving to Gitlab
I'll try to provide my thoughts on each of the issues you've mentioned, let me know if there's something I missed.

On using blake2b: I chose blake2b as I was looking to use a hash function that is small in implementation, readily available and already optimized. With WebAssembly the solver can achieve (close to native) speeds and be least be an order of magnitude or two closer to optimized GPU algorithms.

Using specialized hardware, image tasks (and even more so audio tasks which must be present for accessibility reasons) have the same issue that they can be solved by GPU algorithms (i.e. machine learning, in which even a low percentage success rate would already be enough). If you search on GitHub you will find there are more ML captcha cracking repos than captcha implementations - they are probably even easier to get started with than adapting GPU miner code.

Image/Audio Captcha vs ML is an arms race that can be beat for split seconds of compute (even on CPU) or cheap human labeling: it's just as broken. FriendlyCaptcha optimizes for the end user (privacy + effort + accessibility) by not engaging in the arms race - I think it makes a better trade-off. Like the sibling comment pointed out the captcha solving can happen entirely in the background so that hopefully it doesn't even make the user wait.

As for rate limiting/difficulty adjustment: it's not perfect and it could lead to problems if you share the IP with a spammer (and let's be realistic: even with a million users on one IP there won't be tens of users signing up to some forum per minute). Also normal captchas have problems here though: users from these locales already get presented with much more difficult+frequent recaptcha tasks (I also doubt they are localized: American sidewalks are harder to label if you've never seen one in real life). Setting a reasonable upper limit to difficulty may be good enough here.

On not using blake2b: I have considered mutating the hashing algorithm every day randomly to make writing an optimized solver for it all that more difficult - but that would mean one could no longer self-serve the JS+WASM and be done with it. I won't rule it out for FriendlyCaptcha v2 if this does ever become a real problem.

Swapping out the hash function should be easy (the puzzles are versioned to allow for this). If you have a different function in mind and someone implements it in Assemblyscript (so we also have a JS fallback) then we can definitely consider it.

protoduction··on Wikimedia is moving to Gitlab
The default difficulty is set to a difficulty that makes sense on websites that have a varied audience (which includes some ancient browsers on old devices).

The solver runs in WebAssembly and is really really fast (~4M hashes per second) - but not every browser supports WASM yet (around 0.3% empirically). The JS fallback is around 10 times slower (more in 5+ year old browsers) - for those users you want at least a decent solve time too.

For Gitlab's audience the difficulty can probably be increased a lot - it all depends on the website and usecase. I'm sure the JS fallback's performance can be improved (it involves a lot of operations on 64bit ints that need to be represented as two numbers in JS), happy to accept PRs [1] :)

[1]: https://github.com/FriendlyCaptcha/friendly-pow/blob/master/...

protoduction··on Wikimedia is moving to Gitlab
I don't have all the answers yet, but indeed rate limiting a larger block (at least /64), or even at multiple prefix sizes with different weighting makes sense.
protoduction··on Wikimedia is moving to Gitlab
It's not enabled yet in production - but the main mechanism is by increasing the difficulty as more requests are made from an IP in a certain timeframe (it's basically rate limiting at that point). Think: every 3rd request in a minute doubles the difficulty with some cooldown period.

With that the cost (and complexity) of an attack can hopefully be in the same ballpark (or higher) than ReCaptcha - without your end user having to label cars or send data to Google.

But in the end a determined spammer will get through any captcha cheaply (for reference: ReCaptcha solves are sold by the thousands for $1) - we just hope we can do better than ReCAPTCHA, especially UX-wise.

protoduction··on Wikimedia is moving to Gitlab
Hi phikai,

I built a privacy friendly alternative to ReCaptcha called FriendlyCaptcha [1], is there a possibility to see this integrated as a more user friendly alternative?

Happy to chat (e-mail in profile)

[1] https://friendlycaptcha.com/

protoduction··on The new Jupyter Book
I built Starboard (https://starboard.gg), its interface is similar to Jupyter, but it runs entirely in the browser (which has pros and cons)
protoduction··on Show HN: Starboard – Fully in-browser literate notebooks like Jupyter Notebook
A late response, but I can add a bit more motivation.

Without lit-html Starboard would be nowhere near as flexible or elegant. Starboard has to be completely customizable at runtime, that would probably be difficult to achieve with anything that has a build step, with lit-html it's "just Javascript" all the way down.

If you usually use a framework like React, Vue or Svelte it is still worth trying out lit-html, they are not mutually exclusive, it's simple to learn and deceptively powerful. It's great to have it on your toolbelt.

protoduction··on Show HN: Starboard – Fully in-browser literate notebooks like Jupyter Notebook
I am not sure yet. I am working on this full-time, but it's my own savings that keep me fed so that's not very sustainable.

I am not sure what makes the most sense in terms of monetization in the longer term. I think it's a difficult for many open source projects.

protoduction··on Show HN: Starboard – Fully in-browser literate notebooks like Jupyter Notebook
Not yet, but it's near the top of my to-do list
protoduction··on Show HN: Starboard – Fully in-browser literate notebooks like Jupyter Notebook
In a way that's an important feature. I don't want to (be able to) track individual users, and I don't think I have to.

I don't have to tailor ads to the visitor or something, I just want accurate enough stats in aggregate.

protoduction··on Show HN: Starboard – Fully in-browser literate notebooks like Jupyter Notebook
Observable is an amazing tool, some of the differences:

* Observable has quite some magic, it's not just Javascript (there is a compile step to make it possible, see the other comment). With Starboard the plan is to make it trivial to export as a static website, in the end Starboard is just away to make a simple website without too much magic that you can instantly share.

* The reactive system they have there is powerful, sometimes probably a bit too powerful. Starboard is just plain javascript (for better or for worse). You can run cells out of order and have complete freedom. Any restrictions on the order that cells should be executed and re-evaluated will have to come from external plugins.

* Observable is a walled garden, you can't export the notebooks into say a text file and use it independently (as far as I know). They own the ecosystem and you can't really move away from it (and fair enough, it's a commercial product!). Starboard notebook is fully open source and it is built with customization, vanilla-ness and portability as main features.

← PreviousPage 4 of 5Next →