Tor Users Might Soon Have a Way to Avoid Those Annoying CAPTCHAs
motherboard.vice.com
motherboard.vice.com
The key insight is the use of blind signatures (https://en.wikipedia.org/wiki/Blind_signature) to provide anonymity: 2 unblinded tokens can't be determined as belonging to the same user, because the server didn't see what it signed.
It's a neat idea but it's just moving the trade-off more towards user convenience and away from security: if a server accepts to sign 100 tokens per CAPTCHA, then solving 1 CAPTCHA now allows a robot to do 100x the work it did before.
This allows us to extend the lifetime of an "approval" across time and domains.
So it's not just changing the trade-off. It's enabling a human behavior that should make the user experience much better, while preserving privacy and not really changing the equation for bots.
(Disclaimer: I work for Cloudflare and wrote the original incomplete draft.)
Hi I am 54.123.22.16 let's open a secure connection my public key is 5468..
connection -> This is FarkUser57, token:fark123
connection -> FarkUser57 -> "Actual message 1"
Connection times out.
Hi I am 55.55.55.123 let's open a secure connection my public key is 5592..
connection -> This is FarkUser57, token:fark123
connection -> FarkUser57 -> "Actual message 2"
Connection times out.
As long as FarkUser57 does not map back to your real you, then you preserve anonymity.You do need to trust the endpoint as. The end point can already easily deanonymize you in under a minute based on the traffic it's sending, if that can be compared with the traffic you get from TOR.
The point of a CAPTCHA is to validate that the person browsing a website is a human being. Once they have done that, the person can click on 1 page, 100x pages, or 1000x pages on the site, and from a security perspective that is the same result. The goal was to identify the nature of the entity behind the request, not limit how many times the person accessed the site.
If a server accepts to sign 100 tokens per CAPTCHA, it means that the server assumes those 100 tokens represent the same human person. It do not allow a robot to do 100x more work than before unless the captcha is so faulty that the assumption is incorrect, in which case the captcha is broken and should be replaced. A captcha that only work 90% of the times is not actually preventing bots, since bot owners can easily just add 10x more bots.
Consuming one token bypasses the need to solve one CAPTCHA. Now you have heard about people being hired to solve CAPTCHAs, right? Well with these tokens, every time a human solves one CAPTCHA and obtains 100 tokens, it's as if he helped a robot bypass 100 CAPTCHAs.
Other example: say a CAPTCHA system is so good that only 0.1% of its challenges can be solved by a robot. With these tokens, if 100 are signed per solved CAPTCHA, then it's as if the robot could solve 10% of them. A 100x improvement.
Unfortunately, because it requires that the user use a plugin, this creates two groups of Tor users: those that are using this protocol and those that aren't. This is more information that can be used---with other information---to aid in de-anonymizing users. (To be clear: using ephemeral JavaScript, as they mentioned, is not a credible option, so they have chosen the better route here.)
CloudFlare stores cookies today, yes, but they can be ephemeral with good client cookie policies. A browser plugin usually persists sessions---even if the tokens don't, the fact that it is _installed_ does.
I understand that this is the case for other plugins as well.
In any case, CloudFlare criticism aside: I'm glad that CloudFlare is listening to the Tor community, and has come up with a protocol that does its best to respect users' privacy.
That also means that if TBB were to be extended with a new plugin, it would get to every user very quickly. Especially if they did some sort of time delay (probably overkill), where the browser updated 2 weeks before the update actually kicked in. Then everyone who has upgraded in the past 2 weeks instantly gets group-anonymity, and everyone who hasn't upgraded has only themselves to blame because the browser gives you a nice flashy warning immediately when you open it up.
I hope that the code is audited for back doors by multiple independent parties, but other than that I think this is fantastic.
That's assuming that the Tor Browser Bundle (and Tails) will include it. I'm curious what they will decide.
A few years back back I used cloud flare on my company website in an attempt to improve speed. When I browsed from public Wifi, Cloud flare showed challenges for static sites; this seemed pretty pointless. What kind of attack is cloudflare preventing when they block people from accesing static sites?
tor traffic should stay on the tor network.
https://geti2p.net/en/comparison/tor
Tor: "Designed and optimized for exit traffic, with a large number of exit nodes"
I2P: "Designed and optimized for hidden services, which are much faster than in Tor"
It is not part of the internet, if servers want to provide "anonymous" access they can provide a tor address. But if they don't, then accessing a standard webserver via tor is very little different than any of the other means of illegitimately accessing computers.
->web access via tor seems like a lot of effort to provide what is little more than providing "hacking" services against normal webservers.
I would much rather see for example: https://thepiratebay.se.onion
where tor contacts the dns for thepiratebay.se to get the underlying onion address to use.
much better all round than either going through an exit node (most of which are malicious) to thepiratebay.se or trying to remember uj3wazyk5u4hnvtk.onion or whatever it has changed to now.
It's a real slap in the face to have to enable JavaScript just to gain access to a site, especially if it's a site that doesn't even use JavaScript, or a site that you don't entirely trust.
The funny thing is that it's entirely possible to open a bunch of these tabs in the background (e.g. search for 'site:somesitethatusescloudflare' and just middle click some results that interest you for later reading) and forget about them while they sit there refreshing over and over again, until you really do get blocked (assuming they keep track of unanswered challenges like they do for failed ones).
Lots of sites use Cloudflare to prevent scraping.
This is nonsense.
https://support.cloudflare.com/hc/en-us/articles/202775670-H...
You don't get cached hackernews page because they use cf in front. As the other commenter said, it's nonsense.
DoS. The Tor network has quite a bit of bandwidth (https://metrics.torproject.org/bandwidth.html) and we see DoS attacks at L7 (e.g. floods of HTTP requests) that attempt to knock sites with poor bandwidth/servers offline.
There are many, many other attacks that come through Tor.
that point is beyond moot as we are talking exclusively about sites already fronted by CF
Just because Cloudflare has a ton of bandwidth doesn't mean the origin web server does. An attacker who wants to hurt an origin server just needs to find a URI that Cloudflare doesn't cache (and that's very often /; think of any site that's dynamically generated by some CMS as an example).
So, if you let traffic through to an origin without running any security it's trivial for an attacker to knock off a web server. Thus Cloudflare has to take different security measures to protect servers.
And we do see L7 DoS capable of knocking origin servers down via Tor.
It's absurd for people to be required to do CAPTCHAs just for read access to a page.
The experience of having to repeatedly perform CAPTCHAs the whole time is a real barrier.
1. Benign GET / repeated 1000 times per second. That's a DoS on the server
2. Shellshock. Looks like a benign GET / but nasty payload in User-Agent header
3. Simple GET but with SQLi in the URI
All these are real examples. All seen through Tor (and, of course, without Tor).
The scheme requires the server to detect nonce reuse with reasonable
reliability. However, there might be no need for a zero false positive rate,
because if an attacker needs to make 10,000 requests to have one succeed,
that's possibly an acceptable trade-off.
Therefore, the server could use data structures such as Bloom filters or
cuckoo filters to store tokens that it has witnessed. The parameters of
these structures can be chosen to ensure a false-positive probability of any
given amount. Cuckoo filters may be more efficient but Bloom filters may be
easier to construct.
I don't think this makes sense. "False positive" for a Bloom filter means it thinks an item was previously inserted when it wasn't really. If the filter represents a set of used nonces, the result of seeing an item as previously inserted would have to be blocking the request as a duplicate: therefore a false positive would cause a fraction of legitimate requests to be blocked, not malicious requests to be allowed as the first paragraph seems to imply.This result wouldn't necessarily be unacceptable either, especially if there was some mechanism for the browser to automatically retry a HTTP request with a new token if it received a "reused token" error. However, this behavior would have to be specified, and it's somewhat tricky: it's also possible for tokens to be (actually) reused by accident, e.g. if the user restores their system from a backup or a VM snapshot. In that case, it would make more sense for the browser to respond to a duplicate token error by throwing away all its tokens, since it would have no way to know which of them were clean.
Then again, if the Bloom filter is big enough that the probability of a false positive is very low, even spuriously forcing the user to complete another CAPTCHA may not be the end of the world.
A couple of months ago there was an article on the state of web scraping in 2016: https://goo.gl/eUtkRA. In it, the author easily identified and integrated one of many captcha solvers.
Worst case scenario, there is also crowdsourced mechanical turk style captcha solving as a service: e.g. https://anti-captcha.com.
I guess this raises the question as to whether captchas pose more of a barrier to users than bots, and whether they should be used at all?
The answer was yes. CAPTCHAs are still present because they act as a throttle. The point of a strong CAPTCHA is to limit the amount of abuse that can get through if the other mechanisms break down, by exploiting the fact that humans are kind of slow. Even though OCR can handle most CAPTCHAs these days, it's still not 100% effective, so by ramping up the number of CAPTCHAs you ask users to solve you can still put a throttle on activity. In this way it acts as a last line of defence.
That's why I'm not sure this is going to work out. CAPTCHAs are not a way to distinguish good users from bad, which is how CloudFlare is trying to use them here. CAPTCHAs are way to slow down and throttle traffic that might be auto-generated when you can't tell if it's good or bad. Building a new way to show you solved a CAPTCHA previously doesn't help if the reason you're being shown CAPTCHAs is specifically to slow you down regardless of whether you're good or bad.