The Trouble with Tor
blog.cloudflare.com
blog.cloudflare.com
Say what you will about cloudflare, that's an impressive move.
Imagine what that was like for the technical support and customer success teams who were helping customers with their sites.
The problem being solved was not "anonymous internet access gets abused a lot", it was "the mechanism we use to combat abuse is too aggressive for some segments of our users and thereby denying service to actual humans, and these challenges are invisibile to our engineers and employees because they don't browse on connections that trigger the system".
As the blog post shows (and this is backed up by my experience as a user of Tor) CloudFlare significantly improved its handling of this in response. So I would say this is a success. I am sure that pressure from employees, including support guys, caused this long-running issue to finally get the attention it deserved.
It sucks. But that's exactly the point. Using the internet is a mission critical thing for many people and for some that means using Tor Browser or similar to get around oppressive governments. This sounds like a really effective way to make sure the things you're doing are impacting your customers minimally. I bet you guys screamed loudly and were heard more clearly than a random customer trying to use Tor day-to-day.
The CAPTCHAs are (mostly) easy to solve, but all of the ones I was presented with were "pick the right one out of 9 different images" and loading the entire CAPTCHA in Tor Browser took several seconds (and many revealed a new image after clicking one of the 9). This is then repeated at least once (I received three on one site, I'm guessing because I didn't know if the picture was a store front or just the front of some building). After completing the challenge I was given a connection error and had to repeat the entire thing again in one case.
There are much lower bandwidth CAPTCHAs out there and those should be favored over these large image-based ones for connections originating from the block of addresses represented by Tor exit nodes.
Once you get onto the site, it loads more slowly than in a non-Tor connection, but a news site I hit loaded everything at about the same speed as the little reCAPTCHA form, so I'm left wondering if it's something related to reCAPTCHA.
Its like the http blocking load RTT problem magnified by 10.
You now get a page with images that have checkboxes next to them. When you submit the form, you get a Base64 key to paste into a text box.
For example, the question was "Does this have a river in it?" with a picture of the Grand Canyon, where you couldn't quite see down to the Colorado River.
But has he such low opinions of his employee that he couldn't have simply told them "captchas are bad UX, try it out if you don't believe me"?
it just seems condescending to me
It was more of a "let's really feel the pain from this and try to come up with a better solution". And at the end of it they didn't. They felt the pain of having to do it for just about everything, and still couldn't come up with a better solution other than "improve the captcha system".
And from there they gave some suggestions on how they could work with tor to give a better experience.
Simply saying "captchas suck, find a better way" would have gotten nowhere because the people involved wouldn't personally know the pain points, they wouldn't know the exact issues, or what exactly about it was the most annoying. (was it the captcha itself, or the frequency you had to fill it out? Did it completely break some of their workflows, and what were the ways they ended up getting around those breakages, and can they prevent them? Perhaps they found the redirect page broke some browsers/software/extensions!)
I would kill to have a one-on-one with a typical "user" of my products. A long sit down meeting to understand exactly what they want/need above "i want it to work better". Making your developers/staff users of your product is a great way to do that and to really internalize it.
But the bigger issue I would wonder about is Google's reputation systems. Google does not treat all CAPTCHA requests equally.
An office full of CloudFlare employees peaceably going about their daily browsing is going to get a much easier CAPTCHA situation than a Tor IP containing a mix of automated nefarious activity and individuals peaceably browsing from infomatically repressed countries.
Tor uses hashes generated with the weak SHA-1 algorithm to generate .onion addresses and then only uses 80 bits of the 160 bits from the hash to generate the address, making them even weaker.
Other weaknesses include use of RSA-1024, people have been complaining to no avail since at least 2013: http://arstechnica.com/security/2013/09/majority-of-tor-cryp...
Basically what that means for Tor, is that while it'd be pretty easy for a Onion site operator to generate two keys corresponding to the same .onion address, an _attacker_ still has to do 2^80 work to attack a site by generating a key with the same Onion address. While that's not great - 2^128 work is considered "standard" in cryptographic work - 2^80 work is still hard enough that there are probably cheaper ways of attacking Onion sites (for reference, the cumulative total work done by all Bitcoin network miners in the entire history of Bitcoin is about 2^80 hashes).
As for the 1024bit pubkeys, I'm not sure what the status of that is; from what I hear Tor is actively working towards a Onion redesign that will fix these issues, and longer pubkeys may have already been fixed.
https://gitweb.torproject.org/torspec.git/tree/proposals/224...
> This isn't actually a vulnerability yet, as SHA-1 is only known to be vulnerable to collision attacks
should be:
* this isn't a vulnerability (there are no reason to believe that we might see it vulnerable to pre-image attacks one day)
* SHA-1 is thought to become vulnerable to collision attacks
> total work done by all Bitcoin network miners in the entire history of Bitcoin is about 2^80 hashes
without thinking of hashes, current cycles done per second by the bitcoin network is around 2^90
SHA-1 is actually known to be, not just thought to become, vulnerable to collision attacks at less than the full bit strength of the hash[0].
The important part is the pre-image resistance, of which there is no known attack.
But there is a difference between the theory and actually finding a collision. And then a huge difference as well on how to exploit that.
I just give up immediately when I see a Cloudflare captcha page on Tor.
Have you tried recently? They have gotten way better. There was a time not long ago where I would have agreed with you (the CAPTCHAs were literally impossible for a human in most cases).
CloudFlare has toned the CAPTCHAs down a lot recently, they're now presenting image classification tasks (select street signs/bodies of water/storefronts/cactuses/...), and in my experience, my success rate is close to 100% on those.
Admittedly they still make me solve three of them and they get tiring, but at least they're not completely cutting off access to Tor users anymore.
Works every time.
The ones that are impossible for my are the street signs, and after that all of the cultural ones (I see US captchas and when asked to select a sandwich or a recreational vehicle I'm doomed to not complete them).
An interesting side effect of failing a captcha is that to Google this looks like proof the captcha is working, that you're likely to be a bot, and that they should definitely give you the hard captchas.
As such, if you cannot complete a captcha the chances increase that you must now complete multiple difficult captchas.
The audio captcha is delightfully simpler though it does take a moment longer to complete.
I could be wrong though
Am I supposed to click the tiny corners of signs in another square? If it is a sign that is not a "street sign", such as a billboard, do I click it?
::confused::
Though it seems that the idea of allowing GETs when a site isn't under load attack is probably the right solution?
https://plus.google.com/104092656004159577193/posts/H2sakaRx...
I'm also aware of some tools/approaches which address the question of fair anonymity -- ensuring well-behaved clients while retaining anonymous status for the client. Best I'm aware these are very experimental. I've also forwarded the information to TK Hyponnen of F Secure, who may have some impression of the approaches.
FAUST: https://gnunet.org/node/1704 (Efficient, TTP-Free Abuse Prevention by Anonymous Whitelisting | GNUnet)
Fair Anonymity: http://arxiv.org/pdf/1412.4707v1.pdf
Assessing these is beyond my skills, but the references may be useful to CloudFare (or others).
Some background - we run several SaaS services for schools, which are politically and socially non-sensitive. The only realistic reasons anyone would want to connect anonymously would be nefarious. Allowing Tor traffic is like a bank having a special ATM round the back with no security cameras - you're giving a free pass to attackers to try anything they want with impunity.
I'm having a hard time seeing what the compensating advantage is. How does not accepting Tor traffic to our "normal" sites lessen the anonymity of Tor traffic to sites where it is important?
Really, no one has any business looking at anything I do on the Internet. I don't use Tor for everything but I may use it when I'm connecting via a network that I believe to be hostile (i.e. just about everything outside of my house)
Which ISP do you use? I would be surprised if a trustworthy ISP exists.
When they're completely consumed by their new owners, TPG, then no, there won't be a trustworthy ISP in Australia. I have high hopes for SkyMesh, though.
Edit:
> I'm having a hard time seeing what the compensating advantage is. How does not accepting Tor traffic to our "normal" sites lessen the anonymity of Tor traffic to sites where it is important?
The anonymity of Tor depends upon diversity. The more people using Tor for more things, the harder it is to correlate any particular person's traffic.
That being said, it is your website, and if you decide to block Tor you have as much right as your users do to use Tor. But I'd ask you to think about whether you are actually attaining any benefit.
Using TOR for everything helps to build plausible deniability, if you are always on TOR then an external watcher can't determine if you are doing good or bad stuff.
Of course good and bad are relative terms, if your country doesn't have free speech obviously the "bad stuff" is just speaking freely
Sounds like security by obscurity.
By the way, some people use a track pad (with a stylus), where the mouse can jump discontinuously.
You might use a trackpad, so you'd "fail" that test, but your useragent is normal, you've been seen before with those cookies, and your IP is good so you are fine.
But if your IP is a known "bad actor", your useragent is something never before seen, your mouse movements are abnormal, and your keyboard inputs are instant, well all of that combined means you are getting blocked.
What if I install a new computer on an IP address freshly provided to me by my ISP? Or what if I just open a new incognito window? Will I get blocked?
> your mouse movements are abnormal, and your keyboard inputs are instant
It seems to me these are really easy to fake programmatically.
If there are enough "red flags" you'll probably get a captcha, if there are an overwhelming number of "red flags" you might just get blocked.
Again, just opening an incognito window or a new computer/ip isn't going to do it alone.
>It seems to me these are really easy to fake programmatically.
I'm sure they are, but they make the bar for "automated traffic" a little higher, and weed out some of the lower hanging fruit.
Conversely, while signed into my gmail account I get passed through immediately, regardless of whether I click the box or tab into it and hit space.
I'd really hope that CloudFlare whitelists Tor for CDNJS.
From this perspective, captchas are a very minor concern. I'm as pro-privacy as anyone, but this expectation that anonymous activity is supposed to be easy or convenient will never be satisfied. Thousands of years of lessons from both military and civilian clandestine operations bears out the critical lesson that anonymity is, by default, very very inconvenient. Nothing is going to change that.
Do uses who publisher their email address on their website that is hosted through cloudfare see a lower number of spam? It should be a fairly easy thing for cloudfare to test, while also testing vulnerability and login attempts. As an aside, it would also be interesting to see if there is a quality vs quantitative differences in the malicious activity (ie, if serious attempts are done through botnets, and script kiddie activity is done through tor).
The last a final test in order to verify a security measure that has such a high cost as this one, is to ask if its has any meaningful impact to the end result. A website with 10000 vulnerability scans per day is not going to be meaningful improved if it was reduced to 5000 per day, even if that is a 50% reduction. If there is a known vulnerability, the site is going to get hijacked either way.
admin.example.com
Such a record is usually not routed through Cloudflare because the last thing a webmaster wants is to solve captchas for their own website. They don't however care much for their visitors if they're subjecting (possibly a substantial amount of them) to captcha solving nonsense.The content in the non-cf sections of a site can still be accessed because the webmaster is lazy and didn't care to check if a visitor can do a DNS DIG on all their DNS records.
Or you can simply use TOR pluggable transports to pretend you're Googlebot, and also hide all your traffic in Google-like traffic.
I would reserve this for rare cases as there are people in censorship prone countries who really need this bandwidth :)
Ideally you're looking to use TOR as the first hop, and then you dial into the wider Internet with a VPN, or as I mentioned: Using various Google services to camouflage traffic instead of a VPN. This is where pluggable transports come in, because Google doesn't like TOR, so you want to choose how you're connecting to Google, and get to traverse the TOR network to find an optimal route.
While I agree that you can mask your exit from the Tor network via an additional proxy or VPN, that's not the role of pluggable transports. They're only for connecting in to the Tor network, not out from it.
And would negate any anonymity offered by using a VPN.
VPNs do not provide anonymity.
> With most browsers, we can use the reputation of the browser from other requests it’s made across our network to override the bad reputation of the IP address connecting to our network. For instance, if you visit a coffee shop that is only used by hackers, the IP of the coffee shop's WiFi may have a bad reputation. But, if we've seen your browser behave elsewhere on the Internet acting like a regular web surfer and not a hacker, then we can use your browser’s good reputation to override the bad reputation of the hacker coffee shop's IP.
I occasionally use a VPN, but I've never gotten a CloudFlare captcha. Is it possible that you might be doing something else other than just using a VPN, such as blocking cookies?
For those on phones, you can still opt for CAPTCHA if you don't want to kill your battery.
It would be interesting to see how fast you could make the JavaScript code. I'm sure my version is just terribly unoptimized. The requirement to support the least common denominator will present a major problem, though.
If the wasm effort works out (it's looking like it will), it would hopefully alleviate the issue you present entirely and make this solution viable in a very sane way.
Additionally, specifically with Tor users, you can expect a large chunk of the user base to have JavaScript disabled completely. You can do many things with JavaScript that could be used to build a browser fingerprint, so someone who's already using software to browse the web anonymously is very likely to disable that.
The JavaScript issue can be gotten around with a browser plugin that does it as well, which would be easy to bundle on the existing Tor browser. JavaScript would still be fine for all the VPN users who get stuck with these things, and the regular users who get them occasionally for whatever reason.
If we assume 5 seconds of CPU time per comment, that's ~17k per day or ~500k per month. The first captcha solving service I found sells 100k solved captchas for $139, so that's about $700 for 500k. As a spammer, I could probably post 5 to 10 times more comments for the same amount of money using your system. This is obviously a very rough estimate, but it should get my point across.
If the delay is only once per, let's say a domain, then you don't do anything against the spammers, they only have to wait a full delay once.
These days a spammer could train a ConvNet to pass reCAPTCHA with > 90% accuracy and very little processing overhead if they really wanted to. The only reason it works is because the bar is "high enough" that it's cheaper to spam somewhere else that has a lower bar.
[1] https://s3-us-west-2.amazonaws.com/excredo/hashrate.html
I don't get how this is different than a super cookie. Anyway, I think globally that's a well balance reaction to the TOR issue.
The supercookie would survive or be detectable across multiple browser sessions (keep in mind that the Tor Browser automatically deletes regular cookies when you quit). The behavior that CloudFlare is describing here works within a single Tor Browser session but not across multiple sessions.
I believe the Tor Browser is willing to send some first-party session data to a site after changing circuits, so that you wouldn't be logged out of an account if you logged in over Tor and kept that session active for long enough that Tor switched over to a different circuit. This behavior is basically what should allow CloudFlare to recognize that a particular Tor user has recently passed a CloudFlare CAPTCHA (on a particular site). However, if the user quits and restarts Tor Browser, CloudFlare will no longer be able to detect that it's the same visitor (if it could, that would be the supercookie case).
I was pleased when, recently, I found out they switched to the image based one. Sure, sometimes it still refuses to accept that I selected all the street signs but at least I don't have to give up in frustration after 30 consecutive failed attempts
Wouldn't something like rate limiting or proof of work achieve the same result? If you're simply allowing someone to browse, you don't actually care whether a user is real or not. You care about stopping automated comments/spam.
Is this just another tentacle of the advertising industry?
On the other hand, anonymity is also something that provides value to online attackers. Based on data across the CloudFlare network, 94% of requests that we see across the Tor network are per se malicious. That doesn’t mean they are visiting controversial content, but instead that they are automated requests designed to harm our customers. A large percentage of the comment spam, vulnerability scanning, ad click fraud, content scraping, and login scanning comes via the Tor network. To give you some sense, based on data from Project Honey Pot, 18% of global email spam, or approximately 6.5 trillion unwanted messages per year, begin with an automated bot harvesting email addresses via the Tor network.
Rate limiting would block legitimate users, and pow doesn't impede the malicious uses of tor.
The only vaguely reasonable one I can see there is 'ad click fraud', and I think that fundamentally restricting the usefulness of a site for advertising purposes is awful.
Yes there is: "we don't provide censorship as a service".
By using reCAPTCHA, mentioned in the article as a preferred solution, visitors from China are routinely blocked, as reCAPTCHA it is now hosted by Google.
The trouble with CAPTCHA hosting: And if you intend to do anything with China based orgs.
1. It's a charity tax; you have to convince people to incur the cost of Tor (i.e. CAPTCHAs everywhere) for activities that don't require Tor.
2. You can't neutralise a poison by diluting it.
Firstly, from the operators' POV, if there's a widespread agreement that people use Tor even though they don't need to, then they know voluntary users can be pressured not to use Tor through sheer inconvenience. Even if you wanted to boycott a service that blocked Tor, it's notoriously hard to make good on that threat unless you wield a lot of power or annoyed a very large number of people. So the consequences are minor.
Secondly, the percentage of malicious Tor traffic is a red herring. What operators care about is the origins of malicious traffic. If 50% of your attacks come from one particular country (or Tor) and the cost of losing that traffic is less than the cost of that malicious traffic, there is a real incentive to block that traffic. Combined with the first point, the cost of losing voluntary Tor users is insignificant if they can easily choose not to use it.
People are willing to invest personal ressources for charitable purposes. Why not here?
People are willing to fight against discrimination. Why not against discrimination of Tor users?
> 2. You can't neutralise a poison by diluting it.
There is also poisonous traffic from non-Tor adresses.
> Combined with the first point, the cost of losing voluntary Tor users is insignificant if they can easily choose not to use it.
People would strictly avoid restaurants that don't serve coloured people. Why don't they avoid services that don't serve Tor users?
Some people will (and do) do it. You're right that you won't convince everyone to run Tor all the time, but you won't need to.
Also, mozilla have been floating ideas such as integrating Tor into firefox for use in a new kind of private browsing mode. This affects things considerably.
> 2. You can't neutralise a poison by diluting it.
Yes, you can. Both in the metaphorical as well as the direct sense. At some point the solution is too dilute for the poison to cause harm.
I use Tor all the time. I know of local web shops that have rejected the idea of blocking Tor because they looked at their logs and saw that they get actual sales through it - from people like me.
This could be a benefit for the website (lower server load) or a harm (fewer people appear to be reading their content, fewer people see their ads). Whatever the case may be, I'm caught in the crossfire between crackers and servers. I don't care about their war at all. As far as I'm concerned, I'm winning.
Cloudfare wanted me to solve a CAPTCHA to read their article. I tried to archive it, but arhive.is already had a copy of it. This happens to me quite often. So, obviously, I'm not the only one who has figured out a way around their war.
> Security, Anonymity, Convenience: Pick Any Two
Nah, I usually have all three.
...
> Firefox (Tor Browser) has plugins to get a copy from arhive.org, archive.is, or google-cache. So if the page asks me to solve a CAPTCHA, I don't visit them.
You don't have convenience.
Honestly, I don't know how I can make the process as complicated as you described it, even if I wanted to. In reality, it is no more complicated than right-click, open in new tab.
Whitelisting is not convenient.
If people are pretending it is, they're doing a disservice to the security community. Kinda how like people pretend GPG is usable and convenient, thus holding back progress in the security UX front.
Sorry for the late reply.
I cannot imagine how Cloudflare could distinguish VPN traffic routed through Tor and standard traffic routed through Tor. The only difference is a hop on the front end, no change to what comes out the exit node.
Edit: This is good if you are trying to maintain an Internet profile (i.e. Facebook, twitter etc.) that isn't tied to your true identity.
If that's not your aim (like you say - being signed in to the same facebook account all day suffers you the same problem) then this isn't an issue.
But this isn't what a lot of tor users want tor for.
for example having a backend algo offer certain captcha's that show up only in certain areas of the world?
I feel like this is entirely possible and is part of the reason I will not complete and captcha's moving forward.
Here's what I found:
Large Image based CAPTCHA: Tor is slow. I'm on a 75/25Mbps internet connection and it loads images similar to 24k dial-up. The CAPTCHA I was presented with was the highest bandwidth CAPTCHA I've ever seen. I was given a 9 pictures and needed to select "Bodies of Water". Each click yielded a new square. I had to wait for 4 additional images to load before I could click "Validate". This took over a minute. Then, repeat, this time "Store Fronts", which were hard to discern (is it a blurry Apartment Building front or Store Front?). I received a connection error on one site so I had to repeat the process. With Javascript features turned off, it was a little easier, but included the extra step of having to paste a Base64 encoded string into a text box which failed twice. Every site I tried gave me this CloudFlare CAPTCHA page.
One of the sites I pulled up had no images. I set my privacy settings to the least protected, enabling JavaScript and HTML5, assuming this was the problem. Nope. They used images from another site and I had to grab the image URL and paste it into a browser to see what was going on. It was yet another CAPTCHA. A few minutes later, the previous site displayed images properly.
On to Google. Privacy settings are still at the weakest setting. Type "Google" into the search engine and I get a "wavy text in an image" CAPTCHA. This loads quickly and is easy to answer, but just results in another CAPTCHA. I gave up after 10 tries. Bing, Yandex and Yahoo all worked with Yandex only presenting a CAPTCHA once after the third search I did (simple, like Google, but worked).
This is a terrible experience for people seeking to get around oppressive governments. While I applaud CloudFlare for dogfooding their CAPTCHA system, I doubt they did it in a way that truly simulates the experience via an extraordinarily slow internet connection which is what I ended up with when using Tor Browser. I wonder how much slower things would be if my internet connection was 1Mbps or being interfered with by government infrastructure. I understand the trade-off between securing a site from "evil traffic" that is more likely to originate from a Tor exit node, but why must they use such a bandwidth intensive CAPTCHA? A browsing experience that would have taken seconds to complete took me almost 5 minutes (and a lot of frustration) not including actually consuming the content I was looking for. Are the text in image based CAPTCHAs not good enough for this task? Are there other reasons I'm missing?
Maybe browsers shouldn't request <img> src with Accept / and cloudflare should use that to detect whether it can actually serve html?
unfortunately I do not trust tor.[0][1] I'm unsure the internet will ever be anonymouse unless large completely private networks start gaining popularity.
[0]http://fusion.net/story/238742/tor-carnegie-mellon-attack/ (11/2015) [1]http://www.theguardian.com/technology/2014/jul/22/is-tor-tru...
Treating TOR traffic the same as non-TOR traffic makes no sense; read the main link for confirmation they do.
Case in point, and for starters, STOP repeatively requiring a user from a session to keep passing "I'm not a robot" tests. Set a global cookie that's valid for the session, across all of Cloudflair's network, and honor it.
If the "I'm not a robot test" doesn't work unless it's repeatively given, then that is the problem, not TOR.
Please address this issue; thanks.
The possibility of exploitation does not mean exploitation or making exploitation easier is fine.
I think human generated traffic may have priority but blocking bots entirely is nonsense: ultimately the user agent is always a "bot" acting on behalf of an actual person: by blocking this traffic you may always break some user workflow.
IPs are cheap, if you let someone try 20 times in an hour before banning an IP, there are targets that people will cycle through IPs that quickly for.
I've just posted top-level in this thread listing two projects of which I'm aware that provide this, though as experimental protocols only. I've been mildly agitating for further development of such tools. Looks as if CloudFlare are working in a similar direction, which I see as positive.
Nah, a human isn't going to waste their time refreshing a page manually 50,000 times.