Wikipedia’s Switch to HTTPS Has Successfully Fought Government Censorship
motherboard.vice.com
motherboard.vice.com
Wikipedia publishes database dumps every couple of days[1]. So it shouldn't be that expensive for smaller governments to create and host their own censored mirror. You'd maintain a list of banned and censored articles, then pull from wikipedia once a month. You'd have to check new articles by hand (maybe even all edits), but a lot of that should be easily automated, and if you only care about wikipedia in your native tongue (and it's not english) that's much less work.
The academics will bypass censorship anyway, since it's so easy[2], so an autocrat won't worry about intellectually crippling their country by banning wikipedia. Maybe they don't do this because the list of banned articles would be trivial to get.
Better machine translation might solve this by helping information flow freely[3]. We have until 2018 I guess.
[0] https://news.vice.com/story/china-is-recruiting-20000-people...
[1] https://dumps.wikimedia.org/backup-index.html
[2] https://www.wired.co.uk/article/china-great-firewall-censors...
[3] https://blogs.wsj.com/chinarealtime/2015/12/17/anti-wikipedi...
Chinese Wikipedia has 940,000 articles, baike has 6 million articles.
With 6x the articles on baike I can't imagine that there is that level of quality control. Unless there are 6x as many things worth documenting in China vs rest of the world.
An interesting statistic none-the-less.
If you see that somebody spins the Russian Wiki, you should definitely try to make it right to the extent suggested by Wikipedia norms.
As another exhibit: https://en.wikipedia.org/wiki/War_in_Donbass
This article in RU denies any involvement of Russia in Russian-Ukrainian war, however weird that may sound. They are either complicit or so deep in denial that it is impossible to talk to them about the war.
Currently Russian Wiki segment can't be trusted except for bare facts and non-political entries.
The russian version of that article is currently Киевская Русь [1] (Kievan Rus), though Дре́вняя Русь (Ancient Rus) is listed as a synonym. So it seems that specific change has been reverted, right?
[1] https://ru.wikipedia.org/wiki/%D0%9A%D0%B8%D0%B5%D0%B2%D1%81...
https://en.wikipedia.org/wiki/Wikipedia:Mirrors_and_forks/Ba...
I was using https://www.dotcom-tools.com/website-speed-test.aspx?se=1403....
My assumption is that because wikipedia has a known plaintext and a known link graph it's plausible to identify pages with some accuracy and either block them or monitor who's reading what.
I also assume that the traffic profile of editing looks different from viewing.
At least in theory, the latest versions of TLS should not be vulnerable to a known plaintext attack. TLS also is capable of length-padding, which would reduce the attack surface here as well for an eavesdropper.
My understanding is that HTTP/2 makes it even more difficult to construct an attack on this basis, because HTTP/2 means multiple requests can get rolled into one.
Of course, all this is assuming an eavesdropper without the ability to intercept and modify traffic. In practice, governments will probably just MITM the connection - we have precedent for governments abusing CAs like this in the past - and unless Wikipedia uses HPKP and we trust the initial connection and we trust that the HPKP reporting endpoint isn't blocked, then it's still possible to censor pages, without anybody else knowing[0].
[0] ie, the government censors will know, and the person who attempted to access the page will know, but neither Wikipedia nor the browser vendor would be able to detect the censorship automatically.
TLS1.3, which is still a draft, does have support for record-level padding, but I haven't seen any of the experimental deployments using it.
HTTP/2 does have support for padding, but again, it's not common to see it being used, at least not in the kind of sizes it would take to obscure content fingerprints.
Wikipedia is a particularly hard case for traffic analysis fingerprinting. First, the combination of page size and image sizes are just highly unique, even modulo large block/padding sizes. But more importantly, anyone can edit a wikipedia page, so if the size of a target page isn't unique, it's very easy to go ahead and edit it to make it so. It would take very large amounts of padding to defeat this.
So it's definitely possible to fingerprint which wikipedia someone is browsing. But it's probably not easy to block it; the fingerprint is only detectable after the page has been downloaded. So it's not very useful for censorship.
Well, it's detectable after the request has been made and Wikipedia sends the response. Assuming that a government has the capabilities to block delivery of that response (which they do), they can still implement censorship at this level, before the page reaches the end user.
If they routinely MITM connections they will quickly be found out, and the CA would be removed from browsers.
E.g. if every wikipedia page, plus all of the content it includes, came to exactly 10K, 20K, 30K, ... in size, then you could obscure what the user is reading.
In the event this was even tried, it would presumably be trivial to defeat with injection of random content somewhere in the server responses anyway. This of course all assumes we can trust the root certificate authority though :P
Though you shouldn't be compressing things over TLS. I think the only proper solution is to pad out all articles (and images) to the nearest 2kB or something so that you can't figure out the length (randomness can be thwarted by forcing refreshes).
Certificate pinning in the HTTPS client would mitigate TLS MITM (HPKP).
Last month there were some articles on the English Wikipedia about ISIS-Erdoğan (I don't care true or not). Then they have blocked all Wikipedia (all languages). Because they were unable to block those individual pages.
Fired up my VPN, accessed the page, thank you very much.
"The Net interprets censorship as damage and routes around it." - John Gilmore
Since the rest of the request, including the URL is hidden, governments and other malicious agents between you and the server cannot actually see what pages you're requesting directly. They can only see that you are accessing wikipedia.org and transmitting some data. You may still be somewhat vulnerable to timing attacks to try to identify what pages you're viewing, but censorship can't happen at the page level over HTTPS; you have to block the whole thing in one go.
Although countries like China, Thailand and Uzbekistan were still censoring part or all of Wikipedia by the time the researchers wrapped up their study
The top comment might be asking about the "were still censoring part" of the article.
I know that certain companies (like Google and Microsoft) will actively censor themselves to continue to operate within China, but I figured Wikipedia would be against that practice on principal. Now I'm curious as to how it's done.
https://googleblog.blogspot.com/2010/01/new-approach-to-chin...
If your internet traffic is going to flow through infrastructure that a curious government owns, then you'll know that they're monitoring the traffic but there is no way to keep them from seeing what you're doing.
That is, assuming you don't click away your browser's security warning.
https://security.stackexchange.com/questions/8145/does-https...
Or in other words it is vulnerable.
China can (and probably does) issue a certificate that all Chinese browsers must install, they can then do MITM https using their certificate to sign the new versions.
Companies do this routinely BTW. Since it's their equipment, it's considered just fine. (But be aware of it if you are using a company computer.)
I've never seen or heard of this (at least across all browsers), so I find this unlikely.
"On Friday, March 20th, we became aware of unauthorized digital certificates for several Google domains. The certificates were issued by an intermediate certificate authority apparently held by a company called MCS Holdings. This intermediate certificate was issued by CNNIC."
Although there have been lots of concerns about CNNIC, I don't believe that the Chinese government currently either (1) routinely uses CNNIC to perform MITMs for censorship or mass surveillance purposes, or (2) purports to require UAs to trust CNNIC or another Chinese root in order to be used by Chinese users. I'm happy to be corrected if someone knows otherwise.
What's more dangerous, and much more likely, is that they might use forged certificates against specific individuals for a short period of time, for example, to intercept login credentials. The attack will go unnoticed as long as they also block the corresponding HPKP reporting URL (if the targeted site uses HPKP at all).
Hopefully the risk for the attacker from the two kinds of attack are gradually converging, due to pinning and especially Certificate Transparency.
Lots of people travel in and out of China with all sorts of computing devices. China does care about the reputation of their root and of their highly profitable electronic exports.
https://en.greatfire.org/blog/2014/oct/china-collecting-appl...
Given the recent Shadow Brokers release of the NSA tools, it seems to me that this was not only possible, but probable (not necessarily with Wikipedia, but any website).
Cert Pinning or HPKP is one type of solution, but it's tricky to get right especially for a large site like wikipedia.
"In Turkey, Wikipedia articles about female genitals have been banned; Russia has censored articles about weed; in the UK, articles about German metal bands have been blocked..."
Critics of this plan argued that this move would just result in more
total censorship of Wikipedia and that access to some information
was better than no information at all
I'm no critic of this plan but I still don't understand why this wouldn't result in more total censorship. Someone explain please?There's a rather interesting analogy to be made with the GPL here. Critics argue that companies shy away from it because they cannot control it. Yet its entire goal is to not be controlled, and it draws its strength from the conviction that the body of GPL software is too useful to ignore. And again, that's self-fulfilling.
It takes courage, but it's important to know when you have the power to say "all of me, or none of me".
No, they don't. Critics point out that companies avoid it, and non-critics ascribe this avoidance to "can't control it", which is false, because nothing under a third-party copyright under any non-exclusive license can be controlled by the licensee, but businesses avoiding the GPL don't generally avoid all non-exclusive licenses.
For the increasing number of companies that do participate in the GPL ecosystem, they do so because the opportunity cost of not participating outweighs the concomitant behavioural constraints. This produces a strong network effect as GPL software gains contributors, making GPL software more useful.
Wikipedia's anti-censorship strategy is analogous in that the switch to HTTPS raised the opportunity cost of censorship to the loss of the entire Wikipedia "ecosystem", which for many regimes is more severe than the "cost" of not censoring. This too produces a network effect as Wikipedia gains more contributors, thus further increasing its value.
HTTPS encrypts the URL and the content, but does not mask the DNS lookup nor the server being connected to.
If they don't want their population to access a Wikipedia topic/article and can't block/determine if someone is accessing it, the easiest thing to do would be just block it right away. So why they won't do it?
(PS: I'm in no way in favor of censorship, I'm just trying to understand such mindset)
To the average citizen, it won't look much different than going to actual Wikipedia.
Last time I checked, Wikipedia had HSTS enabled. So trying to forge their DNS without also forging their SSL certificate would be equivalent to total censorship for anybody who has previously visited Wikipedia.
Certification negotiation happens before the GET request happens, which means that the "URL" (or, rather, everything after the domain) is encrypted.
You can also see some of this process with curl. So:
curl -vvv https://www.google.com/WireShark also provides a good visualization of the HTTPS negotiation process and the various layers of HTTPS requests and responses. It does take a lot more to figure out than telnet though.
HTTPs encrypts basically the whole protocol, this includes your request (the URL, your fingerprint -- e.g. browser, plugins installed, preferred languages) and the response (the content, type of the response (text, video, audio file), and some other not some important things).
What HTTPs does not encrypt is the domain and ip. The domain is leaked through DNS. DNSSec will not help either because it will not encrypt the DNS request. It rather signs it so that you can be sure it is authentic (not tempered with) but everyone can read it. This includes the wifi hotspot you use, your ISP, your government and anyone who tampers with the wires (theoretically even your neighbor and nearby people if you use mobile data since the connection from your device to your ISP is not really strong[1]).
Even if you would encrypt the DNS traffic (or you use just use the host's ip directly), the person who intercepts your traffic could just build a database with IP addresses that correspond to DNS entries (or do a reverse lookup, however, not every IP address has a reverse lookup configured to the domain you are visiting).
In wikipedia's example, this can still be pretty bad. For instance, if an oppressive government realizes that you visit wikipedia version of a particular language pretty frequently (compared to the rest of the population), they might make assumptions about you and profile you. When you visit the German wikipedia site, you are actually visiting de.wikipedia.org instead of en.wikipedia.org which can be intercepted and seen.
This gets worse for static file servers which serve different images at different subdomains (e.g. static512.domain.tld). So, if a DNS request is made to static523, static123, static721, and static132, an attack might be able to guess which article you are reading (or narrow down the choice) because their will not be many articles which have images served by those particular file servers. Thankfully wikipedia does not do that. Everything is served through upload.wikimedia.org but newpapers/forums, etc might not do that or they even have a unique domain for that article (e.g. embedded chart/video, which comes from a unique their party and is loaded automatically).
So all in all, HTTPs is pretty good but you still leave a lot of metadata (the DNS requests are just the tip of the iceberg) that can be used to learn a lot about you. If you want to be safe, use Tor or a VPN. If you use a VPN be aware that you just shift the trust from your current location to another one (so that the VPN provider, their ISP, and the government where the VPN server is located can read all those metadata, which might be not a big deal or even worse, depending where you actually life. Furthermore some VPNs have been known to be broken easily and your ISP/government still sees that you are using a VPN or even Tor).
[1]: One exception is LTE internet but you could still downgrade the connection to 3G or edge to intercept the domain
Currently also with SNI.
https://en.wikipedia.org/wiki/Server_Name_Indication
This is important for some censorship circumvention schemes and also because some people have suggested that encrypting SNI is useless because DNS leaks the hostname [however, not necessarily along the same network path!!], while some people have also suggested that encrypting DNS queries is useless because SNI leaks the hostname.
'...Time is the virgin killer. A kid comes into the world very naive, they lose that naiveness and then go into this life losing all of this getting into trouble. That was the basic idea about all of it' Different times... https://en.wikipedia.org/wiki/Virgin_Killer
AFAIK, whether or not the image is actually illegal under English law is somewhat unclear (the definition of "indecent" is rather woolly), though it's certainly a poor choice for an album cover.
Edit: "to its blacklist" -> "on"; added "a non-governmental organisation"
Anyway, the idea is unworkable as the user's client could simply lie about what URI it's going to send after the encrypted connection is setup.
It's not about making things easier for the censor. It's already easy. It's about making life easier for people who have to live with censorship (pretty much the entire world, I guess?).
> Anyway, the idea is unworkable as the user's client could simply lie about what URI it's going to send after the encrypted connection is setup.
Good server should response with error, I guess.
- Unlike the host, URIs are a property of the request, not the connection, so sending it as part of the connection handshake doesn't really make sense.
- Unlike the host, there is a very long history of putting secret things into the URI. Even if the extension is built with this in mind, the number of security breaches that will result is greater than zero, with probability one. That's probably not the correct price to pay for convenient censorship infrastructure.
Even when using SNI (the optional extension that sends the domain name in cleartext), the web server fully entitled to ignore it.
Any numbers/figures?