Enabling Secure HTTP for BBC Online
bbc.co.uk
bbc.co.uk
Not to be snarky, but haven't people written tools to help with this? This seems like a common issue. I mean, there's `sed` and similar tools, obviously, but something that could go, validate that the link works over https://, and update it. I don't see why that would need to be some monumental amount of work.
HTTPS is more than just privacy. See https://certsimple.com/blog/ssl-why-do-i-need-it and https://www.troyhunt.com/ssl-is-not-about-encryption/
AUGH! Seeing this "SSL is just for private things" mindset in 2016 is really disheartening. It's to keep people from screwing with your connection, not just snooping on it.
I really hope the browser vendors start treating HTTP the same way they treat broken certs sometime soon. This will change once users start asking, en masse, "Why am I getting all these warnings", not before.
Source: http://peter.sh/experiments/chromium-command-line-switches/
See:
--mark-insecure-asThey could put in place redirects, and then use HSTS to tell browsers to only visit the HTTPS links.
They could leave the old HTML unprocessed and pointing at HTTP and HSTS will fix it for modern browsers.
Only the first request would be via HTTP, and Chrome and other browsers can be told to use HTTPS when they see the links even then: https://hstspreload.appspot.com/
If the BBC "channels" stopped working, but other providers' content continues to work, the BBC would be blamed.
haven't people written tools to help with this?
Let's say you have a web page with a javascript slippy map that imports openlayers from a CDN; and openlayers then retrieves map tiles from openstreetmap.If you serve that page over https but the javascript CDN url is http, the javascript library won't load. And if the js CDN supports https and you switch to it, the library might still compose a http URL to retrieve the map tiles - causing some browsers to block the tiles as mixed content. Other browsers are willing to load http images on https pages and will work. Unless the tool understands how the map library composes its URLs, someone will have to fix this manually.
To detect bugs like that automatically, after changing to https you'd have to spider every page in your site with several different browsers / browser configurations looking for errors and bad links. And if your archived site had a bunch of errors and bad links to start with, you'll need some way to compare the before-and-after error reports too.
TLDR: It can be more complicated than you think.
[1] http://news.bbc.co.uk/nol/ukfs_news/hi/uk_politics/vote_2005...
Not as trivial as you'd think: if there's an HTTP URL on the page when it should be HTTPS, how did the URL end up there? Dynamically from PHP code? Dynamically from JavaScript code? Did the URL come from a database? Did the URL come from an environment variable? It can be a lot of work to track all these down and a lot of them you won't be able to find using grep/sed e.g. URLs might appear as relative URLs in code with the "http" part being added dynamically.
You'll get insecure content warnings as well if you try to load HTTP images, css, iframes or JavaScript on an HTTPS page. Likewise, the URL for these can come from lots of places.
I think this shows how valuable it is to use incentives to get people to Do The Right Thing(tm). Perhaps more things should be changed to require HTTPS.
I thought that it hasn't been significant overhead for a while now?
related: https://www.maxcdn.com/blog/ssl-performance-myth/ https://istlsfastyet.com/
The BBC has to deal with machines much older and much less powerful than that.
The servers shouldn't be running on old MacBook airs.
I'm sure the people using these machines that are "much older and much less powerful" than a 2012 macbook air are not expecting sites to load as fast as a newer machine, and probably don't care about the loss of less than 0.1 seconds to load time. If you're running a 6+ year old machine and expecting high performance you'd have to be insane.
Even if BBCOnline cared this intensely about performance, there are more than a few other things they could do to speed everything up. The switch from Apache to NGINX for one. I know that this takes many more developer/sysadmin hours, but if they really cared about a tens of milliseconds then it is definitely something they'd invest in. NGINX has quite a lot of support and is very stable, as well as generally known to much faster than Apache in most cases [1]. It's also not like NGINX is a hipster/unused server, it has quite a respectable share of the 'market' [2].
I also noticed on this page that they docwrite a script (probably to force it async?). This type of 'hack' is terrible for performance [3]. You could just add the 'async' attribute to the script tag and actually move it in the html and reduce the cycles wasted by a hacky solution.
[1]: https://www.rootusers.com/web-server-performance-benchmark/ [2]: http://news.netcraft.com/archives/2016/03/18/march-2016-web-... [3]: https://www.stevesouders.com/blog/2012/04/10/dont-docwrite-s...
What they HAVEN'T enabled is Diffie-Hellman Ephemeral suites, which give older clients forward secrecy at a big CPU hit.
So this is an example of performance-tuning your TLS settings. There's also stuff to do with session tickets, session resumption, and eventually they'd also be served using ECDSA certs, once all clients support it, or there is at least a great way to only show the older RSA cert to old clients.
ON a more serious note, I always use http://example.com. Being reserved and maintained by the IANA for documentation and testing, it's the most stable site I can think of.
or use what Google does when Chrome notifies you of a login gateway to public wifi: http://www.gstatic.com/generate_204
At least this will stop ISPs like BT from doing deep packet inspection and serving stale pages from their cache. Once it's been rolled out to the news site over the next year, of course.
If they use ChaCha-Poly then the load on low power devices shouldn't be much. I did a lot of reading on this for my recent book and it's pretty good for devices lacking hardware AES acceleration.
[^1]: https://unop.uk/block-bbc-breaking-news-on-all-devices
string(240) "https://ssl.bbc.co.uk/dna/api/comments/CommentsService.svc/V... string(40) "Error in cURL request: SSL connect error"
They probably have more websites and hostnames than you'd realize, as well. Take a look at a site like https://dnsdumpster.com/ and search for bbc.co.uk
For each individual product, they need to figure out what modifications it needs to become HTTPS-enabled (lots of links and identifiers are hard-coded to HTTP, and third-party CDNs might not support HTTPS by default), and update their testing procedures to ensure that it remains HTTPS-compatible, before they can enable HTTPS. Given that this is the BBC (a publicly-funded entity), they also have to ensure that everything continues to be fully supported on browsers going back to IE6, Firefox 3, and Safari 3 - with partial support for some browsers older than them.
In my opinion, a year is doing pretty well.
A blog post about spending several years updating to a protocol that's been around for 2 decades and has been standard for full sites for years. This makes me feel like anyone who has an account on BBC should be afraid of their security practices. Calling a plaintext password leak from BBC right now.
EDIT: People are taking this comment more seriously than I intended. I don't actually think you should distrust BBC's security practices because of this, but I do feel that major websites should have side-wide SSL by now. It is clear that a lot of people below me disagree with that, that's okay, I'm glad I spawned a debate here.
It hasn't been "standard on full sites for years", and still isn't now. Only recently with the 'HTTPS everywhere' move has the idea that public sites with no authentication should support HTTPS. And even now, that's not a universally supported opinion, because of its effect on caching.
The BBC has used HTTPS on pages with forms that submit secure data, as has been the historic standard.
Moving a site as massive as the BBC, which spans multiple domains and subdomains and has millions of pages is a big task. Note how you can still see news articles from the late 90s at the same URL. So, yeah, I can understand why writing a blog post about it is worthwhile.
https://www.google.com/transparencyreport/https/grid/
For example, the following are all in the world's top 100 websites and none of them support any form of HTTPS. The link includes quite a few more.
* alibaba.com
* ask.com
* ask.fm
* baidu.com
* cnet.com
* cnn.com
* dailymail.co.uk
* ebay.com
* globo.com
* go.com
* goal.com
* goo.ne.jp
* imdb.com
* live.com
* mirror.co.uk
* naver.jp
* nytimes.com
* onet.pl
* pornhub.com
* telegraph.co.uk
* uol.com.br
* weibo.com
* wikia.com
* wikihow.com
* wp.pl
* yahoo.co.jp
* yelp.com
* youporn.com
New York Times still dosent have HTTPs.
Instead, the burdens are on testing and developing the migration. For example, they'd have to inventory and edit everywhere they use http:// (hardcoding the scheme in your front-end code) instead of //. Furthermore they have to support third-party ad networks deliver active scripts (like javascript) over HTTP. Having HTTP while on HTTPS will create mixed content warning and for active contents browsers will block these violations immediately, thus breaking the website.
To me, the decision of not migrating to HTTPS because of infrastructure capacity is always a myth. Someone has to prove that with data.
I'd estimate about 75% of the time I'm on an HTTPS website.
> The BBC has used HTTPS on pages with forms that submit secure data, as has been the historic standard.
This is insecure as the HTTP page can redirect to a malicious HTTPS page from a different domain.