Securing Web Sites Made Them Less Accessible
meyerweb.com
meyerweb.com
The problem is, they're taking an experience that is fundamentally ridiculously heavy, and then spending thousands of man-hours trying to optimize it. No one even considers that maybe its the experience itself that is too heavy, and no optimizations can help that.
Youtube Home Page. Load it up, and you'll find its making over 200 requests, transferring megabytes of data. Google's most obvious solution: Lets speed up TLS, make each request go faster, lets invent new image and video compression algorithms to lower the size of each response, lets batch requests to reduce latency, technology, complexity, more code, more code.
No one actually takes a step back and asks if the Youtube home page should make 200 requests. What if it only made 20 requests? We gotta load some thumbnails, so there's bound to be a lot of requests there, but otherwise what the heck is all this JS?
TLS on one request isn't the problem. The problem is the hundreds of requests a typical website leans on.
Uncomfortable opinion: The only reason the internet has survived for so long is because of Moores Law. We've developed all of this technology and SDLC process in an era where another 20% jump in performance is just around the corner, so who cares if it's slow today. Yeah, that era is done. And we, as an industry, are completely fucked. Its not an understatement to say its a "back to the fundamentals" moment, and its going to cost us billions of collective dollars engineering for it.
In 10 years, the fixes made to existing websites will have been replaced with a new set of inefficiencies, but the benefits and drawbacks of today's physical infrastructure, protocols, and algorithms will endure.
The point I took from parent's comment is that we should consider not thinking this way, that it's not a fact set in stone. That individual sites can and should be built on better foundations that don't involve building our own transport layers
X can be a few things. Growth. Data. Analytics. Usage. Engagement. Security. But the point is the "at any cost" part, because we rarely consider the costs of the engineering decisions we make across any of these.
Over 30 years, we've been able to sustain this because there's always been "more" around the corner. Facebook and Snapchat can sustain "growth at any cost" because there's always more people in some (moving target) country that the service hasn't hit yet. JS, Python, or Ruby is used to support "development speed at any cost" because we'll always have faster computers next year. Google wants "more data at any cost" because there will always be new applications of AI to find value in it, there will always be advancements in storage technology to store it, and users will never care that we're slurping up so much.
But... none of this is true over a long enough timeframe. And when this idea of "the cost doesn't matter because we'll figure that out later" permeates the core of our tech stacks and people, pivoting that is ridiculously difficult. Especially when the negative changes are beginning to happen so quickly.
Don't get me wrong; we need the people working on optimizing TLS, new media compression, leaner HTTP request formats, all that. But I'm saying that we need more; we need a mind-shift in the way we build software and our software companies.
I don't know what that should look like, but I think what we'll discover is: Keep it simple, Abstraction layers kill performance, Making it work exactly how you want is less important than it working, working fast, and working for everyone regardless of physical, technical, or geographical disabilities, and Backward compatibility might just have to be sacrificed every once in a while.
* Websites are big and take a long time to download.
* They had previously solved this problem with a caching server but theirs broke with TLS.
The author is apparently unaware of the options they have.
* Run a proxy server that caches pages. Basically all software supports proxies. Secure and well understood. No PKI to manage. And you can allow access to the proxy with no TLS or get a public cert. Works very well with those old devices too.
* Run a HTTPS caching server and add the CA to the systems. Little more effort but transparent.
Their problem is a real one, and I don't think we have a good solution for solving it.
It'd be nice if their problems could be solved without requiring gross hacks or degraded security.
Squid supports HTTPS caching, the TLS connection is terminated at the proxy and the client gets a cert from a locally trusted CA. This is not very trivial to configure but the Squid Wiki should get you through it.
Do they really have the options you outlined? The author appears to dealing with a situation involving students in rural Uganda. It might not be as simple in that context as you suggest.
At the same time, I’d have to Google to figure it out, and I’d want to see if I could optimize for viewing new pages. I think optimizing new pages could also be really important for learning to code. Maybe mosh into a server with a fast, reliable connection and then use that server to access the web with lynx. The idea being that less data would need to be transferred between the mosh client and server than between the mosh server and website. I would suppose that every interaction like scrolling down would have some latency, but new page requests would have less latency. I’m not sure if bandwidth usage and latency would be good enough for their circumstances though.
Sure, basically everything supports proxies. But if you're not running a MiTM attack the proxy doesn't see much. The proxy server gets CONNECT <domain> and relays the request without ever even getting the full URI for the requested resource, so it certainly can't cache it.
* Run a HTTPS caching server and add the CA to the systems.
I think there's several challenges here:
- This CA empowered MiTM attack isn't just proxying Wikipedia and other pubic sites anymore. It's also proxying banking information, and other things that should be secure. We've meaningfully reduced the security for everyone in that lab.
- Where previously a computer playing the part of the proxy could just sit between the computers and the internet at large, allowing new machines to be added at will. Now every new computer will need some operating system & browser specific setup work. Machines that move around (laptops, rasberry pis are super portable) will end up with a variety of random certs on them.
Even with a local SSL caching server, we're still looking much slower pages. The high ping & packet loss means that the extra requests required to set up the connection could add seconds of latency to even loading a page.
TCP doesn't perform well with 5% packet loss. 50% packet loss coupled with the tremendous latency of the link makes it close to useless. The long and lonely signal path needs a link layer between the terminals and the sat. Unfortunately it's probably a bent pipe transponder.
I'm a librarian and I talk to hundreds of other librarians a year about technology and security and all this stuff. In my experience it is incredible rare to find a library that is running anything THAT old.
- wiki, old.reddit, most news sites, github ... - css and js are loaded from my tampermonkey scripts and i update them every once a half year or so so. Then in umatrix i block loading css and js from their servers so that only my own is loaded(also has benefits of using custom themes/fixes)
- google(youtube)/facebook sites are a major pita. You can use youtube-dl to download videos, you can even perform some basic search like `youtube-dl ytsearch5:keyword --get-title --get-description` but I haven't researched if there're any better youtube-alternative sites because on 64k it's unusable anyway. Otherwise using mobile apps instead of sites is the only option here because google changes these assets quite a lot and the compression/obfuscation changes the names of css classes
- use RSS (inoreader) as much as possible - RSS you can get all the updates and especially inoreader has a neat feature called "Load mobilized content" which only grabs the text from the site and sends it back - also using it for my youtube subscriptions
Is there was a really easy way of mimicking all the effects of this type of latency so I could periodically test the stuff I set up?
Also, if it is just HTTPS, then it is possible to proxy through something that downgrades the protocol, but it feels dirty.
Your browser's developer tools can simulate latency and constrained bandwidth, at least in Firefox and Chrome. Firefox instructions: https://developer.mozilla.org/en-US/docs/Tools/Responsive_De...
At a system level, Clumsy (https://jagt.github.io/clumsy/), Comcast (https://github.com/tylertreat/comcast), and Network Link Conditioner (https://nshipster.com/network-link-conditioner/) are relatively user-friendly and work at a lower level. Okay, Comcast isn't as user friendly, but it has a really cheeky name. Also, the GIF on Clumsy's homepage is brilliantly well-done.
Apparently Charles (https://www.charlesproxy.com/) and Fiddler (https://www.telerik.com/fiddler) can also simulate bad connections, if you're already using one of those tools.
> Also, if it is just HTTPS, then it is possible to proxy through something that downgrades the protocol, but it feels dirty.
Not necessarily. Consider HSTS, HPKP, Expect-CT, etc.
> Not necessarily. Consider HSTS, HPKP, Expect-CT, etc.
Yeah, I'm aware of those (and use them on my own site) but the reality is that the vast majority of content sites out there do not use HPKP. Even if they use HSTS, many do not pre-load and as a worst case scenario the MITM can just switch the domain to something like google.unsecure.com or something like that.
There's also the problem, generally, of one-size-fits-all security so far as websites are concerned: there's really very little content I receive that's specific to me,and much of that is Hacker News and a few other forum sites. The content itself is almost wholly public. But I cannot cache or otherwwise proxy this.
Locally, I've setup both Squid and Privoxy, mostly for shins and grits, but also to explore the use and viability of proxies these days.
Squid caches less than 10% of my traffic.
Privoxy can filter by hostname, but little within pages -- no path or content actions work for HTTPS URLs.
I've looked at the SSL options of each -- Privoxy seems a lost cause, but Squid looks as if it should be aable to MITM. TLS traffic, but I can't sort out how, or sensibly verify it. And I understand browsers will start screaming bloody murder if they detect this as well.
The notion of a trusted delegated proxy seems potentially useful. As with the author of the article, I'm wondering if there is any movement in developing HTTPS-friendly proxy tools, in a sane manner?
Well duh, ideally Squid without a TLS MITM would cache 0% of your traffic.
> And I understand browsers will start screaming bloody murder if they detect this as well.
Sorta, you need to configure your clients to accept your proxy's CA but browsers should be good after that.
I haven't really caffeinated yet so I may be missing something important here but this seems like a few hours worth of work?
That said, I'm sure the weight of typical pages will grow leaps and bounds too, as they've been doing.
Isn't signing and verification arguably just as complex as HTTPS? I'd argue ignore signing and verification entirely, and make it immutable.
That's why I like immutable designs like IPFS. No trust needed when viewing content, your client can verify the content is what you're requesting.
Even if it was a completely static site, I would still argue for HTTPS. It prevents ISPs from injecting tracking scripts or ads into the site.
It's not a good thing to have your entire web usage history recordable and attributable to you anyway. What requests were made should only be between you and the site in question.
When your spouse dies under suspicious circumstances, you don't want that one time you were idly curious about corpse disposal to be discoverable.
That I agree with but we're talking about people who have very marginal access to the internet already like the article is talking about where the benefits of allowing local caching while maintaining content security is a huge boon compared to the possibility of some tracking.
Can you clarify what you mean by "security overhead"? In 2018, there is virtually no performance difference between HTTP and HTTPS. For example, see performance evaluations and comparison overviews [1] and [2].
It's also very easy to set up HTTPS. If you have the technical capability to administer your own Apache or Nginx server, setting up HTTPS doesn't require much more configuration. Let's Encrypt is straightforward to use. If you're setting up a blog on WordPress or Squarespace, it's a matter of flicking a switch (or the choice has been conveniently made for you).
As for your idea regarding content digests, how do you propose to implement this? They need to be signed, sure, because an attacker intercepting my connection could just supply their own hash digest corresponding to the content. But then where do I obtain the public key corresponding to the site? How do I grab that key in such a way that I know it's the correct one? You'll find that any protocol you can device to safisfactorily resolve this problem will substantially reimplement TLS itself.
____________________________
To be fair, this is in conjunction with HTTP/2.
HTTPS is now relatively easy to implement and doesn't carry a significant performance overhead in most cases. The set of legitimate MITMs is vastly smaller than the set of malicious MITMs. It offers a multitude of benefits with very few downsides. There may be cases where HTTPS is of no benefit, but that's very difficult to accurately discern. It makes a great deal of sense for the major web players to push for HTTPS by default.