Enforcing TLS protocol invariants by rolling versions every six weeks
ietf.org
ietf.org
The idea is to catch them early by simulating future versions now, which in their belief is best done by perpetuating traffic that contains bullshit versions on the wire. It's fairly clever, in that any implementer who doesn't want to break significant portions of web traffic now must code their implementation in a way that anticipates Chrome's future versions, which is most sanely done by coding their implementation in a way that happens to match what's in the spec. Meanwhile, noncompliant implementations will likely break for a huge amount of users, causing widespread uproar and, presumably, pressure to fix the problem.
Google's clients (e.g. Chrome, Android) and servers originate and receive a high enough portion of traffic that they could pull this off on their own without anyone else's buy-in, and still achieve the intended effect, but they're reaching out to the community to build consensus. It's a hacky idea, but one that backs noncompliant implementations into a corner now, instead of them sticking around forever and causing incompatible changes later down the road.
I guess we just have to relearn all the lesson from the 80's
I sure hope that's not a lesson we learn in any decade.
Importantly, the big reason for TLS to be at the endpoint and not use IPsec/tcpcrypt/etc. is because the important thing to verify in TLS is the hostname, not the IP address, and DNS resolution is in userspace everywhere (as far as I know). You could imagine a design which pushed both DNS resolution and TLS verification into the kernel, but it would be too monolithic of a kernel for even UNIX's tastes.
TCP is a relatively simple protocol (in the sense that it doesn't have too many moving parts, not in the sense that it's uninteresting) and fits fine in the kernel. But even for reliable transport, there are plenty of apps using their own thing over UDP (RTP, Mosh's SSP, QUIC, etc.) that are never going to get into a kernel. And even for TCP alone there's a bunch of potential complexity you can put into the kernel (DCTCP, New Reno, fq_codel, etc.) that have been hard to deploy precisely because they require a kernel patch.
if you keep conflating responsibilities, you'll keep ending up with the same, flawed, "solutions".
besides, everyone and their dog now has a "certificate" from the likes of let's encrypt, so even their role as identity trust is basically moot.
What's the point of encrypting something if you don't know who you're encrypting to?
> besides, everyone and their dog now has a "certificate" from the likes of let's encrypt, so even their role as identity trust is basically moot.
Yes, everyone should have a certificate, and I'm not sure why you think this means "identity trust" isn't working. Certificates mean that you probably are who you say you are. They don't mean that you're a good person.
https://datatracker.ietf.org/meeting/87/materials/slides-87-...
https://shader.kaist.edu/mtcp/
https://github.com/google/netstack
The first was BLACKER VPN. NSA has their Type 1 Link Encryptors. Mikro-SINA put it in on Intels with more open components. Now, I'm pushing for Wireguard variants in both app and network deployments. There's people laying groundwork for expanded use of it right now. Once that's done, I know some people who might put it in some low-privilege architectures.
Very few developers (outside of Google) are going to write randomized fuzzers to test for compatibility with theoretical future extensions. They're going to test that it works with existing stuff and then ship it.
So, the fuzzing has got to become part of the public ecosystem.
This is about basic programmer competence, not time a consuming feature that might impact your development costs relative to your competitor. You are not going to make more profit by leaving out the "default:" case to your switch/case statements that skips parsing for unrecognized elements.
[1] https://archive.org/details/The_Science_of_Insecurity_
[2] https://media.ccc.de/v/31c3_-_5930_-_en_-_saal_6_-_201412291...
default:
return DROP_CONNNECTION_AND_BAIL; for (item = params->head; item; item = item->next) {
switch (item->type) {
case KNOWN_PARAM_TYPE_FOO:
// do normal stuff
break;
/* ... etc ... */
IGNORE_KNOWN_PARAM_TYPE_BAR:
// fallthrough - BAR explicitly uses default handling
default:
continue; // skip unknown parameters
}
}
is evidence of incompetence, not a strategy that will "make more profit than your competitor".Also, as the BAR constant suggests, you probably already have code that skips unrelated fields. While the difference in programmer time is almost always trivially small, sometimes it might be zero.
Until then, you have to make the financial incentives in the short and long term such that they lead to desirable behavior, e.g., producing non-barfing middleware in this case.
Then it ends up in production in too many customer sites with IT that avoid updating stuff (because the updates frequently break things, which is a pain) that it's too late to fix it now. The concerned forward-thinking developer moves on to (slightly) greener pastures to rinse and repeat and the project gets handed to an off-shore team who don't comprehend the difference between current behavior and specified behavior of a product - yet another bold cost-saving measure for pointy-hair to use to argue for his bonus.
Didn't we learn that from icmp blocking break path MTU discovery already?
Clever! I like it.
There were probably some problems with odd dumb custom-coded gadgets. But mostly, the problem is with "value-add extra-security" boxes in the middle that don't _do_ the protocol at all. They are not a TLS client or a TLS server. They are a product which promises to spy on the traffic and "stop the hackers". They really can't do anything useful. They just see that the handshake packet had (1,3) instead of the familiar (1,2) and drop it. For good measure, they also block packets between those endpoints for 24 hours (so fallback retries don't work). Also they didn't add (1,3) to their recognition system a couple years ago when tls1.3 was in the works, because anyone buying or selling these things is just not very on top of things. So, all the middleboxes which are popular enough to matter, need to be tricked with obfuscation, so they're _really_ not doing anything.
This was done because the version code ploxiln describes is so thoroughly rusted shut in middleboxes that TLS 1.3 would be undeployable in practice without such a changes.
Also, "fallback retries" aren't a thing. The reason is downgrade protection. If I have a new protocol (TLS 1.3) with better security, but I'm happy to retry with an older protocol (say TLS 1.2) if that doesn't work, obviously bad guys on the path will just ensure my TLS 1.3 connections all fail so that they can attack the weaker TLS 1.2
TLS 1.3 guards against an attacker who tries to downgrade a TLS 1.3 connection, if both sides know TLS 1.3 and yet somehow the packets when they arrive from the client say TLS 1.2 on them, the Hello "random" value sent from the server will have the message "DOWNGRD" scribbled across part of it, and the client sees this and aborts because somebody is tampering with the connection. If the middlebox tries overwriting the bytes with "DOWNGRD" written in them then the random data doesn't match up and the connection fails.
If google deploys 1.3 and 50% of corporate users can no longer reach google's services, that would be seen as google's mistake. Moreover, that would hurt google quite a lot.
The issue is that these boxes are already deployed, and it wasn't noticed just how shitty they were until we tried to deploy 1.3 . The system has essentially rusted shut.
Raymond Chen wrote some articles about how this impacted Windows 95. It doesn't matter that program X completely ignored the documentation in Windows 3.x and so it "makes sense" technically that in Win95 that program crashes, the user experience is that Windows 95 broke Program X, and they just bought Windows 95, so they will demand a refund and moan to all their friends.
There is a limited tolerance for Chrome versions that don't work, because the lesson customers get is "Don't run Chrome" not "My middlebox is garbage". The tolerance is increased for security problems, so Google is more willing to lose say 0.1% of users because they enabled TLS 1.3 (the remaining 99.9% of users get improved security) than to lose 0.1% of users because they added a cool 3D logo and it crashes on a specific model of video card or version of Windows due to a driver bug. But losing 10% of your users is a disaster, and that was the ballpark for TLS 1.3 in earlier drafts (before it was taught to sidestep more middleboxes).
Hopefully someone will notice and bring it up with the IT or ISP, especially if the message says to do so, and that Google/internet could stop working in the future.
Actually,get more popular sites to do it, like Facebook, Twitter, Apple and Microsoft (heck, they could do it as a part of the operating system): if a lot of websites say it, users will tend to think that something is wrong on their end.
This was surely already brought up as a solution, so I wonder what the catch was, if any?
Since this would be browser-independent, the browser wouldn't get blamed, and if it's only a mild inconvenience, it shouldn't bother people that much (they have been using the web despite more invasive cookie notices). I would expect it to allow a critical number of non conformant devices to be quickly disabled as a result.
As far as I know, Javascript doesn't have low-level access to connections, not enough to run even a stub TLS 1.3 implementation.
As you say, "fallback retries" are bad. But that is what browsers did, some years ago, when tls1.2 was less common ... and they added the TLS_FALLBACK_SCSV "not-a-cipher-suite" to try to detect attacker-caused downgrades in a dumb-server-compatible way ...
https://blogs.msdn.microsoft.com/oldnewthing/20040211-00/?p=...
Summary: a video driver cheated by having its implementation of the "do you support this DirectX feature" API always return true no matter what feature was asked for (and it didn't support everything, obviously). This made things crash and led to Microsoft (who didn't write the driver) getting complaints.
The solution, since DirectX features used GUIDs as their identifiers, was to take the MAC address of a new network card, use it to generate one GUID, then smash the card. Since they knew that MAC would never generate another GUID, they put in a check that would ask video drivers "do you support the <GUID from smashed network card> feature?" And if the driver claimed to support that feature, DirectX would know not to trust the driver's claims of feature support.
Sometimes these things are required by law and/or industry standards. IIRC banking sector companies are required by law to record every communication metadata of the employees... which only works with said middleboxes.
https://community.letsencrypt.org/t/adding-random-entries-to...
(Let's Encrypt creates random protocol extensions on every connection in order to ensure that clients are tolerant of protocol extensions that they don't understand. Breaking changes to the protocol, on the other hand, will be served from a separate API endpoint.)
The famous "Alice and Bob After Dinner Speech" mentions
"Now most people in Alice's position would give up. Not Alice. She has courage which can only be described as awesome. Against all odds, over a noisy telephone line, tapped by the tax authorities and the secret police, Alice will happily attempt, with someone she doesn't trust, whom she cannot hear clearly, and who is probably someone else, to fiddle her tax returns and to organize a coup d'etat, while at the same time minimizing the cost of the phone call.
A coding theorist is someone who doesn't think Alice is crazy."
HTTPS is like Alice. In cryptography a theoretical attacker is often given seemingly outrageous abilities, like they can send you huge numbers of arbitrary messages to see what happens, they can time everything, they can see messages you were sending and try sending other messages that are just a tiny bit different, they can collect your messages and re-send them later, and so on. In many systems a real attacker would struggle to pull these things off, but in HTTPS thanks to things like cookies and Javascript it's actually not difficult at all.
Your internal stuff almost certainly doesn't have arbitrary clients running code from arbitrary other participants like the Web does. It also almost certainly doesn't have a dedicated reliability team who can go change everything every six weeks to keep up. If you do such changes every six weeks for a few months, then get bored and stop, the last set rust shut and you've gained nothing. Google is essentially promising their teams would undertake to carry on indefinitely.
Google essentially proposes an artificial Red Queen's Race, with the goal being to tire out middlebox vendors and/or their customers and have them choose to exit the race.
[1] http://hg.openjdk.java.net/jdk/jdk/file/f36d08a3e700/src/jav...
1. So called "Security" companies (middlebox vendors) advise customers to write what are effectively firewall rules that bake ossification into their systems using, I kid you not, regular expressions
https://www.fidelissecurity.com/threatgeek/2018/02/exposing-...
This type of nonsense is why one of the optional TLS 1.3 features is certificate compression. It seems like a no-brainer to offer this for earlier versions, but it turns out that middleboxes snoop the certificate and make decisions about it so compressing it causes them to freak out. Why can we (hopefully) fix that with an optional extension in TLS 1.3? Because in TLS 1.3 now the certificate is encrypted, so the middleboxes can't see it in the first place.
2. In about 2016 when it originally looked like TLS 1.3 was almost finished, the middlebox vendors finally noticed something was happening and began yelling about how their products had "legitimate" ‡ uses that would be blocked by these improvements and it all needed re-thinking.
Fortunately (or perhaps inevitably) the middlebox vendors have no idea how the IETF works, so they spent a lot of effort on trying to "win votes" which should have anybody who was somehow unaware of this drama but involved with the IETF smiling since there are no votes and you can't win. They also, like typical business people, figured they could fly in to meetings in say, Singapore or London, and spin up at the meeting, but of course those meetings are just a temporary physical incarnation of the IETF, it exists all the time as mailing lists, so if you aren't following those lists you're basically always a kid who wandered into a room where people are having grown-up discussions you can't understand.
3. After a surprisingly long time the middlebox vendors got the hint and went to ETSI to make their own alternative. ETSI is a traditional SDO which is perfectly happy to work on a standard without any messy ethical considerations. It is also, like most traditional SDOs, closed door, so we have only limited visibility of what they're up to. So far it looks like TLS 1.2 (so, bad) but with the ability for an arbitrary number of middleboxes to interact with all the packets on path (so, worse).
It's OK though because under their ETSI proposal the user will "Consent" to this. If you've spent the last month mindlessly clicking through GDPR-inspired boxes on US web sites you regularly visit, or indeed if you're the new hire who has just this moment realised why the big boss decided he needed her with him on this trip, and is now calculating whether to say "No" and risk losing everything or to close her eyes and pretend this is happening to somebody else you have a very good idea what "Consent" means in this context.
‡ "Legitimate" is a word you use when it's important for people to accept that what you're doing is OK, without them thinking about what you were actually doing because they might throw up. "Bribe officials to ignore flagrant safety violations" makes you angry but "Legitimate facilitation payment" sounds very respectable.
(edit: or at least will start getting bug reports as soon as a Chrome/Chromium user enters the population)
First, it's coercive. Second, it's in principle spam, even if the cost is affordable in this case. Third, it sets three bad precedents:
* Google can do whatever they want
* Hijacking user computers for ulterior purposes (by utilizing Chrome) is ok
* Spam and coercion are ok. And if you argue it's ok in this instance, you aren't thinking past your nose.
Imagine if Symantec or Microsoft tried it; how would people react? Well you don't have to wonder, because they will.
Good question.
No, I wouldn't, because it serves that user directly. Google could argue that the TLS 'GREASE' benefits users, but that benefit, and let's assume it exists, is very indirect: That user won't see any benefit that day, that week, and maybe never; Google is using their users' computers to advocate something that Google thinks is a good idea. To demonstrate GREASE can be taken too far, imagine the extreme case where Google has Chrome send messages to U.S. Congresspeople advocating against some policy - Google could argue that they believe it's in the users' interests, but that would be highly disingenuous. Again, that's extreme and not going to happen.
More realistically, imagine Microsoft had Windows 10 insert something in network traffic to compel compatibility with some proprietary technology of theirs. Imagine vendors disagreed about some standard, and different products used GREASE to compete for different outcomes.
Where do you draw the line? I think GREASE violates end-user control or autonomy. They become pawns in a standards competition. It's very arrogant of Google and support of it betrays an arrogance perspective here at HN: users are pawns, and their computers are ours to use as we see fit. (It's also intentionally bad engineering, which makes me very uncomfortable.)
I will benefit from this even though I don't use Chrome (or other Google product/services). Reducing traffic manipulation by middleboxes and eliminating passive traffic snooping are important goals, and preserving forward compatibility should increase interoperability with future versions. I benefit from the internet moving back towards the end-to-end principle (smart hosts connected to a dumb network that only routes packets).
The additional bandwidth cost is trivial and by definition will not affect anything that follows the spec.
> imagine the extreme case ... Again, that's extreme and not going to happen.
Speculating about hypothetical problems you admit are not going to happen is a distraction and waste of time. It's possible to make up extreme hypothetical problems afflicting any plan.
> imagine Microsoft had Windows 10 insert something in network traffic to compel compatibility
I spent a lot of time in the late-90s/early-00s fighting against Microsoft (and others) trying to embrace, extend, and extinguish[1] open standards. That strategy relies on extending an existing protocol with features that are not supported by existing implementations. The goal is to create the wall around a public garden by creating interoperability problems.
The only people that might be affected are middleboxes that want to manipulate traffic (good; that was the goal) and incompetent implementations of the spec. The latter probably needs to be fixed (or replaced) anyway because not following the spec that is important for security is a sign that the software probably has other serious bugs and vulnerabilities.
I have spent decades fighting for Free Software and against proprietary control of data formats, which are usually a form of rent seeking. Encouraging implementations to follow the spec and ignore unknown elements is unrelated to those concerns.
> different products used GREASE to compete for different outcomes.
How, precisely, is that supposed to work when (according to the RFC[2]), "Servers MUST correctly ignore unknown values in a ClientHello and attempt to negotiate with one of the remaining parameters."
> It's also intentionally bad engineering
This is very good engineering. The software industry has been negligent in learning good engineering practices like the importance of designing in tolerance[3] and failing safely.
[1] https://en.wikipedia.org/wiki/Embrace,_extend,_and_extinguis...
[2] https://tools.ietf.org/html/draft-ietf-tls-grease-01#section...
Sending messages to Congress would be very different - that's squarely outside of the purview of a browser. TLS is not.
As for Microsoft having Windows 10 insert things to support a proprietary technology, again, this is a flawed analogy - it's simply not what's happening here.
Why draw the line if we know that we haven't crossed it? You have put forth two cases that seem obviously on the other side. This case seems obviously on the 'we good' side.
Would you say A/B testing means 'users are pawns'? I really disagree with that. Again, this is all in the interest of end users who have voluntarily installed this software, at least in part because of the secure reputation.
in this instance, they're doing a good thing. are you arguing that google shouldn't ever do anything, no matter how good it is, because if google is allowed to do things then eventually they might do bad things? that seems absurdly defeatist.
The day that a programmer can insert additional data as "optional fields" in TLS that can't be inspected without breaking spec? Why can't they be inspected? Because they're "unsupported"
Sounds like a great plan. Definitely a great plan. What could go wrong with clients writing out more data than they need to? What could go wrong with clients writing out "garbage data" specifically to make sure that random unsupported data is supported and "ignored"? Nothing wrong at all! /s