Preparing Rustls for Wider Adoption
abetterinternet.org
abetterinternet.org
Anyway, this sounds great to me. I'm not a huge fan of memory unsafe languages, especially for critical code.
I find it shameful that letsencrypt issues statements like this one:
"we hope to replace the use of OpenSSL and other unsafe TLS libraries in use at Let’s Encrypt with Rustls."
OpenSSL has made companies billions (trillions?) of dollars, now they drop it like a hot potato.
1. Making money.
2. Making backwards incompatible changes to products and services in the name of best practices.
Instead NIH solutions are now popping up after having exploited OpenSSL for more than two decades.
> Instead NIH solutions are now popping up after having exploited OpenSSL for more than two decades.
Building new crypto libraries to replace OpenSSL is not caused by NIH, it's caused by rationally examining the existing codebase and coming to the conclusion that sometimes starting from scratch is the sane choice.
Google is doing the incremental fixing approach (in BoringSSL) too, but it's very clearly a stopgap.
1. The EverCrypt primitives are formally proven, whereas ring has no such formal proofs. Also it seems that all the EverCrypt primitives have portable (non-assembly) fallbacks, while ring has several primitives that are assembly only and have no portable fallback.
2. MiTLS is written in F#, which is harder to integrate with other languages than Rust.
Ring uses assembly to make sure it can do constant-time operations, to avoid leaking information through computation time. What does EverCrypt use for constant-time operations?
AFAIK Rust also doesn't guarantee timing. That's part of why Ring uses assembly for the constant-time bits (really Ring uses code from BoringSSL for that, and BoringSSL uses assembly for that reason and for performance.)
(This sort of thing is not my area of expertise, but has always made me uneasy...)
Especially note his guideline to look at the assembly output. The best guarantee is to simply check that the compiler did what you wanted.
There are also some aspects that a compiler is unlikely to alter. EG if you write code with no secret-dependent branches, no normal compiler will insert any secret-dependent branches during optimization. It's generally safe to rely on the compiler not being actively malicious or de-optimizing code.
The lack of a guarantee in assembly is less of a concern than the lack of a guarantee in C, since the CPU behavior is less likely to change unexpectedly than a C compiler.
Though Intel at least does provide a list of instruction latencies and throughputs, in their "Intel 64 and IA-32 Architectures Optimization Reference Manual"[1].
[1] https://www.intel.com/content/dam/www/public/us/en/documents...
So, you don't have to know how many cycles every instruction takes; you just have to make sure you use instructions that don't take a data-dependent number of cycles.
Personally, I would love to see enhancements to Rust and LLVM that would make it possible to provide such guarantees in pure Rust.
Expressing most algorithms this way is not impossible, but does nor feel natural for many.
It fundamentally isn't, because there is no way (in standard C) to prevent the compiler from introducing performance differences. Even if you avoid control flow constructions, what if the compiler e.g. generates a separate codepath for when one particular variable is 0 (perhaps as part of arithmetic optimizations)? There's just no (standard) way to stop that happening, no matter how cleverly you write your code.
Jokes aside, I wonder if some kind of -Ox option could be standardized to switch off all optimizations, like -O0, and also guarantee no code rewriting for whatever other purposes.
Hand-crafting constant-time algorithms in assembly is even more error-prone.
for the guaranteeing const execution time aspects. More or less any language compiled by LLVM, GCC or similar optimizing compilers can't do so. (And even wrt. assembly the cpus could do optimizations which also brake this but it's less likely).
Same is true for any language using a JIT with optimizations like e.g. Java, JavaScript, C#, etc.
Practically the degree of how much compilers could mess up const execution time guarantes is language dependent but for many languages especially including such using compilers created for C/C++ basically you have no useful guarantees.
At best you can create some code which just with the current version happen to not be optimized in a bad way. But this means you will have to review the produced assemply code and do it again and again every time the compiler (or just surrounding code) changes... To make it worse the high level code is force to use all kind of tricks to prevent the compiler from doing certain optimizations, which often obfuscate the intent of the code.
So all in all having a small very well reviewed set of assembly snippets for the core primitives the the easier to review and more reliable way in many cases. (Note that "a small set" means that just a few primitives are assembly only, not all the crypto, which normally is good enough).
Through best would be to have some language which focuses on both making it reasonable to implement such primitives in a more high level language and has tooling for proving code correctness and also code-to-assembly transformation correctness (e.g. see what SeL4 did wrt. proving assembly).
We do use some of the Fiat Crypto stuff for elliptic curve computations. I am not opposed to switching some stuff to use EverCrypt or other things that might be better, as long as the performance is the same or better, and as long as there's a clear path towards that code being in Rust.
> ring has several primitives that are assembly only and have no portable fallback.
Either in the latest release (0.16.20) or the upcoming release, there is a portable non-assembly implementation of everything in ring.
> The stable version of miTLS including the new 0.9 release are written in F#
With F# being a link to https://fsharp.org/
Isn't rustls [1] also built on very unsafe groundwork? Namely ring [2], which, according to github, contains 47.3% Assembly and some C as well.
I'm not trolling here - we were discussing this a lot in my peer group lately.
[1] https://github.com/ctz/rustls [2] https://github.com/briansmith/ring
Main reason I stopped using it.
Now for all I know, ring has never fixed any bugs and it just loves adding new API features so that this pruning has no desirable security properties at all, but in principle I can see that this is the equivalent of the standard boilerplate Linux release text which tells you that you should update to the latest kernel because they fixed bugs.
If you have a complete threat model and if you are capable of the insight needed to examine all changes and determine how they impact that model, you could successfully choose whether to upgrade based on whether a new version fixes a bug you care about. But chances are you don't have such a model and even if you did you aren't capable of the inhuman levels of insight needed, even in a language like Rust (and forgetting that we're talking about this because large parts of ring aren't even in Rust).
If ring wants to notify me that I should update, they should send an email to a security mailing list, open a CVE, register the cve in any of the rust services to notify users with those dependencies (there are some, like crev), etc.
Pruning your releases from crates.io just means that I am going to be annoyed the first time it happens, will start looking for a solution the second time it happens, and it won't happen a third time (and it didn't). If you want to wake me up a Saturday at 4 am, the world better be on fire.
This is probably the only dependency I can remember as being... more than annoying, toxic. I still prevent any of my dependencies from ending up with ring as a dependency. If that shows up in our dependency tree, CI fails, and that change cannot be committed. Unfortunately, this pruning of old releases was only one of the issues with ring (there were others, like cross-compiling it wasn't easy, etc.). All in all it was a no brainer to drop it as a dependency.
I don't think I've ever met a rust dev with something nice to say about `ring`. In a meetup a couple of years ago another rustacean said: "`ring` is so secure that it protects you from using it in your projects". Sums it pretty well.
The library has couple of thousands of daily downloads so for the latest version, and like 25k daily downloads for other versions, so maybe things changed now.
In general, my initial thinking was based too much on the assumption that people would help maintain the things that depend on ring to update them to the latest release. It turns out there's less cooperative maintenance like that than I expected.
This is why projects like https://github.com/RustSec exist.
Bugs? No. Security bugs? Yes.
Not yanked.
Vulnerable to https://github.com/RustSec/advisory-db/blob/main/crates/hype...
I think that's a reasonable assumption.
What isn't reasonable is to expect people to "upgrade right now". Not everybody lives in your time zone, so when you yank a dependency, you might be breaking a workflow in the other part of the world at 3 am, and if some webserver doesn't deploy or whatever, somebody will get a call.
I'm not suggesting this is an easy problem to solve, but there is a wide range of options before "never upgrading" and "force an upgrade right now". Some of these are supported by Cargo via Cargo.lock, etc. so the responsibility for how this is handled doesn't fall on one library or person.
Building a secure system is also not the exclusive responsibility of `ring`. If I'm building a secure system, I have to assume that `ring` will have a bug that's exploitable at some point, and that someone will use it in a zero-day, and my system needs to be secure even if that happens.
So "updating right now" doesn't buy me much. Its something that can wait until after a meeting, or after my vacation, or until monday. Its not a "the world is on fire" situation, even though it would still be pretty severe.
I'd still like to get "notified" ASAP and asynchronously somehow. While updating ring is low effort, the update still needs to "internal QA,..." etc. at companies, and that takes people's time that must be planned on.
I'm not sure if you were affected by this, but Cargo introduced a (regression) bug a couple years ago that caused it to fail when a crate got yanked when it shouldn't have. This bug was eventually fixed, but lots of people blamed ring for this bug. If this Cargo bug hadn't been introduced then most people who were using Cargo correctly wouldn't have been negatively affected by ring's old policy.
In a previous role I actually have had things set so that I might be woken at 4am on a Saturday, IIRC specifically under certain conditions it'd play "Straight out of Compton" at full volume on my Hi Fi to ensure my attention, which gave me about 10 seconds before it gets real loud.
I was leak hunting, specifically looking for a huge leak in a production system that we couldn't reproduce on smaller test systems - so I needed to wake up, attempt to diagnose the leak and then (regardless of whether successful or not) mitigate it (kill the bloated process, one transaction fails but everything else will auto-repair) and go back to bed.
But what I don't understand here is, why are you constantly rebuilding and alerting on failure? A CI flag can wait until Monday stand-up, are you auto-deploying any change of state even when there aren't any humans around to cause that? Why? That strikes me as up there with Apache's "But your Good OCSP response expires in 18 hours, so I stapled this newer Bad one instead" in terms of terrible mistakes.
If instead your new builds fail after pruning, the human who is causing a new build can decide what to do about that when it happens, no need for anybody to be woken at 4am.
[1] As always there are exceptions to this rule
So you have a team (i suppose could be just 1) or teams across few time zones.
Each team should deploy in a way that they have people ready to fix thing if they go bad.
You can do that by having large enough team in similar timezomes (+/- 2-3 hr), or by paying extra for having people on standby at 1am.
My personal opinion is that having large enough team (again could just be 1 guy) in similar time zone is preferable.
Bottom line is, deploying new stuff is always risky. That's why people spend so much time trying to reduce this risk (various test, CI, staged rollout ...). And sometimes all of that still fails, and you need people to either rollback or fix it on the spot.
You can yank crate versions, but that doesn't affect builds which lock a version via cargo.lock so failing builds shouldn't be an issue.
Also, the goal is definitely to bring all of that code into Rust, unfortunately Rust lacked the features to do that safely (things like const generics).
The security audit referenced by the post suggests offering EverCrypt as an optional alternative to ring. Are you going to act on the audit's recommendation, or continue to only offer ring?
Depending on what you mean by "groundwork" literally everything is. Hardware doesn't obey Rust's rules, and you need to interface with hardware to get input, and do output, so literally every program will have unsafe code at the base.
The key difference is that Rust gives you the tools to explicitly demarcate what is safe, and what is not, and build safe abstractions on top of (hopefully validated) unsafe foundations.
Neither does the OS where rustls is running.
I think Rust will have more adoption and more libraries like Rustls will be developed. I also think that when this happens, also more exploits targeting Rust code will exist too. I guess the excuse (sorry for using this word) will be: "In fact, the Rust code is still safe. What happened is that a pointer returned by (or used in) an underlying C library got messed up with a very clever timing attack, and somehow the pointer emerged into Rust code... etc.".
Now there definitely are some tricky requirements in crypto code that application code doesn't need to deal with, like constant-time requirements. But auditing for those isn't really any harder in assembly or C than it is in Rust. In the end, porting these sorts of core crypto algorithms from C to Rust tends to be more interesting from a build systems and tooling perspective than from a correctness perspective.
The assembly code in ring is some of the most heavily-tested code in the world. It's fuzzed pretty much continuously in various projects that use it, and a bunch of testing has been done on it. It is from BoringSSL, and much of it is shared with OpenSSL and/or Linux kernel. As we are able to replace the assembly code with safer code, we'll continue to do so, just like we've replaced most of the C code with which we started.
Not a fan of assembly for cryptographic code.
Are we likely to see stdlib changes in python, ruby etc?
The site mentions:
> Enforce a no-panic policy to eliminate the potential for undefined behavior when Rustls is used across the C language boundary.
Is this a thing? Does panic open up for undefined behavior when using a rust library via C?
I can't recall seeing it mentioned before/in other rust threads/projects?
I whish you the best of luck- libopenssl is terrifying :)
Nonetheless, there is a working group, proposing to make it defined under certain circumstances. https://blog.rust-lang.org/inside-rust/2020/02/27/ffi-unwind...
In the Rust case, IIRC, Rust annotates functions with the LLVM attribute no_unwind. I'd expect bad things, thanks to optimisation passes, if an unwind were to start. Doubly so if mixing GCC with MSVC ABI.
Yes.
- panic=abort (uncatchable)
- panic=unwind.
With the latter, you would need a catch_unwind at the FFI boundary.[1] https://doc.rust-lang.org/cargo/reference/profiles.html#pani...
AAre the panics being discussed for removal in rust tls "soft" errors that should not crash?
https://github.com/ctz/rustls/issues/447#issuecomment-820719...
It seems like there's been the most effort on category 4, the ones that can't be reached but can often be proven away with refactoring and typestate. Removing these ensures that the APIs don't expose Result types to callers unnecessarily.
Category 1-3 have different severity, but they are all planned to be folded into Result error values.
> Make it possible to configure server-side connections based on client input.
The other three bullet points I can immediately understand what they're achieving (the "No-panic policy" is trickiest but since I'm learning Rust I knew what that means, and it also links a ticket describing more)
But for this one I haven't a clue, maybe it should be obvious, and somebody else will explain.
For example, changing the certificate used based on the ALPN identity, the SNI server host, and the advertised client ciphersuites.
Though you can't use it to pick other properties, like making an ALPN protocol conditional on SNI.
Long term, is this OK? How do we ensure the assembly and C code in Ring is as safe as Rustls itself?
EDIT: I found issue #423², but the reporter is not making a compelling case, since the use case presented is for a prospective new application, yet to be written.
That being said and from what I understand Rustls is not a drop in replacement and it is not that easy for all Rust libraries which use TLS.
Which brings me to the point that I think a big leap forward for wider Rustls adoption in the Rust world itself would be to make it easily usable with all popular and widely used Rust libraries that depend on a TLS implementation.
What I would like to see is sort of a global (not per dependency)
features = ["rustls-tls"]
so that all libraries and their dependencies use Rustls automatically
and OpenSSL is completely out of the picture. I know that this is not on Rustls alone but
on the library writers too, but still it would be really cool to switch the TLS implementation
like that.Then we could even dream to make Rustls the default and use OpenSSL (or one of its relatives) only if need be.
"Right this minute" is a common phrase, but "does right this" isn't (at least in UK English).
I don't think you can get away with not treating HTTP as a first class citizen anymore. IMO, Go got this right.
The ecosystem is part of rust.
Rust will have support only when it is included in the standard library.
I think people are fairly apprehensive towards a bigger std because firstly, dependency management is almost trivial (C++ needs the kitchen sink in std because dependencies don't exist to the language, for example) so it's not that big of a deal, and secondly because a bigger std means a bigger commitment to APIs and maintenance by a relatively small group of maintainers for all eternity.
At the level of the stack where it makes sense to use rust, it doesn't make sense to put HTTP in std. At least not like Go.
Rust has included TCP and UDP in the standard library* because it was rightly recognized that these protocols, like stdio and filesystem io, have become fundamental to how most modern software interacts with its environment (system, IPC). HTTP today has also become fundamental in a similar manner. I argue that it is time that we treated it as such.
* as apposed to leaving them as an external dependency, like mio.
[1] https://doc.rust-lang.org/std/net/struct.TcpListener.html
Far better to let the experts do it.
I should also add that most applications won't even be using the TCP/UDP primitives directly. They'll likely be using a third party library. Rust's stdlib is intentionally minimal and lacking a lot of higher level features that are practically necessary for many applications.
It's the wrong abstraction too - this is handled better by dynamic linkage. Rust just doesn't have a great dynamic library story (a C API/ABI is not exactly what I'm talking about here).
What you'd really want is a
trait TLS {
// ...
}
in some base tls crate and leave the implementation up to other crates. And any upstream dependency would require something like: pub fn initialize <T: TLS> (tls:T) { ... }It's pretty bare bones though, because to expose things through that, all of the underlying implementations need to support it.
One problem with Rustls is that the error messages are not very informative. The situation described above just generated an "Invalid certificate" message. More use of anyhow::Context would be helpful. I don't disagree with Rustls disallowing decade-obsolete crypto. It's the "silently ignores" part that's a problem.
Because of how X.509 certificate validation works, in general it's not possible to tell you why an issuer couldn't be found, because there are many possible reasons.
Regardless https://github.com/briansmith/webpki/issues/206 tracks improving the situation.
In a sense it's going to have a big undigestible list of reasons the certificate wasn't found trustworthy, like if you asked grep to tell you why all the non-matching lines in a file don't match a regular expression. "The first letter on this line wasn't a match, and then the second letter wasn't a match, and then the third..."
However, as that ticket says, one relatively easy thing the code could do is notice if there was only one consistent reason and if so tell you what that was.
Also I agree with several commenters that webpki's current behaviour, in which it says "Unknown issuer" even when that's not the problem at all is undesirable and an even vaguer error might actually be better for these cases. See also, languages in which the parser too easily gets confused and reports "Syntax error" when your syntax was correct but something else is wrong, "Parse failed" is vaguer but at least doesn't gaslight me.
In some cases, e.g. the end-entity certificate is signed with an algorithm that isn't supported, or an RSA key that is too small, we could add special logic to diagnose that problem. However, all this special diagnostic logic would probably approach the size of the rest of the path building logic itself. It doesn't seem appropriate for something in the core of the TLB of the system. Perhaps at some point in the not-too-distant future we can provide some mode with more diagnostic logic, and find a way to clearly separate this diagnostic logic from the actual validation logic to ensure that the diagnostic logic doesn't influence the result.
There are a variety of interesting strategies to decide that an end entity certificate you were shown is trustworthy, and you're presumably aware that Mozilla (in Firefox) eventually chose to go with a strategy of
* Requiring all root CAs to disclose intermediate CAs as part of their root programme
* Bundling this set with Firefox
Whereupon Firefox gets to decide whether the end entity certificate it was shown was issued by one of these trustworthy intermediates and short cut to a "Yes" answer regardless of which certificates, if any, the server included in the supplied "chain".
In the general case this isn't very applicable, but I note that webpki is named "webpki" and not "General purpose certificate validator" so actually it could go the same route (with the caveat that this needs frequent updates to avoid surprises with very new CAs)
Mostly though if you're quite determined not to introduce more complicated logic into webpki (which is understandable) I specifically don't like the gaslighting of saying "unknown issuer" as I said, when the reality is that you don't know why you don't trust the certificate, so say that.
If std::fs::File::open() gives me Result with an io:Error that claims "File not found" but the underlying OS file open actually failed due to a permission error, you can see why that's a problem right? Even if this hypothetical OS doesn't expose any specific errors, "File not found" is misleading.
I love that they did that; it was actually my idea (https://bugzilla.mozilla.org/show_bug.cgi?id=657228). I believe the list is pretty large and changes frequently and so they download it dynamically.
> short cut to a "Yes"
Do they really do that? That's awesome if so. Then they don't even need to ship the roots.
> I specifically don't like [...] saying "unknown issuer"
https://github.com/briansmith/webpki/issues/221
> If std::fs::File::open() gives me Result with an io:Error that claims "File not found" but the underlying OS file open actually failed due to a permission error, you can see why that's a problem right? Even if this hypothetical OS doesn't expose any specific errors, "File not found" is misleading.
A more accurate analogy: You ask to open "example.txt" without supplying the path, and there is no "example.txt" in the current working directory. You will get "file not found."
Regardless, I agree we could have a better name than UnknownIssuer for this error.
I didn't know that. Congratulations, I see this survived an early WONTFIX that tried to downplay the privacy implications of AIA chasing, which is even more impressive in today's "Who cares about privacy" world.
> Do they really do that? That's awesome if so. Then they don't even need to ship the roots.
I don't in fact know if they do that. They will always conclude that a typical Let's Encrypt certificate from R3 is trustworthy via ISRG Root X1, even when the "chain" provided by the server leads to the (still trusted) DST Root CA X3 but they could actually be choosing that path on the fly rather than just short cutting.
They do need to ship the roots still because
1. The UX does actually show roots, I still have Certainly Something installed here, but the built-in viewer also shows them.
2. Users can manually distrust a root. It would be weird if either: we expected users to go in and manually distrust dozens of weirdly named intermediates that chain back to that root, or, disabling the root was possible but just silently didn't work in most cases.
3. Some trustworthy intermediate CAs can exist that aren't captured. Imagine if Let's Encrypt spins up R5 tomorrow because of some disaster that makes both R3 and R4 unusable, they can sign it with the ISRG root, and it'll work for lots of people - it's not great, but it's workable, however even if they tell Mozilla immediately, there's just no way everybody's Firefox learns about this instantly. So it's good that if your server shows leaf -> R5 -> ISRG X1 (or indeed leaf -> R5 -> DST Root CA X3) the Firefox browser can still conclude that's trustworthy even though it didn't know about R5.
I look forward to seeing issue 221 resolved, thanks.
Ideally you want a formally verified SPARK or CompCert C implementation of TLS, but those are not open sores compilers so they won't fly.
Having a SPARK-verified version would of course be great anyway.
That's a hefty judgment. If we look at major attacks on TLS endpoints I'd summarize them as, in order,:
1. Memory safety
2. Weak / old configurations
3. Invalid state machine transitions
Rustls addresses all 3 of those, or at least it attempts to. (1) is obvious - it's rust. (2) rustls only supports the subset of TLS versions that are considered safe. (3) rustls avoids issues like gotofail and smacktls by encoding state machines as types, turning invalid state transitions into type failures.
Plus, the actual crypto primitives are extremely well tested and built off of other existing libraries.
So yeah, maybe some aspect of the crypto is incorrect, but, while interesting from an academic perspective, the real world ranks those other 3 things as way more important.
Half of these are caused by C problems not present in Rust like: null pointer errors, buffer overflows, integer overflows.
The rest are logic errors (e.g., parsing errors) or crypto vulnerabilities (e.g., side channel attacks). There's nothing about Rust that magically prevents these errors. These vulnerabilities are discovered through testing and more importantly real world usage. Rustls has not been used nearly as much as or as long as openssl, so critical bugs could be in Rustls. You could encode some logic in the type system, but the types are still code that have to be tested. It's not a formal proof. It could still be wrong.
It depends on your threat model and your risk tolerance whether you want to depend on relatively unvetted code.
> These vulnerabilities are discovered through testing and more importantly real world usage.
FWIW I disagree that real world usage is necessarily more important. It is quite important, but at the same time the real world usage isn't going to hit all sorts of weird quirky paths. By contrast, Rustls has 97% line coverage, OpenSSL has ~65%. Of course, coverage isn't everything, but that's notable.
> It's not a formal proof. It could still be wrong.
So long as the type system is sound, I'm not sure this is true. If you embed a state machine into your type system your type system guarantees that the state machine executes as defined. Granted this is not checking against an external model.
I will say that rustls still will benefit from wider adoption and more evaluation but imo it addresses the largest concerns today.
One verified TLS implementation is miTLS, written in F*.