Read Schneier and Ferguson describing in _Cryptographic Engineering_ the steps required to validate a bignum arithmetic library. That's one of the smaller problems in getting a TLS stack right, and it's complicated.
If you demand that your browser implement protocols the same way ActiveRecord implements SQL join generation, you should just stop using browsers; none of them are built the way you'd want them to be.
I have a hard time not being dismissive about TDD's role in improving systems software security. If you're going to propose a radical change in the way systems code is written (or, worse, hold systems developers to a nonexistent standard), be intellectually coherent about it: demand that security-critical code be implemented in a rigorous language, like idiomatic Haskell. Don't propose monumental changes that promise only that the resulting code will be asymptotically secure as Ruby on Rails.
What I read him as saying (and agree with) is that he's surprised that GnuTLS isn't tested against known-bad certificates in a simple integration test that doesn't require a Ph.D. to set up, as you imply.
The fallacy that you could have easily written a test case for this bug appears neatly encapsulates the weakness of "TDD" as a mitigation for programming flaws.
Your résumé is well-known around here and you don't have to set the terms of the discussion about every security-related issue on Hacker News, particularly when it's this heavy-handed and you get recognition upvotes. All I'm saying.
Him: Can't you test for this somehow? Even without TDD...
You: TDD is bad because math and wall of text.
Go read a book and realize why you can't, Web hacker.
Me: You jumped on TDD unnecessarily, there.
You: Now I'm going to redirect the argument to you!
You're browbeating anybody that comments here, rather unnecessarily, almost as a display of expertise. - Testing distributed cryptography is difficult.
- Read a book.
- You're a web developer and don't know the first thing about
testing TLS stacks, clearly.
- Testing does not improve systems security, as evidenced by
Ruby on Rails: Rails uses testing and it's insecure.
Seriously, re-read your comment. More than half of it is just unnecessary and dilutes what little point you've made into something unrecognizable in the snark.This is an important and valuable point. He wasn't discouraging new ideas, but rather pointing out how hard it is to make new ideas practical in the security arena.
It's hard to see how that test suite couldn't help improve the reliability (and therefore "reduce the insecurity") of said new implementation (say a library with only support for a subset of tls1.2 -- without any support for fallback).
Define "known-bad" in a general enough way that a specific test can be created to cover the entire range of "bad" certs. That's quite difficult, and, probably isn't realistically possible to go through all the "known-bad" if you want your tests to run quickly.
Realistically, all you can do is have regression tests to make sure that the found bugs aren't repeated in future releases.
Feature: *Certificate Validation*
In order to *keep NSA from reading my emai*
As a *TLS X.509 validation library*
I want *to never erroneously validate a certificate*
What are the "Scenarios"? Given: *???*
And: *???*
When: *???*
Then: *the certificate should be rejected*
Remember, if we're switching topics to the "goto fail" bug: that bug didn't affect every instance of certificate validation. You had to be in a particular set of ciphersuites.But I agree that you & I are better off not discussing things.
If you'd permit me a brief bit of my own concern trolling: I remember when I looked forward to reading your comments, several years ago. Now I see your nickname and say "bah, again?" What changed? Was it me or you?
Compared to that:
- Better unit testing is much easier to integrate into existing projects. Yes, it can only prevent a bug if the developer generally thought of the class of error, but at least it sort of forces them to spend some time thinking about possible failure cases, and can detect cases where their mental model was wrong. Also, it helps detect regressions: "goto fail" wasn't a strange edge case the developer didn't think of, it was a copy paste error which good unit tests could have caught.
- Functional testing can be independent of the implementation and written by someone unrelated. They can only do so much in general, but they might have caught both of these bugs.
Yes, audits are another option, but I'd say they should complement tests, not replace them.
ed: oh, and if you want to be really intellectually rigorous, you could try to formally verify your C code; model could have bugs but could also be implementation independent. But I hear that's rather difficult...
I don't all together disagree with your points, however I'd like to pint out that Haskell has a great C FFI:
http://www.haskell.org/haskellwiki/GHC/Using_the_FFI http://book.realworldhaskell.org/read/interfacing-with-c-the...
Stipulate that we're just talking about X.509 validation. (You can still have "goto fail" with working X.509, but whatever).
Assume we can permute every field of an ASN.1 X.509 certificate. That's easy.
Assume we're looking for bugs that only happen when specific fields take specific values. That's less easy; now we're in fuzzer territory.
Now assume we're looking for bugs that only happen when specific combinations of fields take specific combinations of values. Now you're in hit-tracer fuzzer coverage testing territory, at best. The current state of the art in fault injection can trigger these types of flaws (ie, when Google builds a farm to shake out bugs in libpng or whatever).
Does standard unit testing? Not so much!
Would any level of additional testing help? Absolutely.
But when we talk about building test tooling to the standard of trace-enabled coverage fuzzers, and compare it to the cost of adapting the runtime of a more rigorous language --- sure, Haskell is hard to integrate now, but must it be? --- I'm not so sure the cost/benefit lines up for testing our way to security.
For whatever it's worth to you: I totally do not think code audits are the best way to exterminate these bugs.
Also note that Haskell doesn't exactly help in avoiding timing attacks. In a sane cryptosystem, you might be able to implement AES, ECDSA and some other primitives in a low-level language and use Haskell for the rest; but as you know, TLS involves steps like "now check the padding, in constant time" (https://www.imperialviolet.org/2013/02/04/luckythirteen.html). You could certainly implement those parts in C, too, and then carefully ensure that no input to e.g. your X509 parser can consume hundreds of MB of memory, and so forth, but you're going to lose some elegance in the process. (Those problems would admittedly be smaller in OCaml, ADA or somesuch.)
I'd be more interested in something like Colin's spiped - competently-written C implementing a much simpler cryptosystem. If only because even a perfect implementation of TLS would still have lots of vulnerabilities. ;-)
(I think the case for writing applications in not-C is considerably stronger, if only because TLS stack maintainers tend to be better at secure coding than your average application programmer. Like you, I do like writing in C, though.)
Any known-bad cert at all would have been quite sufficient to catch this bug apparently. A simple ARE WE ACCEPTING BAD CERTIFICATES LOL sanity-check would have found it, which is the kind of unit test it should be possible to think of in advance rather than in response to a specific bug found earlier. A little can go a long way.
EDIT: Additionally, the difficulty of catching all bad certs is good reason to develop and continually update a torture-test of invalid certs (and valid ones) to test SSL clients against. The suite would be much too slow to check against once per recompile, but testing once before each point release should be useful enough...
If I implement a network protocol I will make sure to write automated test involving clients and servers. Setting this up on localhost or on a virtual network using TUN/TAP is not that hard. And this has made me find TONS of bug ahead of time.
I hear a lot of arguments as to why testing networked or otherwise distributed things isn't necessary. But just look at Aphyrs complete destruction of well-known distributed systems by using realistic testing.
And don't claim that the kernel network stack isn't tested. It's just tested by hand. Plus, there are projects like autotest. Without testing, I don't think there's any chance that the kernel devs could release any more versions.
Safe languages, formal verification, using proven methodologies and testing are all separate methods of improving code quality. But I don't believe any one precludes the other.
I'll end with a famous quote from Knuth: "Beware of bugs in the above code; I have only proved it correct, not tried it."
I agree that it's helpful, but the problem is coverage. Your test code, most likely, doesn't cover EVERY single possible condition that can happen with a simple TCP/IP connection. Especially once you get out of localhost land, where you're dealing not only with your code, but all the hardware and software between the two systems
The fundamental problem is that even with a rigorous test suite, you're probably going to run into things you didn't even think were possible once the code is out in the wild. For example, we just ran into a scenario where we were seeing corruption through a TCP connection. Knowing the wire, it was impossible for the packets to be appearing in the way that they were(there was some packet level corruption). After going through multiple wireshark logs, we found that the culprit was a hardware firewall in between the server and client. Thankfully, our code didn't crash, but it's also something that we never tested for, because(in theory, at least), it should never be possible for that specific corruption to be sent in the first place.
Most CPAN Perl modules for almost two decades have such tests for much less critical stuff. There's no excuse for not having such methods for security critical code.
Tptacek is muddying the waters by claiming it's hard. It's not. We have more different implementations of the same protocols so it's even easier to cross-verify.
Apparently just a self signed cert. It was accepted as the "CA signed." Since 2005.
1) http://arstechnica.com/security/2014/03/critical-crypto-bug-...
- Here's a library that claims to validate certificates.
- Here's a TLS httpd with a forged commonName.
- for ciphersuite in $supported; do
Completing this test is an exercise for the reader, yet we're waxing philosophical about the perils of Ruby on Rails in this thread for some reason. If I'm being told by a security expert that such a framework would not be helpful, that's concerning and makes me wonder how many bugs such a framework would uncover (given that this one remained untouched for a decade).Some of the stuff has assertion files, too, so you can check the way it fails or check some conditions the success, but that's not common.
Overall, this took like a day or two to setup and it catches a lot of errors already, especially because people can just drop errors into the bad input folder and be done with it until someone has time to handle it.
For certificate verification, you craft various correct and incorrect certificates to cover the various scenarios involved in certificate verification.
For communication related code, you talk to the tested code over TCP or internally as a transport layer implementation (assuming the TLS library supports user defined transport layers). Same as above, you identify the various code paths and attempt to test them all.
> for the same reason that the kernel TCP/IP stack isn't built with TDD
I believe the primary reason why testing kernels is hard is that they are intended to run on real hardware and often with strict timing expectations. This is not the case for TLS libraries, which normally run in user space.
But in this case this is not about that. It's about testing code already written. All you need to do is go through that code and verify that every branch does what is supposed to do. Which is pretty straightforward, if time consuming.
If we define a correct format there are likely an infinite number of incorrect formats, no? A test explicitly checking for this bug would prevent a fix from regressing, but it seems to write a test that exploits this bug before understanding the bug itself would require quite a bit of luck.
EDIT: I'm re-reading the initial advisory and trying to decide if this applies to any cert with a ROOT CA fail or just a specifically crafted one. If it's the former my initial comment is garbage.
It's true, that doesn't prove the implementation is perfect, which seems to be what you're holding up as the goal. But it would have caught this bug.
But finding bugs like the above is TDD's bread and butter. TDD dictates specifically that the test must be written first, that it must be isolated to a particular spot as much as possible, and the dumbest piece of code possible must be written in order to allow the author to move on to writing the next test. Someone TDDing this method would stub out any required external state in order to focus only on the piece of code in front of them.
The final system may not be correct––you must still perform high-level testing as you always would, and you must understand the rules of the system you're building. But you're apt to avoid the simple stuff.
I know I could come up with a very basic set of test cases, and my only experience with TLS is reading parts of some book I found in my company's library. Heck, OpenSSL has a ton of tests it runs on every build, I know this much.
And even in Haskell you'd need tests. At no time will you ever be able to write code without testing it.
A TCP stack can be written in a way that makes it simple to test, and getting full branch coverage isn't that onerous a task. Does that catch all bugs? Of course not. But it'll still catch many simple logic errors like this one. Such logic errors do creep into code quite often, can be hard to trigger or detect during manual testing, and will be hellish to debug if they manage to get all the way through to a live system.
(This isn't idle speculation. I'm the lead engineer for a high performance userspace TCP stack that's handled petabytes of data, probably for millions of endpoints. We have deterministic unit tests for a large proportion of TCP behavior. Getting into a position where we could do proper testing took a while since the initial version was not written with testability in mind, but that work was definitely worth it.
But I'll take your word on TLS.)