GnuTLS considered harmful (2008)
openldap.org
openldap.org
[1] The OpenSSL license is incompatible with the GPL, making it technically illegal to distribute binaries of GPL programs linked with OpenSSL (so Debian refuses to do so), unless the GPL program has an OpenSSL license exception.
Networkmanager → libsoup → glib-networking → GnuTLS
Eclipse → WebkitGTK → libsoup ...
VLC → FFmpeg → GnuTLS
kdelibs (core KDE libs) → upower → libimobiledevice → GnuTLS
I have a feeling there are greater dependencies within GNOME distros and, as you said, Debian. Networkmanager is an especially annoying one because it uses NSS directly and gnuTLS indirectly. $ apt-cache rdepends libgnutls27 libgnutls26 libcurl3-gnutls
$ apt-cache --installed rdepends libgnutls27 libgnutls26 libcurl3-gnutls
The latter invocation only displays packages which are installed locally, the former prints all packages apt is aware of.[1] https://developer.mozilla.org/en-US/docs/NSS [2] http://cve.mitre.org/cgi-bin/cvekey.cgi?keyword=nss
https://wiki.mozilla.org/NSS_Library_Init
RedHat has misguidedly chosen to base all of its security infrastructure on NSS, even though NSS was never designed for such a use. It is completely inappropriate for servers, multi-user workstations, etc.
It reminds me a lot of environmentalists going crazy to ban nuclear power in the 70s before we had as clear a grasp on the impact of dumping carbon dioxide into the air.
Why do you jump to blame the GPL and rms, when one could just as easily fault the OpenSSL authors for using the 4-clause BSD instead of the far more common 3-clause?
> It's technically illegal to use a better solution because of something as relatively unimportant as a license.
No, it is technically illegal to distribute compiled binaries that use OpenSSL, because the OpenSSL authors wanted to retain the advertising privileges. But it is not illegal to use the software as long as it is distributed in source and compiled by the end user.
I would not call licensing unimportant. As long as software is copyrightable, licensing terms are highly important.
The problem is that the GPL willingly refuses to permit advertizing clauses. Is there a congent argument about why an advertizing clause is a limitation of freedom? The GPL is more often than other free licenses putting restirctions on usage of diversely licensed software. It is an impediment. And, as we see, it has real-world consequences. There is more risk for freedom using bad software security than wielding to innocuous clauses.
The advertising clause is not a limitation on freedom. The 4-clause BSD license is a free software license; it just happens not to be compatible with the GPL (not all free software licenses are).
The reasons for this are very practical: not only does it place additional restrictions on the software (which is not permitted by the GPL), but if multiple 4-clause BSD projects are used, each project requires its own separate advertising statement (the 4-clause license does not permit combining these into a single sentence): https://www.gnu.org/philosophy/bsd.html
> The reason the GPL is annoying is that free license with an advertizing clause have existed for a very long time and are actually widely used.
Most modern projects using permissive licenses use 3-clause BSD, MIT/X11, or Apache, all of which are compatible with the GPL. In this day and age, choosing a 4-clause BSD license is a fairly conscious decision to make the project incompatible with the GPL.
Complaining about this seems a bit strange, since GPL is deliberately incompatible with everything else when it comes to sharing. OpenSSL's license, although kooky, is freer than the GPL in terms of who can use the stuff covered by it.
For the record, I'm not complaining. I'm just saying that it's unfair to blame the incompatibility solely on the GPL (as OP seemed to be), when the developer is the one who chooses the license for their software. (And I presume the OpenSSL authors are experienced enough to be familiar with the compatibility differences between the 3-clause vs. 4-clause BSD license).
> OpenSSL's license, although kooky, is freer than the GPL in terms of who can use the stuff covered by it.
No, both are equally free. Both of them respect the four freedoms, so they are both free licenses.
(The 4-clause BSD is arguably more permissive, but on the other hand, the GPL permits one to advertise the software without any restrictions, so it really depends on which of those two one values more. Generally the copyleft clause is what people care about more than advertising, but it's important to note both).
(Also, remember that the developer could always dual-license - ie, "GPL or 4-clause BSD - if you want to use my software in proprietary code, then you have to advertise me").
Thanks for your even-keeled comments here; helpful and refreshing.
You have mistaken what "advertizing clause" means. The GPL requires that the about box list the copyright holders, so that can't be the types of advertising at issue.
No, the complaint is about:
* 3. All advertising materials mentioning features or use of this
* software must display the following acknowledgment:
* "This product includes software developed by the OpenSSL Project
* for use in the OpenSSL Toolkit. (http://www.openssl.org/)"
If you have software which uses OpenSSL, and to promote it you send out a tweet, then the license requires you to include the above two lines in the tweet.In practice, a project might have 20 such advertising requirements. It gets boring.
An advertising clause can be used as a weapon. Suppose I distribute "free" software to you, but require you to include a 100 page manifesto every time you make an advertisement. Is that really "free"?
If my project is "SecureTalk" with the tag line "the NSA will never know", and it's secure because of OpenSSL, then will I have to mention that text every time I use the word "SecureTalk" in a tweet/ advertisement?
What about "HushTalk"? "MumsTheWord"? "SafeBanking"?
If I add optional rot-13 encryption, so there are now two cryptosystems, then can I pretend that SecureTalk doesn't "really" require OpenSSL, so I don't need the advertising?
No. If you have software that uses OpenSSL, and to promote it you send out a tweet that says "Use our product instead of our competitors, We use SSL to make things secure", then you must include the above two lines
For the clause to apply 1. It has to be an advertisement 2. It has to advertise the features that use openssl
I pointed out that the edge cases are fuzzier than I would like. If my product is called "SecureTalk", and uses OpenSSL for secure connections, then it sounds like almost any mention of the name which might be advertising needs to include that line.
As in, "Secure Systems, the developers of the NSA-proof SecureTalk, are hiring."
Isn't that "mentioning features" of OpenSSL? If so, it needs that line. If not, why not? What does it mean to mention a feature? Can I get away with
"Secure Systems, the developers of SecureTalk, are hiring."
After all, the only reason it's secure is because it uses OpenSSL.
Copyright is sticky. The hypothetical "SecureTalk" program might only use 500 lines of OpenSSL, where that 500 lines was security audited by crypto experts, static code checkers, and formal program analysis, and run in a chroot'ed jail.
A clueful re-use of OpenSSL for secure connections still needs that advertising clause, even if the software really is more secure than anything else out there. In that case, the required advertisement is a false clue to experts, no?
I also doubt that anything that uses OpenSSL as the primary crypto could possibly be "more secure than anything else out there". This isn't so much a slam of OpenSSL, which may overall be doing a better job of implementing TLS than anything else available right now (at least open source) but of TLS in general which is complex and not designed with current best practices. Using TLS is often an easy way to make things a lot more secure than they are without much effort and as such is often a good choice, but it is unlikely to result in the most secure thing possible. OTR is a well known alternative in chat that has a number of advantages (and some disadvantages too). Various others are under construction. Importantly, there are significant tradeoffs involved and it is often not a simple matter of X is more secure than Y.
What constitutes "mentioning features of this software"? If I use another package for SSL and advertise that my software has SSL support, but have OpenSSL in my code for other reasons (let's say, the SHA-1 digest code), then do I need to mention OpenSSL? After all, SSL is a supposed feature of OpenSSL.
No, it's not as bad as I make it out to be, but that's in large part because we are generally lazy when it comes to the particulars of licenses. Just look at the number of GPLv2 software distributions which don't follow the letter of the license. (Section 3 assumes physical distribution, not network. GPLv3 clarified this problem.)
It's also because license holders are lazy. Enforcing the GPL takes a lot of time and effort. Many violations occur because few actively enforce the license.
If your expectations are based on what people do in a lazy world, then you are perhaps a realist (or a cynic), but it still violates the license.
The "pages and pages of advertisement clauses" affects only to those who actually follow the license. These might be nitpickers like me, or organizations with lots of money and who are easy pickings and worried about liability.
These also happen to be the people who are likely to give acknowledgements, especially when the license so requires it (as the GPL does).
The GPL doesn't specifically set out to prevent advertising clauses. It is a side-effect of being incompatible with "other restrictions" - for example, a requirement that you license some third party software or patent in order to redistribute GPL-covered code. Instead of trying to specifically enumerate and disallow all such restrictions that someone might come up with, which is a fool's errand, the GPL disallows any other restrictions.
As a minor quibble, section 7 of GPLv3 allows a few other restrictions. That is, there's a general blacklist, as you say, with a specific whitelist of what additional restrictions are allowed.
For example, "b) Requiring preservation of specified reasonable legal notices or author attributions in that material or in the Appropriate Legal Notices displayed by works containing it;"
The thing is, relicensing isn't likely to happen any time soon, regardless of what RMS says.
How is that irony? I Don't think rms or anyone who pushes for Copyleft does so because they believe it always results in a superior solution.
https://en.wikipedia.org/wiki/BSD_licenses#4-clause_license_...
I also don't understand the environmental anecdote. That seems less about dogma and more about imperfect scientific knowledge. Were the environmentalists opposing nuclear energy on principle or because at the time the evidence made nuclear power look unsafe and detrimental to the health of the environment?
Rev. Dr. King had this to say about pragmatism: http://www.africa.upenn.edu/Articles_Gen/Letter_Birmingham.h...
There are a BSD vs GPL discussion about once every week on HN. Out of those several hundred threads and thousands comments, has a single users been convinced about the preference of either license type? Has a single person said "o, sorry, I will now change my opinion and use your license of choice because your arguments is so good".
Hate or love RMS, but can you keep it in your pants and do it elsewhere?
However, some other distros such as Arch do not have wget depending on it, so you do have a point about Debian.
It's rather silly that the news of a critical bug in GnuTLS that was caused by a goto somehow makes non-news and factually wrong information from 5 years ago popular.
Not that this affects your point, but the critical bug was not caused by a goto. Rather, it was caused by a mismatch in return value semantics, where a variable was used to store a value where 0 meant success, and then later used to return a value where non-zero meant success.
It's also discussed later in the thread: http://www.openldap.org/lists/openldap-devel/200802/msg00100...
> You note that there's really a small number of instances of strcat() in the code. That's true, but that's because you've provided your own _gnutls_str_cat() function instead, which is also heavily used. Assuming that strlen() isn't going to SEGV on you (which depends on dumb luck) this becomes just a question of efficiency.
I also think that in the rebuttal the example is extremely poorly chosen since the code is equivalent to the much simpler
char str[256] = "PKIX1.CRLDistributionPoints.?1.distributionPoint.fullName";
Assuming you even need str to be 256 char long, otherwise you would use 'char str[] = "...";' or 'const char *str = ' if you don't modify the string.Maybe the use of the concatenation is legitimate in the real code but I cannot judge that since it appears to have changed since the article was written:
https://gitorious.org/gnutls/gnutls/source/d9ce82a4ce690857f...
No strcat in there. Maybe it wasn't such a good idea after all? :)
EDIT:
Actually, I dug into the git repo to find the old code and looked to revert to a commit around the date the blog post was written. Obviously I don't intend to get any work done this afternoon. I found this commit on the same day (2011/05/10):
https://gitorious.org/gnutls/gnutls/commit/3df051196838f4f43...
"eliminated last instances of strcpy() and strcat() to keep pendantics happy."
So the reason he says the problem is not here anymore is because he fixed it just before writing this blog post, more than 3 years after the openldap rant. I'm sure nmav was well-intentioned but it does weaken his rebuttal somewhat.
Only that this rebuttal completely misses the point. They misunderstood the criticism being about buffer overflow vulnerability
>> So what is the issue? Howard claims that GnuTLS makes liberal use of strcpy(), strcat() and strlen(). Those functions are known to be responsible for several attacks via buffer overflows in current programs.
while it was in fact about the nature of the data to be processed, namely that it may not be NUL terminated strings but arbitrary binary data for which the whole bunch of `str…` functions and any other string processing that expects to operate on NUL terminated strings will miserably fail
> Looking across more of their APIs, I see that the code makes liberal use of strlen and strcat, when it needs to be using counted-length data blobs everywhere. In short, the code is fundamentally broken; most of its external and internal APIs are incapable of passing binary data without mangling it. The code is completely unsafe for handling binary data, and yet the nature of TLS processing is almost entirely dependent on secure handling of binary data.
"It turns out that their corresponding set_subject_alt_name() API only takes a char \ pointer as input, without a corresponding length. As such, this API will only work for string-form alternative names, and will typically break with IP addresses and other alternatives."
Yes, an API designed for strings will break if you pass it a struct in_addr or something, but it should be fine with a dotted-decimal string, right?
My understanding of RFC 3280 is pretty old, but the relevant ASN.1 type describing a subjectAltName seems to be :
SubjectAltName ::= GeneralNames
GeneralNames ::= SEQUENCE SIZE (1..MAX) OF GeneralName
GeneralName ::= CHOICE { otherName [0] AnotherName, rfc822Name [1] IA5String, dNSName [2] IA5String, x400Address [3] ORAddress, directoryName [4] Name, ediPartyName [5] EDIPartyName, uniformResourceIdentifier [6] IA5String, iPAddress [7] OCTET STRING, registeredID [8] OBJECT IDENTIFIER }
The IP address case is represented as an octet string, and the octet 0 is legitimate, making their API broken...
My point was that it's not reasonable to expect an interface that appears to be accepting a string to also accept random bytes; "10.0.0.8" isn't the same as 0x0a000008.
Except that it's not. At least on glibc, strlen() is declared "pure" to the compiler and (unless otherwise defeated by pointer aliasing) repeated calls will be optimized away.
That's not to say that this is the best way to write the code, but the concern seems poorly informed.
Maybe you're saying it's a "code smell" kind of thing and that being sloppy here indicates more subtle problems elsewhere? Which then hits the argument about whether this is really "sloppy" or just intentionally simple.
Shrug. My point was just that this needs better evidence. There is no demonstrated bug in the linked code, and the assertion that it is "gross" (OP) or "insecure" (you) seems poorly justified.
[1] at least according to the post. The fact that gnutls added a binary interface later seems to support that reading.
If this was C++, and the function took a std::string, would you say it was horribly broken because you serialized a 4 byte IP address into a 4 byte std::string buffer and the function didn't handle it correctly?
In practice calling any other non-pure function, which the compiler can't see in to, will defeat this optimisation, whether there's aliasing within your function or not.
if ((len_len + (int) strlen (str)) <= max_len)
with int's I immediately start worrying about integer overflow leading to buffer overflow. Nobody seems to have mentioned this though.No wait! Ruby, that should be perfect, can't cause buffer overflow there....
Ermm.. wait I got , well use python!
Hint: it starts with a C.
If you build a house in a swamp, you are in trouble.
If C is the wrong choice, what is the right choice?
Why, Python, of course![0]
[0] - http://pypy.org/
(I jest, your point is entirely valid)
> Bogus objects — properly tagged words with invalid addresses that pointed at uninitialized memory or into the middle of object of a different type — which would cause the GC to corrupt memory would be left in registers or on the stack. These sort of problems were everywhere in the microcode.
Go build your cabin in the woods now.
Something like a buffer overflow was very very rare. This was not a systemic problem like it is in C.
And there are Java compilers and runtimes written in Java, and Lisp compilers written in Lisp, etc.
There's a bit of research in this, but I don't know of anything that's ready to use.
For that, we'd need: (1) A set of properties to verify for an implementation (2) An implementation to verify
Looking at http://coq.inria.fr/related-tools, there may be some related tools that could do the job. Still, there's a research project there.
A few substantiating links from my search results for "coq generating code":
http://coq.inria.fr/V8.1/refman/Reference-Manual021.html
http://research.microsoft.com/en-us/um/people/akenn/coq/LOLA...
http://research.microsoft.com/en-us/um/people/nick/coqasm.pd...
The issue is developers are not using the tools or following the best practices because they think they know better than 30 years worth of experience or get caught up in bikeshedding about ideology, licenses and which line the curly braces go on.
http://www.amazon.com/The-CERT-Secure-Coding-Standard/dp/032...
Online:
https://www.securecoding.cert.org/confluence/display/seccode...
Others:
http://www.amazon.com/Secure-Coding-Edition-Software-Enginee...
http://www.amazon.com/Style-Guidelines-Programming-Professio...
http://www.amazon.com/C-Traps-Pitfalls-Andrew-Koenig/dp/0201...
Nevermind the copious undefined behavior, the fact that C programmers sometimes struggle to figure out what a valid C expression actually does, the fact that C programmers have to choose between code bloat and using "goto" for finalization, the fact that there are no standard error handling constructs, the fact that strings are null terminated, the lack of a standardized way to determine array lengths at runtime, etc., etc., etc. Even something as simple as this:
int f(int x, int y) { return x + y; }
Can lead to undefined behavior in C:
https://www.securecoding.cert.org/confluence/display/seccode...
Basically C should be at the bottom of the list of languages that programmers choose for cryptography or security software.
Discussion of the GnuTLS bug is summarized here https://www.debian-administration.org/users/dkg/weblog/42
And people still wonder that GnuTLS certificate verification bugs continue to surface?
The fact that there are still certificate validation bugs in GnuTLS today indicates that the GnuTLS developers still haven't learned the essentials of X.509 certificates. Even with a rapidly deployed fix for this most recent CVE, you'd be a fool to rely on GnuTLS for anything. The code and the developers have proven themselves not to be trustworthy. Multiple times.