Curl is C
daniel.haxx.se
daniel.haxx.se
>C is not the primary reason for our past vulnerabilities
>There. The simple fact is that most of our past vulnerabilities happened because of logical mistakes in the code. Logical mistakes that aren’t really language bound and they would not be fixed simply by changing language.
So I looked at https://curl.haxx.se/docs/security.html
#61 -> uninitialized random : libcurl's (new) internal function that returns a good 32bit random value was implemented poorly and overwrote the pointer instead of writing the value into the buffer the pointer pointed to.
#60 -> printf floating point buffer overflow
#57 -> cookie injection for other servers : The issue pertains to the function that loads cookies into memory, which reads the specified file into a fixed-size buffer in a line-by-line manner using the fgets() function. If an invocation of fgets() cannot read the whole line into the destination buffer due to it being too small, it truncates the output
This one is arguably not really a failure of C itself, but I'd argue that Rust encourages a more robust error handling through its Options and Results when C tends to abuse "-1" and NULL return types that need careful checking and can't usually be enforced by the compiler.
#55 -> OOB write via unchecked multiplication
Rust has checked multiplication enabled by default in debug builds, and regardless of that the OOB wouldn't be possible.
#54 -> Double free in curl_maprintf
#53 -> Double free in krb5 code
#52 -> glob parser write/read out of bound
And I'll stop here, so far 7 out of 11 vulnerabilities would probably have been avoided with a safer language. Looks like the vast majority of these issues wouldn't have been possible in safe Rust.
That bug is only triggered by an unrealistic corner case (username longer than 512MB). Run-time check in debug build won't help unless you realize the possibility before hand and put a unit test for that. I am more interested in the second part of your comment: "the OOB wouldn't be possible". How does rust protect against such integer overflow caused by multiplication? Thanks in advance.
What makes you call that an "unrealistic" case? You're probably imagining some sort of world where a randomly distributed set of usernames are sent to this function drawn from the distribution of "real usernames that people use". Since 0 of them are longer than 512MB, you intuitively assume a 0 probability of exploit.
But in the security world, that's the wrong distribution. You have to assume a hostile, intelligent adversary. It isn't hard at all to construct a case of some web service allowing you to specify remote resources accessible with a username & password of your choosing (not corresponding to the username you're using to log in), an attacker specifying one of their own resources, reading the incoming headers to notice that you're using a vulnerable version of libcurl, and stuffing 512MB+exploit into the username field of your web app. If you don't add any other size restrictions between the attacker and the libcurl invocation, they may well pass it right in. (And your same intuition will lead you to not put any size restrictions there; "why would anybody have a multi-megabyte username?" You won't even have thought the question explicitly, you just won't put a length restriction on the field because it'll never even cross your mind.) By penetration standards, that doesn't even rise to the level of a challenge.
The use of "unrealistic" was meant in the sense that you wouldn't think of it, not that you shouldn't guard against it once known.
I think there's two components here: yes, this might lead to a _logic bug_, but it should never lead to a _memory safety_ bug. That is, to get that CVE, you need both, and Rust will still protect you from one.
Edit: I believe I confused rust with another language.
> How does rust protect against such integer overflow caused by multiplication?
To me, the question was whether or not Rust would have been able to natively protect against this, considering it was a runtime issue?
Thus, you could have caused a denial of service by crashing a rust-based Curl, but the crash would have been modeled and just uncaught.
Anyway, the buffer overflow itself is:
>If this happens, an undersized output buffer will be allocated, but the full result will be written, thus causing the memory behind the output buffer to be overwritten.
In Rust's case the output buffer would be returned into a dynamically sized type, probably a Vec<>. Attempts to put more data in the Vec than it can hold would either cause a runtime error or cause the Vec<> to reallocate (which could cause a performance issue, or maybe even a DoS if you could cause the system to allocate GBs of memory, but it wouldn't allow access to invalid memory).
So maybe something like:
/* XXX potential multiplication overflow */
let mut buf = Vec::with_capacity(insize * 4 / 3 + 4);
while let Some(input) = get_input() {
// Reallocs when capacity is exceeded
buf.extend_from_slice(input);
}
So even if the multiplication overflow is not caught it's by design impossible to have a buffer overrun in Rust.https://medium.com/@blakeross/mr-fart-s-favorite-colors-3177...
All bugs are logical errors at various scale.
However for others it's not immediately obvious how a safer language would've helped, for instance #59: Win CE schannel cert wildcard matches too much
This is clearly a logical error due to a badly implemented standard. There's no silvet bullet here.
* A preprocessor #if statement is altered with defined(USE_SCHANNEL)&&defined(_WIN32_WCE) which is OR-ed with the other conditions, so that code that was previously not compiled on WinCE is now potentially defined.
* A local buffer in the code is increased from 128 to 256 characters. A comment in the patch refers to a "buffer overread". So there is a C issue in here!
* In a call that appears to be a Microsoft API function CertGetNameString, a flag argument that was zero is now specified as CERT_NAME_DISABLE_IE4_UTF8_FLAG. Unless I don't understand something in the patch comment, it doesn't appear to be remarking on this at all.
* Code that was taking on the responsibility for doing some matching logic is replaced by something that appears to be using proper API's within curl (Curl_cert_hostcheck).
Why didn't the programmer know about the existing function? One possibility is that it didn't exist yet at the time that code was written. Other such ad hoc matching code may have been refactored to use the matching function; this wasn't found. Or maybe the function did exist, but wasn't well documented. A review process isn't in place that would allow someone to raise a red flag "this should be calling Curl_cert_hostcheck and not itself using string matching at all, let alone be checking for a * wildcard and incrementing over it."
* makes code smaller with less distracting repeated boiler-plate, so things like that stand out more.
* programmers waste less time on fighting language-related ergonomic pains, and so more of their attention is available to spot these errors.
If your language is such that you shout "hooray" and pat yourself on the shoulder when it compiles and the code passes the address sanitizer and Valgrind and whatnot, then actual functional problems will slip under the rug.
Permit me, also, to indulge in some argumentum ad lingua obscura: could we spot in a Brain##### program that some certificate wildcard matches too much? :)
One thing that C programmers tend to love about C is that it's simple. C doesn't have that many language constructs and that makes C programs easier to reason about, read and audit, which also makes it easier to find these non-obvious errors. A C programmer approaching an unfamiliar codebase can feel assured that it is made out of the same concise set of constructs as any other C program. On the other hand, modern languages like Rust, C# and Go are a lot more complicated. They have most of the language constructs of C, plus some extra, so not only do you have to understand the concepts in C, you also have to understand things like ownership, variance or pattern matching. Every new feature increases cognitive burden on the programmer. Adding something like exception handling, for example, means that control flow suddenly becomes more complex. Now, instead of only leaving at return statements, a function has a potential exit point anywhere it calls another function. "Smart" features that encourage terse code and reduce boilerplate can also result in code that is totally obtuse to anyone other than the person who wrote it.
I'm not saying that C is at a sweet-spot for language complexity (Brainfuck is probably too simple, but OTOH it only has eight symbols!) I'm just saying that it's important to understand why a lot of developers really like C, and C is often praised for its simplicity. Any language that intends to replace C needs to understand why C programmers use C.
While I'm sure experienced C developers will be used to that, one cannot simply read The C Programming Language, set a language reference on their desk, and safely learn how C will behave through experience because of all the unspecified or counter-intuitive things which often don't even trigger compiler warnings with -Weverything.
While Rust may have more concepts to grasp and more grammar to remember, I find it far less of a mental burden to code in for the same reason that OOP proponents trumpet encapsulation... C forces me to constantly double-check that I haven't forgotten some detail of the language or GCC's implementation which is sensible if you understand the low-level implementation, but completely counter to my intuition. Rust allows me to audit the heck out of modules containing unsafe code to make sure the unsafety can't escape, then set it aside to think about another piece of the logic.
Rust also places a stronger emphasis on making it possible to reason locally, bolstered by things like hygienic macros.
Finally, simplicity does not automatically make something intuitive. (That's something I hear quite commonly among HCI people bemoaning how Apple has brilliantly used simplicity to convince iOS users that, if they get stumped by the UI, it's their own fault, not a failing on Apple's part.)
Plus a shit load of user defined macros to make it look like a modern language.
Every major c repo will have their own macros for foreach, cleanup on exit, logerrorandreturn, etc. Another example is extensive use of attributes like _cleanup_ from gcc to simulate raii, not plain c89 at all.
Is it not where you're supposed to point to re-implementations or equivalents written in Rust?
He addressed all of those points in the second short paragraph. None of those are C vulnerabilities, they were mistakes made on the part of the developers, not the language. Avoidance of problems in a safer language doesn't mean when things happen, it's the language's fault.
The point of type system features e.g. Option types instead of nulls, or linear/affine types avoiding use-after-free is to make programmer mistakes turn into compiler errors. Nothing more. There is no point in talking about whether or not something is the "languages fault" or not. We know C is like juggling knives. It's a tool. It has drawbacks and benefits. Being widespread and fast are the benefits. Not turning many forms of programmer errors into compiler errors is the drawback. That means the programmer can't make mistakes because they will be shipped. But programmers invariably make mistakes.
The grandparent argued that for the sample of issues he looked at, a lot would in fact be avoided by the type system of e.g. Rust - contradicting the argument in the blog post (could be because of the small sample though).
I think it's perhaps less important to focus on the number of issues of each kind, and instead look at the severity of them. If the kinds of issues avoided by better type systems are typically trivial issues, but the kind of issues coming from logic errors are severe security issues - then perhaps the case for stronger type systems isn't so strong after all. But I doubt that's the case.
It is not even that laborious, it just requires some discipline.
They are absolutely the fault of the language, given other languages would have made these bugs impossible.
Errare humanum est is hardly a new concept, blaming humans for not being computers is inane, and in fact qualifies for in errare perseverare diabolicum as out of misplaced pride you persevere in the core original error of using C.
I doubt the library would have reached its current adoption levels if it had been written in any other language (and I presume an other C library would've taken its place).
The former sense would be useful... if you want to sue someone or something I guess? But the latter is more useful in real life, so that's what we're discussing
#55 -> OOB write via unchecked multiplication
#54 -> Double free in curl_maprintf
#53 -> Double free in krb5 code
#52 -> glob parser write/read out of bound
Yes ...
> and contradict the statement regarding zero policy static analysis errors.
Not necessarily - I've found that static analysis tools (Coverity etc.) have limitations, and I'd expect them to find less than half of these kind of bugs - serious fuzzing tends to find more.
The static analysis tool has to work with the type system of the language, and C's type system isn't particularly helpful, so the tool has to find a balance between flagging almost every pointer dereference as "potentially a problem", most of them being false positives, and flagging only the small percentage of actual problems that can be unambiguously proven to be problems with the context available inside a single function definition.
(Just stating generalities, haven't looked at the fixes for these bugs.)
1. object should not exist and the second free is incorrect 2. object should exist and the first free is incorrect 3. object existance is uncertain and the second free must somehow check that
Although I can double-free only in unsafe languages, the wrong logic behind it can be the same in safe languages. It just have different consequences.
That is the whole concept of safety.
1. object should not exist and the second free is incorrect
In C# I can have to variables pointing to the same object, I null only one of them. The second should be nulled too, but it's not. That's a logic error. So in C# I end up with some object that should not be there, but it is. Which is better - doube-free or undestroyed object - depends on use case.
It can crash right away, in a few seconds, minutes, hours later, or never and just keep generating corrupt data.
Having a reference that the GC doesn't collect doesn't lead to memory corruption, just more being used than it should be.
Of course, you don't get it all for free; you have to wrestle with mutability and lifetimes, which can get hairy.
Actually that's exactly what it means.
That's the only way people mean "It's because of C". That a safer language would have prevented those classes of errors"
That's regardless whether a more careful programmer would also have prevented them in C too.
Would you ever declare a bug to be the language's fault, other than compiler bugs?
Sure, rust is "safer" than C, but is it a better language? that's arguable.
Obviously there are more aspects to a language than safety, so I'll give you that, but yes Rust is a better language than C in this aspect.
Don't forget that programming languages are meant for humans, not for computers. One of the primary goals of a programming language is to prevent humans from making dumb mistakes.
That clearly wasn't the case when C was designed.
That said, colloquial history seems to say that there were plenty of languages which did a better job at that than C.
My definition of fault is "what should be changed to prevent such error". And you are not going to change developers.
- Tony Hoar
I think this is a particularly solid point by Tony Hoare (made during his talk on 'the billion dollar mistake', null).
We find it very easy to blame developers for mistakes that really shouldn't have been possible to make ta all.
Now, I do in general prefer languages "without" null as they generally have more warning about when you can actually encounter null. Though my personal experience says that in sensibly written programs dealing with null isn't that big a deal. Is it a billion dollar mistake? I would have little trouble believing that. However, that would likely make dynamic typing an order of magnitude or two larger in the mistake dept.
I would also say there really is no comparison between null, which escapes typechecking, and unwrap, which does not. But that's not my point, nor is it Hoare's. It's that we blame people for problems that are better solved by languages.
On the one hand, Curl is a great piece of software with a better security record than most, the engineering choices it's made thus far have served it just fine, and its developers quite reasonably view rewriting it as risky and unnecessary.
On the other hand, the state of internet security is really terrible, and the only way it'll ever get fixed is if we somehow get to the point where writing networking code in a non-memory-safe language is considered professional malpractice. Because it should be; reliably not introducing memory corruption bugs without a compiler checking your work is a higher standard than programmers can realistically be held to, and in networking code such bugs often have immediate and dramatic security consequences. We need to somehow create a culture where serious programmers don't try to do this, the same way serious programmers don't write in BASIC or use tarball backups as version control. That so much existing high-profile networking software is written in C makes this a lot harder, because everyone thinks "well all those projects do it so it must be okay".
But in mid-way applications, dangerous languages are dangerous, and should be avoided.
I'd think the most "correct" way to handle this is to create small bits of network code with proof of correctness that handles down parsed and tagged (typed) data into code in high level languages. In that case, Curl type of code would still be in C.
In order for this argument to hold, we would have to live in a world where no one has created a safe language with no runtime.
We don't live in such a world, fortunately.
I don't think we are talking about the same thing at all.
Many people say "no runtime" to mean "small runtime like C or C++ have" rather than literally no runtime. This is due to the confusion around interpreters sometimes being called "runtimes."
The Rust runtime is larger than C's (I imagine it's not much larger than C++, but this is already way too large). AFAIK, nobody ever proved it works as designed. The fact that it's design is not very stable is also a problem.
Now, with some more thought about it, C itself is not a great target for correctness proofs. Too much is left unspecified, and there is little agreement over some of the lower level stuff. A simplified version of Rust might be a better target than C, if somebody ever creates it, but a lot of the safety would have to be left out of it.
Rust also lets you write a custom runtime by specifying a start function if you want.
> The fact that it's design is not very stable is also a problem.
Rust is stable.
I believe that it is smaller than C++'s but larger than C's, but I'm not a mega expert here. I used to know where those files were located in-tree, but don't anymore.
I think there's also several distinctions you can draw here; for example, strictly speaking, almost no runtime is required, but if you're using say, the standard library, you still might end up with a bunch of code. The actual requirements are very small though.
> AFAIK, nobody ever proved it works as designed. The fact that it's design is not very stable is also a problem.
Sorry, I don't understand what you mean here. The runtime?
> A simplified version of Rust might be a better target than C, if somebody ever creates it, but a lot of the safety would have to be left out of it.
I _think_ with the above comments, you mean the Rust langauge overall, not the runtime, right? So, Rust _is_ very stable; post 1.0 we have made extremely small breaking changes, corresponding to the way new C++ standards sometimes introduce minor technical backwards incompatibilities. As for correctness proofs, there are multiple academic institutions working on it; and they don't have to "leave the safety out"; that's what they're trying to prove!
As for Ada, you can selectively disable which parts of the runtime actually land on the generated executable via compiler pragmas.
Deallocation can be done via pools or controlled types.
What about:
MODULE Boom
VAR Foo : POINTER TO INTEGER;
BEGIN
Foo^ = 123;
END Boom.But if you still want to get your point through and have a boom.
MODULE Boom;
IMPORT SYSTEM;
VAR Foo : POINTER TO INTEGER;
BEGIN
Foo := SYSTEM.CAST(POINTER TO INTEGER, 43414);
Foo^ := 123;
END Boom.
Notice the use of IMPORT SYSTEM and SYSTEM.CAST, explicit, easy to search for, and to forbid via compiler switch (no unsafe code).I never did much with Ada or Modula, so I won't claim to know anything more than their syntax, etc.
Now Rust as a target for compiling proofs into a programming language is pretty nice. Safe code can have no aliasing due to borrow checker (which should also be formally verified). Threading issues are vastly simplified, though fairness still has to be verified, as well as a few other properties.
So using Linux, Windows, BSD or MacOS servers are malpractice? I think you might have over stated your case. So are you waiting for a memory safe Herd re-write? A memory safe any OS will be decades away if someone wanted to start tackling it now.
Latest examples, Windows 10 secure kernel and Device Driver protection.
https://myignite.microsoft.com/sessions/36925
Or the new Windows USB stack, written in the P language.
UNIXes, not so much beyond patching C exploits.
If that system exists, it sounds disgusting and unethical. Paying for something changes the nature of the transaction.
Windows 7 was released in 2009. There are substantial proactive improvements that Microsoft cannot feasibly backport to older OS versions even if they're still within the support period, and the corporate culture was simply different when Windows 7 was released.
If you sell it as "secure", yes absolutely. And there are many applications where you would not be even considered as a contractor for doing so. There is increasing number of situations where you want to provide hard guarantees (written in contract) about quality of the software and service you are providing. "we are using Linux" sometimes is not good enough.
I like your idea of a default mindset of internet tools should be written in a safe language with proper engineering techniques. I'm not sure I would go as far as malpractice but it might be a good stick to use to force makers into better practices.
First major security breech through buffer -overflow was in the late 80s.
So when Java came out, they played it "safe" - the language, despite having pointers will be absolutely safe. NullPointerExceptions and ArrayOutOfBoundsException will cause the program to crash rather than corrupting the stack.
Perfect.
Except it wasn't. It ended up being so "holy" that it's now banned in browsers.
So, everyone said to move to JS. Another "perfectly safe" language.
But it's too slow.
So JIT it.
Now it's no longer "perfectly safe".
Rinse and repeat.
And rust won't help here, because while the " compiler " can be guaranteed safe, the code it outputs can't (think of a C compiler written in Rust).
Maybe the solution isn't to rely on language (except for the Kernel) but to make of easy to spawn OS processes that simply have no rights to call any syscalls and limited amount of memory (or a white-listed amount of syscalls).
Take it like this:
Firefox (the browser) has full rights. It starts a process (which can only connect to the network to IP RemoteHost).
If process dies (for whatever reason) or takes too long, tell user that "sorry, sites broken".
Now, malicious code causes the attacker to run arbitrary code? Who cares? You can't overwrite the browser's code and can't break out.
The browser just has to ensure that its subprocess gives you good output.
Same with JS, CSS, or image libraries.
Java being an allegedly safe server-side language has nothing to do with it being a very difficult language to sandbox when running untrusted code. It doesn't matter what language the malware I just downloaded was written in, because it's malware and it's running on my computer. What makes Java different from C in this case is that when I run a Java application on my server, I can be more confident in the fact that I will not have a buffer overflow, even when I expose the service to the entire internet.
No. The issue is that in C (or "release (not debug) rust" (someone here left a comment that rust programs in release mode don't check for array overflows)) untrusted input leads to untrusted code (through buffer-overflow).
So you can compromise a computer through a bug in imagemagick.
>Java being an allegedly safe server-side language has nothing to do with it being a very difficult language to sandbox when running untrusted code
Java was (in the 90s) advertised as safe client-side, hence applets.
I'm excited because there will finally be a reliable way for me to apply my preferred Open/Save dialogs to KDE and GTK+ apps alike. (It's being implemented at the toolkit level and you can run an app in Flatpak with sandboxing set to permit-all to get the portal hooking.)
We have that [1] [2]. It's cumbersome to use, but it's based on the bpf JIT used for tcddump etc.
We also have a variety of other mechanisms you can combine with that to tie down a process. But even then it's hard to get this right in a way that's both performant and secure, because you have to protect every user every time, while for an attacker a very low success rate can still be worth it.
[1] https://www.kernel.org/doc/Documentation/prctl/seccomp_filte...
A safe language greatly helps to prevent the sort of bug that compromises within-process security.
End users, even those compiling from source, will still only need a C compiler. Only developers need to install the safer language (even Curl developers must install valgrind to run the full tests).
Where can you use generated code?
- For non-C language bindings (this could apply to the Curl project, but libcurl is a bit unusual in that it doesn't include other bindings, they are supplied by third parties).
- To describe the API and generate header files, function prototypes, and wrappers.
- To enforce type checking on API parameters (eg. all the CURL_EASY_... options could be described in the generator and then that can be turned into some kind of type checking code).
- Any other time you want a single source of truth in your codebase.
We use a generator (written in OCaml, generating mostly C) successfully in two projects: https://github.com/libguestfs/libguestfs/tree/master/generat... https://github.com/libguestfs/hivex/tree/master/generator
Programmatically generating C code not without problems. How can you prove that the C you're generating is free from problems solved by the safer language? Cloudbleed came from computer generated C code: https://blog.cloudflare.com/incident-report-on-memory-leak-c....
See quote from the author of Ragel in the comments:
There is no mistake in ragel generated code. What happened was that you turned on EOF actions without appropriate testing. The original author most certainly never intended for that. He/She would have known it would require extensive testing. Legacy code needs to be tested heavily after changes. It should have been left alone.
PLEASE PLEASE PLEASE take some time to ensure the media doesn't print things like this. It's going to destroy me. You guys have most certainly benefitted from my hard work over the years. Please don't kill my reputation!
And I'd like to add that what made this a catastrophic error was that different requests were served in the same address space, rather than using address space isolation as in process-per-request/fork() architectures of old. For years now many network daemon programs have been written in an event-based, single-address space style, but I have never seen the alleged process creation overhead quantified (except for maybe multi-threaded programs). Even OpenBSD's httpd disses eg. CGIs as "slowcgi" (when you'd expect the OpenBSD developers take pride in the fact that their httpd uses ASLR etc. features of the O/S rather than inventing their own ad-hoc mechanisms to defeat deterministic memory allocation in user space, and would take the opportunity to tune O/S process creation). I don't have facts to share either, I'm just puzzled that we're re-inventing O/S mechanisms in user space with performance arguments without backing this up by numbers (or are there any?).
Are such errors less likely? Possibly so, but they're not categorically eliminated. It becomes a risk assessment exercise rather than a simple thing that everyone should do. Note that it also opens the door to Java-style problems, where once the generator becomes ubiquitous, it becomes the most valuable target for exploit-hunting because a vulnerability in the generator gets the keys to all the houses.
Therefore no code written in Rust (X) executed on x86 CPU (Y) is safer than manually written x86 assemby, because Rust compiler (and LLVM) may have errors.
And well, we can actually go deeper. There is CPU frontend that is generating micro code, which may have bugs. There is also CPU backend which is executing micro code, which also may have bugs. All in all there is no hope in programming. There might be bugs everywhere so you can never be sure what your program does.
The mystical process of "programmatically generating code" in also known as compilation. The case you are describing is a compiler bug. The compiler wasn't able to generate target code (in this case C code) with semantics and/or guarantees of the source language.
Sure, you can target one compiler and be sure you'll be generating the desired machine instructions, but it can be much more difficult to ensure that your code will produce safe machine code when compiled with all possible C compilers, and the techniques used may result in a slower end result.
If you go straight from a high-level language to a compiler IR, you have a much lower risk of having to choose between either underspecifying your invariants or overspecifying them at the expense of performance.
TL;DR: C wasn't designed as a compiler IR and that complicates things.
And as a human I have had the same issues acting as a meat-implemented debugger. I had to drill through more layers to figure out why low level things happened.
By formal verification. There are ways to do so and several verified compilers already exist.
how is that different from just writing it in another language? End users who need to compile will be able to regardless of the generated C code, but the end users who need to do a _little_ modification will be given ugly generated C code! Seems stictly worse to me...
The answer is not to modify the generated code. Modify the input to the code generator to make changes.
Even when I output a warning to this effect, that all modifications to target code will get overwritten, not to check the target code into version control, the source code is already checked into version control - invariably developers modify the target code right under the comment that says not to, then they check it into version control. They then wonder why there are bugs, and their modified target code no longer works after the target code gets regenerated after the next build.
The impetus wasn't that programmers make errors but to solve the problem of repeatability. Many instances of issues can be solved once. There is no need to recreate the solution a number of times if it is already solved.
A code generator allows one to focus on the actual meta-problem, which is often smaller and easier to solve.
I know there's a sentiment here on HN against C (as evidenced by bitter comments whenever a new project dares to choose C) but I wish there'd be a more constructive approach, acknowledging the issue isn't so much new software but the large collection of existing (mostly F/OSS) software not going to be rewritten in eg. Rust or some (lets face it) esoteric/niche FP language. Even for new projects, the choice of programming language isn't clear at all if you value integration and maintainability aspects.
I think there's two major against-C groups: those of us who have worked with C for decades and those who never worked with it. I'll try and speak for those of us who've used it for decades. The popular high-level languages that have arrived since ~1995 (Java, Python, JS, C# and friends) are excellent productivity increases. In general, they sacrifice memory and performance in favor of robustness and security. For enormous software problem domains, we just don't need C's complexity or error-proneness.
Until Rust, there's been very close to zero serious competitors for C if I wanted to write a bootloader, OS, or ISR. Not even C++ could do those (without being extremely creative on how it's built/used). The ~post-2000 languages (golang, swift, D etc) can't do that (perhaps D's an exception but it wasn't an initial goal AFAICT). This is huge, IMO.
We've groaned and grumbled about how hard it is to parse C/C++ code for decades. This is a big deal for tooling. Because of the language's design, even if you use something "simple" like libclang to parse your code, you still have to reproduce the entire build context just to sanely make an AST. All of those other new languages above probably address this problem but also add all kinds of other stuff which we can't have for specialized problem domains (realtime/low-latency requirements, OSs, etc).
> collection of ... software not going to be rewritten in eg. Rust or some (lets face it) esoteric/niche FP language
IMO it's not appropriate to lump Rust in with "nice FP language"s. And don't look now but lots of stuff is being rewritten in Rust. Fundamental this-is-the-OS-at-its-root stuff: coreutils [1], "libc" [2], kernels [3], browser engines [4].
[1] https://github.com/uutils/coreutils
[2] https://github.com/japaric/steed
Maybe I should have expressed it better, but I didn't intend to lump these together.
>And don't look now but lots of stuff is being rewritten in Rust.
I'm myself cautiously optimistic re Rust, but having been burnt by C++ in the past I'm not enthusiastic about fighting language idiosyncrasies (though modern C++ certainly deserves a second look). Then there's the issue (some might argue it's a plus) that Rust is at the same time a language, a lib, and the only compiler implementation (unlike C or C++ which give you choice).
This was essentially the case for C and C++ as well for somewhere around 5-15 years a piece, depending on how you measure it.
They're significantly larger, yes -- it's a fair complaint of rust. But it's mostly because of static linkage AFAIK [1] and not "a mess of abstractions".
Things like libunwind, libbacktrace, embedded debugging symbols for backtraces, and the jemalloc allocator aren't free.
If you ask for dynamic linkage (with the caveat that Rust doesn't have a stable ABI yet), you get a ~8K Hello World binary.
It's also possible to prune down the statically-linked size by opting out of various conveniences like jemalloc. (They're working toward making the system allocator default but don't want to regress Servo in the interim.)
...and if opt into static linking with GCC and G++ (and ask Rust to make its link to libc static), Rust can actually outdo them on a Hello World.
Here's a detailed exploration: https://lifthrasiir.github.io/rustlog/why-is-a-rust-executab...
I thought C++ had naked functions and all the things you need to write an OS.
You can of course just ignore those when writing kernel code- they get ignored in application code much of the time! But I suppose at that point it could be argued that you're just writing C with a C++ compiler?
You can lose new/delete and .bss statics and still write reasonable, even "safe", C++. Rust doesn't have .bss statics by design (lazy_static emulates this for you though). new isn't necessary for the "modern" C++ safety stuff and you can write pretty good modern C++ without new. All new gets you is a nice wrapper around allocation, and when writing a kernel you can't and shouldn't allocate anyway. In Rust, too, you would not be allocating, either via memmap/malloc or via Box::new().
So it wouldn't be "C with a C++ compiler", it would be "C++ without allocations", which is a restriction from the problem statement anyway.
Further, even in C I rarely see use of "#pragma interrupt"-like tools- rather, everyone still seems just to use per-platform assembly glue code. (To be fair, my experience is mostly in kernel code for things like Linux, rather than standalone embedded applications where "#pragma interrupt" would be more valuable.)
Ultimately the Rust OSes resort to some handwritten assembly as well. I think that's going to be a constant of writing a kernel. Rust is working to minimize it (e.g. with things like `extern "x86-interrupt" fn`), but at a kernel level there are just some kernel specific asm instructions (like all of the TLB stuff) that either compiler will probably never support generating without inline asm.
So while Rust may be better than C++ at writing OSes (I'm not sure! I haven't looked at all the stuff you need to write an OS in C++), I do think they're in the same ballpark, close enough that if Rust is a "serious" competitor C++ probably is too :)
I honestly think that Common Lisp can do this quite well. It was designed to be a high-level language, but it's completely capable of working at the machine level, pleasantly and easily. Unlike C, most of the time one has safety, but one can disable safety when necessary with a simple (declare (safety 0))).
Performance is extremely good with modern compilers, although I don't know how good they would have been back in the old days.
And it makes parsing the language dead simple. Here's the entire reader algorithm: http://clhs.lisp.se/Body/02_b.htm
From what I can tell of Rust, it doesn't look easier to parse than C (but I've not looked deeply); certainly, it's orders of magnitude more difficult to parse than Lisp.
I believe that Standard ML or OCaml could do similar things as well, albeit at the cost of being more difficult to parse. Smalltalk is maybe a little less capable, but somewhat easier to parse.
They've all been around for decades.
Rust is not 100% context-free, but the feature that is non-context-free (raw strings, a rarely used feature) is still pretty easy to parse, and even if you capped it at 6-level raw strings you'd probably be able to parse all the Rust code out there.
I haven't used any lisp dialects for decades, so I have naive questions: is there really sufficient support from compilers+linkers to write a bootloader in lisp? Do I have to do a lot of bootstrapping in assembly to bring up lisp interpreter before I can execute the lisp code or does the ahead-of-time-build result in executable machine code? Can I do inline assembly (not required but a really key benefit IMO)? Are there numerous examples where someone's already written one in lisp?
The rest of this post is an excerpt from an email I sent 6 years ago.
The following comments on runtime systems are partially based on a long c.l.l thread with posts by Lucid, Symbolics, and Franz alumni.
Franz uses a 3-layer approach: CL, a low-level Lisp, and C.
Lucid started with Lisp that generated assembler but reluctantly added some C.
Symbolics Lisp Machines used bootstrap code in a Pascal-level language with prefix syntax. A Symbolics alum said that in retrospect they should have used C.
Most Lisp implementations have subprimitives - low-level functions that can circumvent the type system, often with a prefix such as % or :.
Assembly language integration dates to Lisp 1.5 and there are several common approaches.
1. turn the optimizer off - this is easy to use and implement.
2. optimize the assembler block - Naughty Dog GOAL did this.
3. annotate the assembler block with pragmas that indicate side effects. This can be error prone and difficult to use. https://www.pvk.ca/Blog/2014/08/16/how-to-define-new-intrins... is an example of this approach.
Edited to add the SBCL intrinsic link.
This should not be true, and we fought hard to keep it that way. There's one spot of Rust's grammar that's context-sensitive, for something very rarely used, and other than that, it's all much simpler.
My larger point is that there are plenty of very good reasons to criticize C/++, and parsing is a minor one since parsing is fast, and even if the creation of the AST isn't context-sensitive, verifying its correctness (is this identifier in scope?) still is.
What is the context-sensitive spot in Rust's grammar?
Ok, fair bit, it's a frustration for me but admittedly not as important as the other differences.
I mention it because it's a wart in C's language design and I figured Rust's safety features are already well-known and heavily discussed. If I want to write a simple tool "ask this tree of .c files how often they use an identifier with name 'X' or type 'Y'", I have to find out the include paths, defines, all kinds of other "noise" just to find out what could be a relatively simple query of the source base.
In what way? It's very C-like in its basic syntactic structure, just more regular.
I am not against Rust. Rust has some great ideas and intent. I just feel they should have created simpler syntax. A more complex and unusual syntax doesn't have any real benefits IMO.
What exactly is the problem with the example you linked to?
Take a look at this. It isn't as simple and clear as C/C++ for example and requires significant investment in parsing and understanding the code.
It's not code rust programmers would normally write, I'm one of those programmers. I'm glad some libraries like serde,rocket and diesel are using it to generate code instead of doing run-time analysis.
It is happening and will keep happening and it is really necessary at some point. Sure, I don't expect large project to be rewritten overnight, but every large project is being redone eventually. Especially when C becomes main source of problems. And you can introduce better languages gradually.
> operating systems, drivers
https://github.com/redox-os/redox
> IP stacks
https://github.com/QuiltOS/QuiltNet
Many of those are often written in userspace, where you are free to use any language. Several of them will provide sufficient performance.
> databases, Unix userland tools, web servers, mail servers, parts of web browsers and other network clients, language runtimes and libs of higher-level languages, compilers and almost all other infrastructure software we use daily.
For all of those you'll find Rust implementations. Some are work in progress, some are already widely used.
Everything else outside UNIX was using Assembly, Algol, PL/I, Modula or Pascal dialect.
C owes its success to UNIX's adoption by the market, as operating system available almost free of charge to universities, with source code available.
I'm reminded of a time a colleague needed something like string.split, and working in c++ he filled a std::vector<std::string> with the result. Using a more C way he'd really only have needed a couple of pointers on the stack.
Ah, yes, the good 'ole C way that only works reliably in English.
You can parse utf-8 character at a time. Some characters advance the pointer by 4 at an iteration and some less.
Yes, you can avoid this if you're careful and you understand the intricacies of utf-8 (or some other multi-byte encoding), but it very quickly stops being elegant.
IIRC C++ is getting slices too, so it might be able to get better APIs around string manip. But I've seen decent string manip code that avoided allocations.
This is pretty much how it's done in Rust too via slices. For example, the standard way to split a string is to create an iterator and it won't do any allocations.
I try to avoid manipulating strings as much as possible. ;)
Actually when I'm reading about new software, it's very rare to encounter C. Usually it's something else.
Didn't know that curl was stuck back on C89, that's really optimizing for portability.
If anyone is confused by the "curl sits in the boat" section header, that's basically a Swedish idiom being translated straight to English. That rarely works, of course, and I'm sure Daniel knows this. :)
The closest English analog would be "curl doesn't rock the boat", I think the two expressions are equivalent (if you sit, you don't rock the boat).
"Sit in boat" is a positive expression of the stability benefits.
In the curl project we’re deliberately conservative and
we stick to old standards, to remain a viable and reliable
library for everyone. Right now and for the foreseeable
future. Things that worked in curl 15 years ago still work
like that today. The same way. Users can rely on curl. We
stick around. We don’t knee-jerk react to modern trends.
We sit still in the boat. We don’t rock it.
I see a lot of inertia in there. While it's a great record to maintain 15-year consistency but in the era of every changing InfoSec outlook, it could be a legacy and baggage if the authors resist to change. One thing we know for sure is that human will make mistakes, no matter how skillful you are. In the context of writing a fundamental piece of software with an unsafe programming language, that means we are guarantee to have memory-safety induced CVE bugs in curl in the future.Some of other points that the author raised are valid too. If there is a trade-off that we can have a safer piece of fundamental software by almost eliminating a whole category of memory safety related bugs, and with the downside of less compatibility with legacy systems, more dependencies etc., perhaps we should consider it? I believe the tradeoff is well worthy in the long run and option is ripe for explore.
What he doesn't welcome is rewriting something that's had those bugs and the types of logic bugs not related to the language already worked out. There's a saying about a baby and bathwater.
Not everything is a dichotomy, and you shouldn't be reading the article as if the author is against newer languages. He specifically says that given a fresh start with the availability of these languages he might use something besides C. Carefully weighing options is wise. Throwing away years of actual progress for the appearance of quick progress is foolish.
I specially quoted the section head "curl sits in the boat" and the entire section ends with "We sit still in the boat. We don’t rock it". Now read it again, and then tell me if that's welcoming changes or resisting changes.
> He specifically said someone has or would write a competitor to curl in Rust or some other safer language and that a good one will take off. He welcomed that.
Sure there might already be some alternatives out there. But those are not curl, they are at most forks.
> He specifically says that given a fresh start with the availability of these languages he might use something besides C.
Nope, he used the word "Maybe. Maybe not." Might is a stronger word.
Not everything is a dichotomy. One doesn't have to be for all change or against all change. One can choose which things need to change.
(Note that I suspect you think that I or the original poster are suggesting that he's resistant to all change, not just change within curl. I don't believe he's resistant to all change, but I do believe he's resistant to change within curl, which is what we're talking about here.)
Could be. Or it could not be.
Somehow, Git has increased in popularity, despite its author's over-my-dead-body insistence on C.
Every few years there is a new batch of programming languages that come out and they all gain a small passionate community that tries to convince the internet how much better that language is.
They inevitably use the argument that code not written in the new language is 'legacy' and 'resistant to change'.
Neither of those assertions are accurate or enlightening unless you can provide a proposed replacement and prove the superiority of the new code.
Simply telling other programmers to rewrite there code in xyz language with such arguments is primarily a case of armchair development.
If you really think it could be done better then do it and prove it.
What I said is my believe and opinion, obviously. I don't have the skill set neither have time to do/prove it, unfortunately. But that shouldn't prevent me from speaking out my opinion, should it? Just like a lot of people really think space travel could be done better but not all of them have the capability and resources to do it and prove it.
Given the fact that programming language fanboy noise is a constant problem in language threads it would be nice to see less posts of higher quality then the same arguments that are always made. Its not really going to change anyone mind that has experience and the people who do argue are usually inexperienced with vapid counter arguments.
So I guess the answer is no. If you have an uninformed opinion then its better to not add noise and let the people who really know there topic enlighten us with a well argued not very noisy discussion.
I think the value of NH is very much quality over quantity.
Even if your language (Rust, Erlang, LISP, Go) is "better", it's still a minimal part of the equation. A maintainer is what makes the tool. It's hard work to decide which PRs to accept (and worse yet, reject), to backport fixes to platforms for which you can't get a reliable contributor, coordinating fundraising/donations, keeping up with evolving standards...
Anyway. Thank you, thank you, thank you Daniel Stenberg. Use whatever damn language you want.
On the other hand, if he didn't want his justifications for that choice examined by the world, he wouldn't have aired them, right?
> If you think Curl would be better in another language then port it, release your alternative, and maintain it for a long time.
That's out there; some languages have URL downloading objects that are not based on Curl.
E.g. Edi Weitz's Drakma client library for Common Lisp doesn't seem to be using Curl as far as I can see.
Drakma sounds great. Tone is tough to get right online. I don't like people doing drive-by suggestions like "you should rewrite X in Y". But if people are really willing to roll up their sleeves, write the tool (in any language) and keep it going for the long haul, I applaud them. I just have great respect and empathy for project maintainers, I think some don't appreciate what a huge PITA it is to BDFL a successful project.
Why would I be vendoring my own copy of libcurl in my project? Who does? This is how I (or rather, the FFI bindings my language's runtime uses) consume libcurl:
dlopen("libcurl.so")
I rely on a binary libcurl package. The binary shared-object file in that package needed a toolchain to build it, but I don't need said toolchain to consume it. That would still be true even if the toolchain required for compiling was C++ or Rust or Go or whatever instead of C, because either the languages themselves, or the projects, ensure that the shared-object files they ship export a C-compatible ABI.An example of a project that works the way I'm talking about: LLVM. LLVM is written in C++, but exports C symbols, and therefore "looks like" C to any FFI logic that cares about such things. LLVM is a rather heavyweight thing to compile, but I can use it just fine in my own code without even having a C++ compiler on my machine.
(And an example of a project that doesn't work this way: QT. QT has no C-compatible ABI, so even though it's nominally extremely portable, many projects can't or won't link QT. QT fits the author's argument a lot better than an alternate-language libcurl would.)
I particularly like the mention of portability. No other language comes even remotely close to the portability of C. What other language runs on Linux, NT, BSD, Minix, Mach, VAX, Solaris, plan9, Hurd, eight dozen other platforms, freestanding kernels, and nearly every architecture ever made?
First, I'd rather appeal to every user than most users. That one user I didn't have to appeal to is going to be a much more faithful and grateful user than the "normal" ones. Most of my software work is open source (remember this context is a discussion about curl), and this encourages active collaboration with users with niche situations. If I choose technologies that make using my software attainable for these people, odds are they aren't going to stop at just porting it to their platform.
Limiting your platforms to Linux, OSX, and NT also stifles innovation. These platforms are all deeply flawed. Their popularity isn't due to having the best design, but rather to having a good enough design and being entrenched. They're old platforms, we've learned a lot since they were started. New or niche platforms bring a lot of value to the table. The BSDs are a great example, as it's the best suited platform for a wide variety of applications.
All a new platform has to do to be able to run nearly all general purpose software is port a C compiler. Not even that - they just need a cross compiler. This is a great thing, IMO.
>Embedded can maintain their own software, they do all the time
This is a pretty silly argument. Most embedded developers don't ship their own implementation of HTTP, they ship curl!
I think one could say the same thing about C's popularity as a language.
Wishful thinking! But I suppose we'll see.
This statement is laughable nonsense. Shall we go into their bug history and point out counterexamples left and right? [Edit:user simias has done this; thanks!]
Every single bug you ever make interacts with the language somehow.
Even if you think some bug is nothing but pure, that logic is part of a program, embedded in the program's design, whose organization is driven by language.
That's wrong. A lot of the C mistakes are indeed "logical mistakes in the code", but most of them would be indeed fixed by changing to a language that prevents those mistakes in the first place.
I very much agree that rewriting existing, stable software written in C is likely not worth the trouble in many cases, but I can't accept claims that the limitations of C aren't the direct cause of tens of thousands of security vulnerabilities, either.
In Rust, even a less experienced developer can fearlessly perform changes in complicated code because the language helps make sure your code is correct in ways that C does not. And you can always turn off the safeties when you need to.
Experienced developers should feel all the more empowered by simply not having to always worry about things like accidental concurrent access, use-after-free, object ownership, null pointers or the myriad other trivial ways to cause your program to fail that are impossible in safe Rust. You get to worry about the non-trivial failure modes instead, which is much more productive.
To just use a library, rust isn't much of a dependency, either. It's designed so you don't even need to know that it's not C.
Rust would obviously be a build dependency, but that's lessened somewhat because it tries to make cross-compilation easy.
(But this point does apply to pretty much any other language. Curl would not be used as widely if it depended on the Go runtime, for instance.)
If I were ambitious enough, I'd do it myself in Haskell, but I think it'd be too much work for a simpler curiosity.
Most languages already have HTTP client libraries. (In particular, Rust has Hyper. Ruby/Python/Node/Go have HTTP clients built-in in the stdlib, Haskell has http-client, etc.) Who uses libcurl really? (Spoiler alert… PHP.)
Of course libcurl does FTP and Gopher and all the things, but these aren't commonly required, most applications just need HTTPS.
Thousands upon thousands of projects rely on libcurl...
Safer code, so half of the vulnerabilities wouldn't have existed.
Rust's ecosystem (package manager and libraries).
Of course you lose portability and you probably appeal to fewer developers, at least for now. So there is a trade-off.
I wish Rust compiled to C. It would be my dream language. The only reason I can't choose rust half the time is because it doesn't support targets I need to support.
My dream is human-readable C as well, which I wouldn't get via LLVM, i.e. a coffeescript-like mapping between Rust and C, as far as that will go.
Like I said, it's a dream. I started working on something like it, but I was side-tracked, as often happens. Maybe I'll start it again...
I don't know that you'd gain much in the real world though. Starting such a project now? Rust, for sure, for me anyway. But is it worth rewriting Curl? I agree with the author there, it most probably isn't.
Probably.
From within the community I don't see any coherent effort to tell folks to rewrite in Rust. We're very happy to see Rust rewrites, but not many people are pushing for it except the folks actually putting effort into it.
Yes, every time a vulnerability pops up someone will say "rewrite it in Rust", but half the time it's not even a Rust programmer (often, Rust programmers come and disagree and say "Rust wouldn't fix this", IME).
Also one has a hard time comparing curl with another language, simply because something with curl's properties (take portability for example) doesn't exist.
And no that isn't in defense of anything, just me thinking thinking that measurable points brought up in the discussions don't make sense or exist.
The topic is also a bit broader, as you can easily add in static code analysis, compiler flags, stuff like W^C, stuff like seccomp, capsicum, cloudabi, pledge which might not work (well) in other cases.
It's a great philosophical discussion topics and I don't wanna stop anyone, just hoping people keep that in mind, when they participate, so we don't end up with new dogmas that get thrown around for the next few year, without knowing contexts or meaning of phrases.
Other than that: I really enjoy this discussion. :)
Are such extensions popular, and if not, why not? I assume there's always some performance hit, but that might not be a big deal in an HTTP client, for example.
It's hilarious reading rust marketers talk about how people should use rust, and yet their software doesn't work as well. It has plenty of bugs.
Then they go on and on about issues which post modern C doesn't have. Guess what? C has a lot of tooling, and yes, it's been improving over the years too. CQual++ exists. AFL exists. QuickCheck exists.
Can your rust project from two years ago even compile? Does it have any users at all?
There's a formally proven C compiler. How's that LLVM swamp going you've built your castle on?
Rust brought a modern knife to a post modern gun fight -- and lost.
Post 1.0's release date, which is just short of two years, the vast, vast majority should, yes. We've had one or two soundness fixes in those times that would take a trivial amount of updating to do, but that only hit a very small part of the ecosystem.
Eh, two years have not passed yet since Rust 1.0.
But yes, it is likely that your Rust project today will compile just fine in 2019.
http://www.cis.upenn.edu/~stevez/vellvm/
> It has plenty of bugs.
Rust doesn't claim to prevent all bugs. It only claims memory safety.
IMHO it might eventually make sense to use other language/tech/whatever but the bar is quite high and it will quite probably take some serious sustained effort.
Don't rewrites, even in the same language usually lead to a better version of the software? I can't really imagine a seasoned C developer introducing completely new bugs in a code base they are already very familiar with
Get a CVE list for some Linux distribution, this will open your mind. Happens all the time.
If it think about my usage, it's like get or post something and see what the returned json looks like. If I need to download something wget usually works without having to remember -O.
But higher level things like httpie are easier to deal with, sane defaults and all that. Maybe they use libcurl...
Are there any re-write userland in ${safe-high-level-lang} projects?
Although, curl in C++ - the naming would became inappropriate...
The fact is, mistakes will happen, but in general if you follow the best practices you'll be fine. Failing to follow the best practices means you could be a better programmer. Just because the language gives you an option to do something, doesn't mean you should.
But in practice, people always make mistakes. Some guns/programming languages limit the damage of those mistakes more than others. All other things being equal, those guns/programming languages are better.
Everyone fails to follow best practices at times, precisely because "mistakes will happen!" Encouraging people to be better programmers can't change that.
Therefore we want languages that catch as many mistakes as possible at build time and minimize the damage caused by other mistakes that slip through.
https://jobsquery.it/stats/data.technologies/C
Stats also showing that average salary for C developers is above average for all tech job openings.
Curl is small enough to make it relatively easy and used widely enough to make it worthwhile.
The community should do it, spend a couple of years stabilizing, and then spread the words to others.
Telling others to use their language instead of putting their money where their mouth is is truly what irks me about the rust community the most.
Want a rust world? Go write it and ship it.
Oh, and you don't get to complain about C until your PC runs more rust than C
Cool?
It's okay to identify flaws in tools, that's how we make them better. It doesn't make sense to say '[car on fire] you can't complain about that Honda Civic until you develop your own better car, sir, and more people are driving it than not, now please leave the service center -- until then consider the fire normal'.
If I cut myself by hasty use of a knife, is it the fault of the knife maker? How is that even remotely rational? If you aren't willing (or don't know how) to use the tool correctly, don't use it.
Completely false. C is a disaster.
For crying out loud, this is patently false.
C isn't as dangerous if you really know what you're doing and you never make mistakes. I would bet fewer people know what they're doing than think they know what they're doing, and the set of people who never make mistakes is entirely empty.
> Every once in a while someone suggests to me that curl and libcurl would do better if rewritten in a “safe language”.
And so he minimally addresses that and some other reasons for sticking with (and originally choosing) C89.