The Curious Case of the Longevity of C
ahl.com
ahl.com
I think C is still so pervasive for 2 simple reasons (I might be wrong):
1) It's a small language. Compared to many newer languages (Swift, modern C++, old C++, Objective-C, Java), it's quite tiny. The language itself can be learned in short time. (Which is deceptive, because properly using what you have learned takes longer than a short time).
2) It's fast. It's just direct memory access, direct and explicit hardware manipulation. And whether justified or not, there is a very large portion of the developer community that remains obsessed with speed. I see it all the time, even when speed isn't not going to matter enough to justify C. I often see people going way out of their way, making their work much more time-consuming, because of their everlasting pursuit of writing the fastest damn program they possibly can.
Problem is you have the complete opposite on the other side, folks who don't even think about speed (usually because their development machine is quite powerful and so they don't notice). Thus we end up with programs that are resource hogs and slow that have no good reason to be.
There is a middle ground to be found I think.
I've never felt that way with Python, JS, Ruby, Java, C++ or any of the other languages I've used. They are at least 10X more complex in terms of primitives and non-trivial interactions.
I love C for its warts and all, but you should treat it like a tamed wild animal.
> The current crop of C and C++ compilers will exploit undefined behaviors to generate efficient code [...], but not consistently or well. It’s time for us to take this seriously.
I assume now it's more aggressive, because undefined behavior should never be relied upon. That's the point.
The trend toward UB-based optimizations seems significant to me because they can and will take a bunch of C code which looks superficially reasonable, and which used to do X as intended, and suddenly (and perfectly legally) make it start doing Y instead. I assumed that's what barrkel was alluding to above.
And yes, I'm sure there are old optimizations which will also do that, but basic things like hoisting and reordering and unrolling and dead-code elimination don't fall into that category.
So then it's not new behavior.
And indeed, I do always use the increment operator by itself. The amount of extra thought it takes to distinguish
while(i--)
from while(--i)
just isn't worth it, when instead I could write for(;i>0;i--)
which is very clear about what the last value of i in the loop is.Actually there are a few other languages that are great for debugging runtime code behavior; Lua comes to mind.
I also find C really handy for parsing binary formats (ASN1, Ms EMF...), parsing binary stuff in Python for example feels miserable. However, for manipulation string based formats (xml, json, yaml), C is quite miserable, even with the help of libs like libjson-c or libxml2.
And to be fair, even for parsing binary formats, C can be quite unforgiving. I still have some traumatic souvenirs from the first time I ran American Fuzzy Lop on my EMF to SVG converter... (by the way, I love the afl logo: https://upload.wikimedia.org/wikipedia/commons/f/fa/AFL_Fuzz...)
You can treat C as a small language, if you don't have to work with anyone else.
http://www.open-std.org/jtc1/sc22/wg14/www/docs/n1548.pdf
That's up from 95 in C90, a factor of 2 not 3 in 20 years.
http://read.pudn.com/downloads133/doc/565041/ANSI_ISO%2B9899...
You can treat Rust as a useful language as long as you don't need to get work done. I rather like C11. I would have liked Rust, I tried, if they'd have known when to say no or if they valued sparseness for its readability because it reads like C+++:
const USAGE: &'static str = "usage: blah blah";
By contrast, I like what seL4 did with C. They wrote in C and proved in Isabelle. const USAGE = "usage: blah blah";
The RHS has sufficient type information, does it not? After all, if this were a let USAGE = ... inside a function, that would be sufficient.(I don't find this particular bit a deal-breaker though; thus far, I quite like Rust.)
const USAGE: &str = "usage: blah blah";
Compare it to the equivalent in C: const char *USAGE = "usage: blah blah"; #define USAGE "usage: blah blah" const char * const USAGE = "usage: blah blah";The C11 standard (ISO/IEC 9899:2011) adds support for threading and atomics, and adds IEC 60559 specifications in its annex.
Yet, the language specification part of ISO/IEC 9899:2011 is less than 180 pages, and the remaining pages are dedicated to libraries and the annex, which on itself takes over 200 pages.
Ah, compared with its early 1980s competition (Pascal, Modula-2, BASIC), C was quite lavishly equipped with library functions.
Now, get off my lawn, kids!
Therefore, technically, C++ has slower compilation time but potential for faster runtime.
One sees plenty of news on the downsides, but on the other hand the lack of security has also enabled console homebrew, iOS jailbreaking, Android rooting, and a bunch of other creatively liberating activities which involve doing something not officially sanctioned. It's worth pondering whether we would be better off, had everyone switched to something much "safer" (maybe even with formal verification) a long time ago. I think it's debatable.
But slowly the pool of exploits will dry up, either because we stop using C or because we finally figure out a way to write C securely, and we'll find ourselves decades behind at the necessary task of producing hardware that is open by design rather than open by accident.
Maybe the most secure language is the one so esoteric, and complex, and incomprehensible that no one uses it to make anything and therefore is never exploited.
I hate that it is okay to lie. Of course any OS can have a virus and malicious code. It runs code.
The problem with C is that it's widely used for critical system-level software, and practically all other tools were (and are) inadequate for the task.
On the other hand, without the pervasive vulnerabilities enabling this hacking, we might have been more motivated to enact a legal framework to require companies to permit people to hack devices they own.
Python ≥ 3.6 requires (some features of) C99.
> Perhaps the reluctance to move to a more ‘modern’ C is that even C99 can’t legally buy a drink
Perhaps it's because MSVC didn't fully support anything newer.
(some of the changes) http://www.open-std.org/jtc1/sc22/wg14/www/newinc9x.htm
C99 support was added in VS2015 to the extent required by C++14, similarly C11 support was added in 2017 to the extent required by C++17.
You don't work in the enterprise domain which is VS’s core market, do you?
Usually it takes one year to upgrade to a new version, after its release.
We are already playing with 2017 for next year upgrade cycle.
For me, Linux is a huge development environment dedicated to C programming. Chapter 2 and 3 of manual pages gives C api. C gives the most direct access to kernel. IMHO we should try to find a replacement for C and the first target for this language should be linux kernel. If you can conceive a language that Linus Torvalds accepts, the rest of the planet will follow ;-)
Good luck!
So for a long time if you wanted to do low level, fast, small, reasonably portable code you simply didn't have much of a choice.
Rust is the first language in a long time that I think could end up replacing C in all of its use cases but there's still some way to go. The main difficulty I can foresee is that like C++ it's significantly more complex than C and harder to interface with other programing languages (you end up making an unsafe C interface).
I think C is here to stay. The billions of lines of C code out there won't be rewritten in a fortnight.
ABIs, APIs etc have to be able to propagate size and type information that is basically how C defines them if you want anything to work (pointers, integers, floats).
C embeds a remarkable amount of cross-object-file information about a function just with 'int f();'. We know the range of valid values, we know how big the object is, etc.
Rust literally has none of these mechanisms and likely never will, interfacing Rust to other Rust, such that you will get any benefit at all from using it would require megs of type, scope, size, etc information to crawl across the calling convention.
There's a reason C++ names look like CintFuncf!!!@34$902, Rust doesn't have or plan to have anything like this yet, and it'd need to be 20x more complicated.
If you get rid of Rust's magic about being able to remove guards and checks by infering things about the data to make it conform to calling conventions instead of using types, you have just invented C.
I disagree somewhat with your assertion that:
> C embeds a remarkable amount of cross-object-file information about a function just with 'int f();'. We know the range of valid values, we know how big the object is, etc.".
First of all this is not embedded in the object file but rather in the header, Rust doesn't need that. The object only tells you the name of the symbol and that's about it. In particular that means that the C linker can't detect ABI mismatch if the prototype of a function or the layout of a struct changed, as long as the symbol is found it'll link just fine.
Furthermore even with the header available a lot of the time C prototypes are not sufficient to know how to use a method. Take for instance:
sometype_t *some_function(someothertype_t *param, int flag);
Is param an input or output parameter? Do I own the return value or is it allocated by the function? Or maybe it's just a member of param so it has the same lifetime? I know that flag is an int but that doesn't really tell me which are the valid values I can put in there. In Rust function signatures tell you all of that and it's enforced by the compiler.So yeah, there's a significant overhead in Rust here, but it's for a good reason IMO. It does make it harder to make quick hacks with the linker though.
> There's a reason C++ names look like CintFuncf!!!@34$902, Rust doesn't have or plan to have anything like this yet, and it'd need to be 20x more complicated.
Are you talking about name mangling? Rust does that too but that's not really the same issue, it's just about generating unique "flat" names for objects that include namespacing and generic info. Like if you have a "fn foo<T>(t: &T)" and you instanciate it with T = i32 and T = String you need to generate two symbols. C doesn't need that because it doesn't have namespaces, generics or overloading.
Granted, pointed about the header, I was talking more about the 'int' part but failed to describe myself well at all. I mean more that you can map it onto the ABI properly and easily. Try shuffling internal Rustisms over the usual SysV ABI.
What Rust means by anything is still entirely up in the air, and last I checked their bug tracker, the issue had been open for years and was basically commented as "probably won't ever resolve due to rust changing too often".
>Furthermore even with the header available a lot of the time C prototypes are not sufficient to know how to use a method.
A keyword could be added for that, you wouldn't be able to enforce it on the ABI level but you could make the compiler check the same as Rust. I'm not here to argue about which is better though, I think I accidentally led you up this path though.
>Are you talking about name mangling? Rust does that too but that's not really the same issue
I need to know how to generate those 'flat names', is the point, I can't call it if I can't name it, and it all plays in from the same systemic issue in that Rust has no idea what it is doing from week to week.
EDIT: I'll use this space to re-iterate, if you get rid of all the Rustisms and shuffle everything over the existing SysV ABI, you're left with C. Keeping the Rustisms will need megs of data sending over the ABI somehow to allow Rust to elide bounds checks, etc that Rust is useful for to begin with.
Do you have a link? There's a good chance that issue was opened and that comment made before 1.0, when things really were changing too often. Change is less common these days, though I'm sure the compiler developers do still enjoy not being restrained by the need to maintain ABI backwards-compatibility (which isn't to say they couldn't be convinced, only that it would take cajoling).
> it all plays in from the same systemic issue in that Rust has no idea what it is doing from week to week.
What do you mean by this?
> Keeping the Rustisms will need megs of data sending over the ABI somehow to allow Rust to elide bounds checks, etc that Rust is useful for to begin with.
I'm confused here. I don't see how the ABI has anything at all to do with bounds checking; bounds checking, where it exists, is all done at runtime by regular code in the standard library. And by "etc" I presume you're referring to type inference (which you mentioned in an earlier comment), but Rust doesn't do inter-procedural type inference; it's not like Haskell or ML. There's simply nothing to infer, and no need that I can see for an ABI to care about that. Can you be more concrete?
Overloading as in operator overloading? Because I don't see how that would affect symbols.
Though along with namespaces and generics, there is one thing that Rust also bakes into symbols: versioning information. This is how, in the case of deep dependency graphs, it's possible for a finished binary to include multiple copies of the same library in the event that multiple versions are transitively depended upon. But that doesn't add any complexity to symbol mangling on its own, because if you already have namespaces then you can just treat it as a namespace that only the compiler can see.
So like:
int do_something(int param);
int do_something(double param, char *param2);
I don't see how you can avoid some form of name mangling since obviously you can't just define two duplicate "do_something" symbols.(This isn't the only way to end up with duplicate symbols, just trying to make it clear that this specifically won't be a problem with Rust.)
Ugly? Yes, IMO, but it also is useful.
rustc --crate-type staticlib file.rs -o rusty.a
gcc file.c rusty.a
./a.out
It totally works (and debug symbols, C debuggers, profilers, code coverage — all works in the mixed executable).Of course C can't directly call Rust's ABI, but Rust can call and export `extern "C"` functions, and there are tools to automate this. You can declare C function prototypes in Rust using references instead of raw pointers, so you even get basics of borrow checking for the C code.
It's true that Rust has no stable ABI, but "likely never will" is hyperbole. If people demand it in sufficient quantity then it will happen in time. I personally hope it does someday, but I'm in no hurry.
> There's a reason C++ names look like CintFuncf!!!@34$902, Rust doesn't have or plan to have anything like this yet, and it'd need to be 20x more complicated.
Rust does have name mangling already. It may not be standardized (again, no stable ABI), but C++'s name mangling isn't standardized either. And I see no justification for why name mangling in Rust would need to be "20x more complicated" than in C++; AFAIK Rust symbols pretty much already imitate C++ symbols in order to play nicely with existing tooling.
This would be more informative if you gave some concrete examples. Any hypothetical Rust ABI would be more extensive than the de facto C ABIs, undoubtedly, but the concern here seems unjustified.
You hit the nail in the head. The go-to tutorial and reference book of C is «The C programming language» by Brian Kernighan and Dennis Ritchie, which provides a complete and very thorough description of C and its standard library in less than 260 pages. That's unbeatable.
As a comparison, the go-to book for C++, «the C++ programming language» by Bjarne Stroustrup, goes beyond 1200 pages and doesn't cover some fundamental aspects of C++, and even «The Rust Programming Language» by Steve Klabnik and Carol Nichols, a book on a programming language which was designed to eat away C's market, is over 400 pages.
This speaks volumes on the effort required by anyone to get on their feet and be productive with these programming language.
Then, you may also consider the framing of "simple" vs "easy", they're not the same thing. And that's even if we agree that C is simple in the first place, which I personally consider not true.
Basically, I don't think that comparing page counts of random documents says anything meaningful about language complexity.
That said, not everyone likes my writing style, so I'm glad that there are other books coming out as well.
ISO/IEC 9899:1999 is 554 pages, and the annex alone takes around 140 pages.
ANSI ISO 9899:1990 is even shorter: 230 pages.
Really though, this just furthers my point; page count is a terrible metric for this.
I'm not counting C++ because it was piggy-backing on C.
Coming from Java, my first experience with C was trying to write a trivial program and seeing nothing but the text "Segmentation fault" at runtime. Expecting a nice Java-like backtrace to tell me where the error was, I raised my hand and asked the TA what a "Segmentation fault" meant, and how to get a backtrace. He laughed, rolled his eyes, and went back to playing Slime Volleyball.
Compared to any other language in use these days, C is anything but approachable. And when it comes to learning "secure coding standards in C", as the grandparent comment mentions, one cannot risk putting that off lest they develop bad habits that are never undone (though other environments share this property as well to some degree, e.g. PHP and client-side Javascript).
God, this is the first time in over a decade that I see a reference to Bloodshed Dev-C++. IIRC, Dev-C++ was just an IDE, and the compiler was actually GCC. The IDE was actually very good for its days, and considering it was a free IDE before free IDEs were a thing.
You end up building up an impressive set of street smarts dealing with C, based almost wholly on actual bugs found the hard way. This period last for a long time, if it ever ends.
That may be some of it's charm. You've bled copiously along the way, due to somewhat disguised behavior. Why give up your hard won knowledge just to go play with legos? It's like learning first to juggle knives and hatchets, and wondering if you should "progress" to safe juggling balls specifically designed not to maim you.
I'm joking, mostly. I love programming in C, but it's definitely not an easy or simple overall journey. You can make some cool things with it, though, which is the point.
Edit: fixed pronoun referencing gp...
Not really. Rust will tell you your program is unsafe at compile time, whereas c will tell you at runtime. The major difference is that c sometimes won't tell you about an unsafe program, whereas rust will call perfectly safe programs unsafe. Note that for a beginner, the former is vastly preferable, because it allows them to iterate much faster. It doesn't matter if your code is buggy if you throw it away, and it doesn't matter if your code has security bugs if it will never be worth hacking.
How much stronger do you want your type systems to be?
If you wish to compare books on specialized topics, the book on C++'s STL by Jossutis spans over 1100 pages and the book on C++ Templates by Vandevoorde, Jossutis and Gregor spans over 800 pages.
I think you mean, they think they learn it in a couple of weeks, and then over the next 10 years are continuously surprised at how little they actually understand C.
Not so sure about that. I started on assembly and then C, though did neither professionally. But I'd be a bit scared of touching a serious C codebase now. There's so much to think about there that only long experience can really prepare you for. Sure you can cover the language very quickly with K&C (and what fun compared to much dev learning). But although I'm out of touch with the C world now, I suspect it would take a long time to get from there to being more useful than dangerous.
C has longevity b/c its compact and provides a straightforward model of memory on the machine. I understand the desire to use safe, garbage collected, memory safe language when you're serving HTTP requests, but sometimes you need to access the hardware: twiddle a GPIO or read from a DMA device. This is where I've yet to see a good replacement for C, and by extension, C++ (b/c its fundamentally still just C). Maybe rust is there, but I don't have experience there to judge.
[edited for clarity]
I agree with your comment but hate this comparison. C and C++ are completely different languages. Idiomatic C++ looks nothing like C and vice versa.
The embedded world may start to move to doing more things in a higher level language to speed up development in areas that may not need the same level of high performance or real time certainty. Stuff like TI RTOS while still C made getting to a proof of concept much faster.
That said, the new AVR support includes 8 and 16 bit chips, I don't know the lowest amount of RAM that has had actual code running on it yet.
https://news.ycombinator.com/item?id=14071282
In my view the industry inertia is a chicken and egg thing. You want a chip that runs your existing code. Then you want to write new code on your existing chip. And you have ongoing projects at different stages, sharing chips and code.
I've also heard that Go may a viable option, at least on ARM, and I'm considering giving it a try for my next embedded side project.
https://forge.rust-lang.org/platform-support.html is said target list today. AVR is coming soon, it works but has bugs and is hacky.
Japaric is quite prolific, he writes over at http://blog.japaric.io/ and his latest posts are about his RTOS-like framework.
Finally if you're into ARM, https://www.tockos.org/ might be up your alley.
I really need to write better docs on all of this...
The three big questions in C are "how big is it?", "who owns it?", and "who locks it?" The language gives zero help with all of those issues. Most later languages deal with some or all of those issues.
Yes, it does exactly what you tell it to do and that's dangerous. It was a high level language 25 years ago, but today the metaphor should be assembly. Nobody would criticize assembly for letting you shoot yourself in the foot, C is much the same.
The biggest threat to the tech sector is Linus dying and being replaced by some charlatan who insists on replacing C and using Jira.
</offtopic>
As it is, lots of major C applications already do this on a piecemeal basis and use one or more of:
-fno-strict-aliasing
-fno-strict-overflow
-fno-delete-null-pointer-checks
with gcc to keep a lid on the effects of undefined behavior (and often more or even stricter options, such as -fwrapv).The problem is that even experienced C programmers can get tripped up by undefined behavior all the time. "Yes, it does exactly what you tell it do" doesn't really cut it when the vast majority of actual humans cannot safely predict the effects because they're so unintuitive. Or, in the words of Douglas Adams:
"But the plans were on display . . ."
"On display? I eventually had to go down to the cellar to
find them."
"That's the display department."
"With a torch."
"Ah, well the lights had probably gone."
"So had the stairs."
"But look, you found the notice, didn't you?"
"Yes," said Arthur, "yes I did. It was on display in the
bottom of a locked filing cabinet stuck in a disused
lavatory with a sign on the door saying 'Beware of the
Leopard'."Speaking of which, a killer feature I'd like to see in a C compiler is a flag for debug mode which SIGABRTs whenever it enters an undefined behaviour codepath. For example, sometimes the compiler knows that if pointers alias, or an integer overflows or something, it can do whatever it wants. I want a flag which will automatically add assertions to the generated code, so at least in debug mode it'll crash if thats ever actually the case.
As far as I can tell, -O0 is the only way to avoid that particular mis-optimization in Clang.
Edge cases that have evolved into it do have plenty of problems, but it's very difficult to create a new backwards-incompatible standard which is just a little bit different.
Then again, the embedded (microcontrollers) world seems to be filled with slightly esoteric implementations of C and everybody gets along just fine.
void copy(int* in, int* out) {
for(int i = 0; i < 10; ++i)
{ *out = *in; ++out; ++in; }
}
Now, in C it is undefined behaviour if either o in or out point into the stack where the variables i, in and out live. Maybe their values will change, maybe they won't. In practice, this code will be optimised to some code that just stores in and out in registers.Now let's imagine we want to make it defined. Then the statement '* out = * in' would have to become something like:
1) Write the current values of i, in and out to memory (in case * in overlaps with any of their locations.
2) Read * in into a register
3) Write *out from that register
4) Re-read i, in and out from memory, in case they just changed.
Now, you could make that more efficient by adding to the start of the function a check which looked to see if either 'in' or 'out' pointed anywhere near the stack, but there is still going to be a significant cost. If you plan for this from the start with your language, you can drive the costs much lower.
The main thing that separates C from assembler is that you do not have to worry about what values are in registers, and which are in memory. However, that means any code which writes to memory provides a place where this can make a difference.
"The compiler can always assume writes to memory via one object don't change values in memory accessed via any other object."
Then, if you know your code can overlap, you can use some invalidate function to mark that "memory accessed via this object was changed by some statement prior, make sure to re-read values cached in registers". Because that's the odd case where you're doing something unusual. In general, "The compiler is allowed to optimise things assuming that spooky action at a distance isn't happening". You already have something very similar for atomics, which might be hidden behind a standard library, but in any case are a sequence of "lock the bus, read/write, memory barrier", telling your CPU and your compiler that memory access cannot be reordered around the memory barrier.
Seems to me that that could only happen if you had some other bug clobbering the values of in and out, or you're passing something to the function that you shouldn't be. In either case, it's an error on the side of the programmer.
> If you plan for this from the start with your language, you can drive the costs much lower.
It shouldn't be the job of the language or compiler to make up for the shortcomings of the programmer besides maybe code generation optimization.
In fact, you can't express this in C (or, restricted, whatever), so I do wonder if your assumption that because it's UB it'll compile to the optimal code is true.
I'm talking about something else, the issue of 'in' and 'out' pointing into the stack, so writing to them changing the value of the local variables. That is basically treated as UB by every C compiler, and no compiler I am aware of provides any switches which would make it easier to handle, it's just horrible and your code behaves in different ways depending on the compiler and optimisation level.
But embedded teams often severely limit the language constructs you may use, in order to avoid the foot howitzers and only have to worry about the foot handguns.
There are also many other techniques (such as pre-allocating all variables during initialisation) which can be considered a mitigation techniques for the many traps C offers.
I'm also skeptical of the quality of a lot of embedded software (and I've seen quite a bit of it).
This is not to say that C is bad, or that I disagree with you. But embedded is a bad example.
See Friendly C:
But I bet they would criticize programs for being written in assembly, if they didn't need to be.
If you could have a language with all the performance of C without the footguns, why wouldn't you want that?
I've yet to see a language that actually delivered on this claim.
Rust delivers on all these claims. And there have been others before it. Rust hits all the sweet spots for me.
Example for the first case: Writing a garbage collector runtime in Rust has most of the same problems in Rust as in C, because you have to write most of it in unsafe code, where Rust inherits much of C's undefined behavior w.r.t. pointers via LLVM. In short, you have largely the same problems and have added a hard dependency on Rust.
For high-level work, almost all [1] of what Rust gives you is memory safety and that comes at the price of dealing with a LOT of extra language complexity. But aside from dynamic memory management, memory safety isn't hard (we did that back in the 1970s and 1980s), and for dynamic memory management, we can get memory safety with a garbage collector and much less complexity. So Rust is primarily of interest for those use cases where garbage collection is not an option.
While that still gives you plenty of interesting use cases for Rust, there are also plenty of programming niches that it serves poorly.
[1] People will also mention "fearless concurrency", but guaranteeing the absence of data races is not hard. That more languages don't do it is partly because they simply neglected that aspect [2], but also because any mechanism – including Rust's – for doing so inherently constrains your options w.r.t. concurrency [3]. Plus, avoiding data races is the easy part of getting concurrency right.
[3] Concurrent Pascal had guaranteed absence of data races in the absence of pointers in the 1970s, Eiffel had done it with pointers in the 1980s, and there was a plethora of research in the 1990s to do it in various other ways.
[3] For example, there are plenty of use cases, such as certain idempotent operations, where data races are not only perfectly safe, but also desired for performance. There are also use cases where you can prove that no data races occur, but a type system cannot easily capture that.
It's too high-level for a lot of low-level work, and too low-level for a lot of high-level work.
This is true in the very specific cases that you gave, but I believe that is the minority of use cases, not the majority.Even the example of writing a GC that requires tons of unsafe code, that is not a good argument for making all the code unsafe. All the unsafe GC code would be abstracted away into a module and would be more obvious to those looking at it that they will need to be watchful for undefined behavior. Now you can proceed writing the rest of the project in safe, simple Rust.
People will also mention "fearless concurrency", but guaranteeing the absence of data races is not hard
Maybe for developers that are very familiar with the race conditions of parallel code, but definitely not for most people. Even seasoned developers will make mistakes with simple multithreaded code.Also, the reasoning behind "x is easy so why do I need my language to check it for me" is questionable. The whole point is that you have a guarantee. Have you never had a compiler catch a stupid mistake before it happened and felt relieved? I doubt it. Now imagine if instead of debugging stupid data races in your parallel code you can spend that time optimizing and improving it. I fail to see how this can be viewed as negative.
Sure Rust doesn't cover 100% of use cases, but it definitely covers more than you're implying. It's low-level enough that Redox OS can be written in Rust, but high-level enough that Firefox is now outpacing other browsers and parallelizing everything with Rust.
That code that could be "abstracted away" would be "virtually all the code" in my example.
> Maybe for developers that are very familiar with the race conditions of parallel code, but definitely not for most people. Even seasoned developers will make mistakes with simple multithreaded code.
I'm not talking about manually guaranteeing absence of data races. I mean absence of data races as a language feature.
> Also, the reasoning behind "x is easy so why do I need my language to check it for me" is questionable.
This is not at all what I was talking about. You completely misunderstood me.
We find that Rust’s safety features do not create significant barriers to implementing a high performance collector. Though memory managers are usually considered low-level, our high performance implementation relies on very little unsafe code, with the vast majority of the implementation benefiting from Rust’s safety. We see our experience as a compelling proof-of-concept of Rust as an implementation language for high performance garbage collection.
[0] http://users.cecs.anu.edu.au/%7Esteveb/downloads/pdf/rust-is...
> We found that the Rust programming model is quite restrictive, but not needlessly so. In practice we were able to use Rust to implement Immix. We found that the vast majority of the collector could be implemented naturally, without difficulty, and without violating Rust’s restrictive static safety guarantees. In this paper we have discussed each of the cases where we ran into difficulties and how we overcame those challenges. Our experience was very positive: we enjoyed programming in Rust, we found its restrictive programming model helpful in the context of a garbage collector implementation, we appreciated access to its standard libraries (something missing when using a restricted language such as restricted Java), and we found that it was not difficult to achieve excellent performance. Our experience leads us to the view that Rust is very well suited to garbage collection implementation.
At a severe cost in performance. Static object lifetimes cover 99.9% of a garbage collector's use cases, without the performance cost of GC, nor the nondeterministic runtimes. We first saw static object lifetimes come into their own with C++'s value semantics; Rust refines and clarifies the idea and makes memory safety an inherent part of the language itself.
Static object lifetime is to GC what static types are to dynamic types. Lisp is 1960s tech. It has failed, and been replaced with something much better.
For starters, we have to assume that we don't deal with value types (which will end up on the stack, one way or the other), but with local variables that reference heap objects. Second, we have to distinguish between tracing and reference-counting GCs.
A modern tracing garbage collector will have cost for such temporary allocations comparable to `alloca()` and those allocations will typically be inlined. The cost of deallocating a short-lived stack object is zero (yes, zero). This is possible because GCs (unlike manual memory management schemes) are compacting. Whether one approach or the other comes out on top is very situational.
More importantly, I dispute the 99.9% as a vast exaggeration. There are plenty of important use cases (such as persistent data structures, shared caches, etc.) where unique ownership is insufficient; Rust requires you to use either copying or reference counting when you run into shared ownership scenarios, both of which are more expensive than tracing GC (naive reference counting is already one of the more expensive memory management methods known, and atomic reference counting is especially expensive).
If you use reference-counting GC, then for any program that satisfies Rust's borrow checker, the optimizer can eliminate reference counts that satisfy the same conditions (assuming that the optimizer knows about them because they're part of the language semantics). This is largely what Swift does, for example.
Finally, there is deferred reference counting, which incurs only trivial overhead for objects with automatic lifetime (on the order of a fraction of a percent). This is because this algorithm incurs real cost only when pointers are written to global or heap locations; this is also why it's seen limited use in practice: it's excellent for objects with automatic lifetimes and does not rely on the generational hypothesis, but those do not constitute 99.9% of all use cases. If they were, deferred reference counting would have a far more prominent role.
This does not even account for the fact that when there is overhead, that overhead is generally trivial in an imperative language with value types.
There are use cases, of course, where a tracing GC is an inappropriate choice, but that is not because of throughput. Tracing GCs make interoperability with other GCs different, for example, and have implicit memory overhead that may be prohibitive in large applications such as a web browser (that can easily consume gigabytes of memory on a laptop or desktop machine). That said, there are alternative approaches to garbage collection that do not have those problems.
Rust is not a replacement for C in the sense of portability. People love simplifying the world into Windows/macOS/Linux, but that is not all you may want to target.
Nowhere in this sentence do I see "for every platform C targets".
Given no other requirement other than C without the footguns (memory unsafe) is there a good reason not to use the safe version? I'd say there are some, but they aren't crazy compelling (ie: developers have to learn rust, maybe harder to hire for, etc).
Tests can't establish the absence of bugs the way Rust can. You only think your C code is reliable, you don't actually know that it's reliable. Rust only appears harder because of the latent bugs in your C program that you're not aware of.
For simple cases, this may be true, for complex cases it but shifts the cognitive load up front. Which may be more frustrating, but also may prevent a large class of hard to identify, intermittent in manifestation, bugs from getting into production. Which also saves programmer frustration.
That's like saying a SawStop (http://www.sawstop.com/) lacks a simplicity of a sawblade. I mean - Yeah, sure, but sawblade also won't stop you from turning yourself into amputee.
I understand Rust can be overly verbose, but main complexity comes from the borrow checker and the effect adding another kind of type, to track lifetimes. The lifetime system is the main selling point.
There are other sources of complexity in Rust, but I am glad to say both Rust/Scala seem to be looking for a way to simplify things.
No it doesn’t.
One reason is rust doesn’t have SSE/AVX/Neon intrinsics.
Without them, you can’t get anywhere near all the performance of C, nor anywhere near advertised performance of any modern CPU.
IMO, this isn't even the right goal. The problem is that in many cases using C is itself a premature optimization. Consider two languages:
* C, which gives you performance by default at the cost of an entire arsenal of footguns with esoteric nonlocal trigger conditions
* Hypothetical language X, which has all the same footguns but engineered for safer triggering and locked up in a safe that you have to choose to open when you want C-like performance
I'd rather have hypothetical language X (which is an accurate portrayal of many real existing languages), because it's got better failure modes. Performance issues are less impactful, in general, and more importantly they are more obvious. It is usually easy to tell when code is not fast enough. The endless parade of CVEs ultimately deriving from memory safety issues, often decades old, is living proof that misuse of the footguns is frequently far from obvious.
A language that avoided it would have to be very close to C while making it only mildly more difficult to foot shoot. Everyone who is trying is overshooting and therefore not really writing an adequate replacement low level systems language. Even if they did, it would be so close to C that adoption pickup would be low, awkward, and the language would fail outright. (Think Python 3 but worse)
>But I bet they would criticize programs for being written in assembly, if they didn't need to be.
Certainly. If you're writing your web app in assembly you are very likely a crazy person unless your goal is to do something ridiculous. If you're writing some assembly in a critical path in a web server to precisely control the network stack, you might not be a crazy person (see Netflix tech blogs).
Says who? A language where every single memory access did NOT have the potential to be a buffer overflow would be much more difficult to foot-shoot with, even if it allowed you to sometimes access memory with raw pointers. Just having the ability, as in Rust, to mark code as either safe or unsafe goes a long way towards preventing footguns.
>Everyone who is trying is overshooting and therefore not really writing an adequate replacement low level systems language.
Using Rust as an example again, I don't think there's anything about being able to say "this code right here cannot cause memory unsafety" that precludes "low-level systems" programming.
Well, I'm quite bored of the endless stream of memory safety vulnerabilities that can be systematically eliminated with memory-safe languages like Rust. Note that every single one of the CVEs disclosed yesterday in dnsmasq are memory safety violations, and the fact that buffer overflows, data races, and other related errors are not only pervasive, but also tend to sit around for years in codebases[1].
> Nobody would criticize assembly for letting you shoot yourself in the foot, C is much the same.
It's completely acceptable to criticize the choice of any tool. I can and do criticize the use of C for any code that touches a network and in some cases consider it willful negligence, as using a non-memory-safe language to parse untrusted data is asking for trouble.
[1]: https://twitter.com/johnregehr/status/914663997647069184
Therefore, it follows that re-writing parsers in memory-safe languages would provide a nice bang-for-buck.
That's fine... but the billions of dollars spent on security and the massive trove of user data that's been exfiltrated due to C is kinda a problem regardless of whether you are bored or not.
> Nobody would criticize assembly for letting you shoot yourself in the foot, C is much the same.
I wouldn't criticize a developer for writing bad code. I'd criticize the language for letting them, or the developer for choosing such a language. Same with C and Assembly.
> The biggest threat to the tech sector is Linus dying and being replaced by some charlatan who insists on replacing C and using Jira.
That is... not the biggest threat. It seems extremely unlikely that that is where linux kernel development would go.
If anything replaces C I'm skeptical that more than a shift in the character of the attack surface (rather than the area) will occur. That is, there will be just as much to attack, it's just that the attacks will require alternate methods. Some attacks will be harder to employ, some less, but on average probably about the same quantity and seriousness will occur.
You can call it a meme but that doesn't change my opinion.
> Some attacks will be harder to employ, some less, but on average probably about the same quantity and seriousness will occur.
This is unfounded, I don't really know how to respond to it other than that nothing has ever really worked this way in security. You don't get less secure in one area because you were more secure in another.
I agree with you, however that doesn't address what I actually wrote.
gcc: “hold my beer while I optimize out that block of code you clearly didn’t mean to be in your program”
Or rather, because the standard "imposes no requirements". A compiler that does the obvious/sensible thing for UB can also claim conformance.
From C99 3.4.3 paragraph 3 (emphasis mine): "Possible undefined behavior ranges from ignoring the situation completely with unpredictable results, to behaving during translation or program execution in a documented manner characteristic of the environment (with or without the issuance of a diagnostic message), to terminating a translation or execution (with the issuance of a diagnostic message)."
UB was introduced to allow skipping checks that otherwise might have added overhead, like array bound checks or null checks or to give more flexibility to how correct programs are optimized (e.g. freedom to evaluate arguments in any order). Therefore each compiler/library/os vendor could implement these cases differently.
However, it was never meant to allow compiler to assume "ok, I can see your code is totally broken, so let's break it more, cause you don't care anyway".
That's nevertheless how compiler writers are interpreting the standard right now. Take integer overflow for instance. Stuff like this:
unsigned u = some_computation();
int x = u; // because of reasons
x += 42; // hmm this may overflow, let's check that
if (x < 0) { // It's a 2's complement machine, this'll work
handle_overflow(x); // now, let's deal with the overflow
}
Now here's how your optimising compiler see that stuff: unsigned u = some_computation();
int x = u; // x is always positive.
x += 42; // signed integer never overflow
if (x < 0) { // x was positive, no overflow... always false
handle_overflow(x); // Let's delete this dead code.
}
Then, what actually happens: unsigned u = some_computation();
int x = u;
x += 42;
This is insane.[1] https://wandbox.org/permlink/dDDAJjOlm50GD1PW
It would be able to (correctly) make that assumption if `x` were a long, assuming 64-bit longs and 32-bit int/unsigneds.
Let's see, that's not an assumption you can make in C, or is it? I seem to remember that a compiler is allowed to present to the programmer at least one or two representations other than two's complement.
3.4.3 undefined behavior
behavior, upon use of a nonportable or erroneous program construct or of erroneous data, for which this International Standard imposes no requirements
NOTE Possible undefined behavior ranges from ignoring the situation completely with unpredictable results, to behaving during translation or program execution in a documented manner characteristic of the environment (with or without the issuance of a diagnostic message), to terminating a translation or execution (with the issuance of a diagnostic message).
-------------
Even if you want to argue over what "imposes no requirements" means, I'd argue that "ignoring the situation completely with unpredictable results" is very clear.
The note also says (and elsewhere in the standard it reinforces this) that since there are no requirements, implementations could absolutely do things like perform runtime checks, etc. Compilers do not do this though. This isn't for no reason, but that's a separate issue.
The galling thing is that they did this to win a benchmark war against Fortran.
I've heard some horror stories and apparently the development team is pretty heavily pressured to fix a lot of things.
Exactly, if you could get all the benefits without the danger why wouldn't you?
I hardly saw any reason to use it instead of Pascal dialects, besides being "the language" on UNIX systems.
Thankfully Bjarne created C++.
As long as computing remains a battleground where rich assholes put vastly different wrappers around the exact same shit so they can wring more money out of us, I will probably be writing in that abominable assembly language w/ turing-complete string-paster for many years to come.
I'd think it's quite difficult for a new language to replace a language whose main attractive points are it's longevity, stability (not the code and programs coming out of it, but the syntax and tools etc), and broad support.
There are many fields where C is easy to beat. But I don't think the next 40 year lasting language will be a C replacement.
This not realistic. No one serious/big will use it in production for the next 10 years easily. There is no certainty that it will be mainstream OS ever in the future.
You will see thousands of problems next year in production that was not seen in testing today. Linux has millions of users and millions of use cases and edge cases, so many drivers written specially for it, tested and proven to work. Stability and maturity is the key for OS today. No one will trade that for some unproven hobbyist OS project.
C is what it is partly because of its relationship with Unix, and also because it was the language that gave most straightforward access to "the machine" -- whatever that is.
And the hard truth is that The Machine was the only real universal platform around. But now we have this thing called The Web, which is not quite as universal but is getting there, and is much more like a single platform. JS has a special relationship with it.
The only thing that might upset this JS+Web applecart is WebAssembly. A "C of the Web" (call it W) would be a language high level enough to be pleasant to use, but which would have little or no runtime beyond what is native to WebAsm, and thus anyone could use libraries written in W.
People will use C or Python or whatever on the web for the same reasons so many people transpile to javascript today, including that they simply don't like javascript and would rather write code in a language they prefer.
But on the web, the basic "built-in" run-time is much richer than was available to those languages. It includes the DOM and many other things. JavaScript already exposes all those things, and no more, so I think it is ideally placed to become (remain) the dominant lingua franca.
So even if the WASM world is multi-lingual, expect APIs to be described in JS terms, maybe with some kind of backward-compatible type decoration thrown in.
Not the mention everywhere else it's been used - WebAssembly definitely isn't replacing Node.js, for instance, or game scripting.
WebAssembly is certainly able to replace JS in game scripting.
You're probably right, but egad, JavaScript is an even worse language than C. It's like jumping from the frying pan into the fire.
At least it's garbage-collected.
The thing I take away when reading this is that in 1987 there were superior languages to C which — while usable — were just a bit too sluggish, and hence lost out. I'm thinking specifically of Smalltalk & Lisp, but I'm sure that there are others (maybe ML was around that far back?).
Well, if a program which once ran in half an hour can now run in 300 milliseconds, I think that maybe we should reconsider using just a little of that extra performance to run a better language environment.
Imagine if we had OSes written in safe languages which catch errors rather than allowing programs to misbehave. Honestly, it's hard for me to do because I've become so accustomed to crashes. Recently I've been doing some X programming with Lisp, and it's amazing — even when I make a mistake, things keep on running properly. I just see an error and that's that.
Well, most of the time, anyway. Nothing's perfect. Still, the experience is orders of magnitude better than writing C. I mean that literally, not figuratively.
(Also, as an aside: when I switched tabs to the article, it displayed for just a second, then disappeared. Reader mode didn't even work. I had to enable JavaScript just to view some text. What gives‽)
Welcome to software development, where the code is shit, and the hardware improvements don't matter.
Hindsight is 20/20.
C was simply the pervasive language when the internet happened. C was designed when the world was a friendlier place, where buffer overflows did not exist. It was embraced by everyone because of its speed and simplicity.
No C programmer was raised/educated to have security in the front of his mind. Hell, many programmers of any language still don't.
Stop bashing C and start using Oauth instead of inventing your own authentication schemes.
"Many years later we asked our customers whether they wished us to provide an option to switch off these checks in the interests of efficiency on production runs. Unanimously, they urged us not to--they already knew how frequently subscript errors occur on production runs where failure to detect them could be disastrous. I note with fear and horror that even in 1980, language designers and users have not learned this lesson. In any respectable branch of engineering, failure to observe such elementary precautions would have long been against the law."
Fran Allen on "Coders at Work," about C.
"Oh, yeah. That would have been fine. And, in fact, you need to have something like that, something where experts can really fine-tune without big bottlenecks because those are key problems to solve. By 1960, we had a long list of amazing languages: Lisp, APL, Fortran, COBOL, Algol 60. These are higher-level than C. We have seriously regressed, since C developed. C has destroyed our ability to advance the state of the art in automatic optimization, automatic parallelization, automatic mapping of a high-level language to the machine. This is one of the reasons compilers are ... basically not taught much anymore in the colleges and universities."
Exactly. The same thing with TCP/IP. People complain about C and TCP/IP because security wasn't baked into these technologies. That's because that wasn't the big concern back then. And the reason C and TCP/IP have legs is because they are very good at what they do.
You get the same thing with people complaining about SQL, HTML and every other language/protocol/technology.
It's not about having security in the front of his/her mind. People make mistakes and always will. In C those mistakes very frequently turn into easily exploitable bugs. That is the security problem with C -- not that programmers somehow don't think about security enough[1].
[1] Of course this is also true, but it's not part of the argument against C.
I wonder if anyone's trying an explicitly adversarial approach to early CS education. Split your class into two teams, each team produces an implementation, each team scores points for finding vulnerabilities in the other team's implementation. Doesn't really matter what they're implementing as long as there's untrusted input involved.
Lately I've been doing some of the stuff I would normally do in C in GoLang instead. But it still can't match C in performance. (beats the hell out of Node.js though which is where I do most of my middle-tier / front-end code)
It also makes an excellent foot gun for the same reason.
People should not have to bend their mind to satisfy some tool like the borrow checker, the tool should do the right thing and allow them to express their ideas in code with as little effort as possible.
Erm, have you ever used Rust, and had much experience with C++? Rust's semantics are far clearer, simpler, and well defined than C++'s, and there are orders of magnitude less rules of thumb that you need to learn in order to avoid. If it looks complex, it's because it needs to be in order to adequately express the essential complexity of systems programming. Hopefully we will have languages that do a better and more elegant job in the future, but at the moment Rust is the best we have.
Example? In my experience, those saying this have either never used Rust, or never used C++. A strict compiler does not a complicated language make. To wit, if the response is "lifetimes", note that lifetime tracking is also critically important in C++ (not to mention C), it's just that C++ compilers give one far less help in doing so than Rust does.
But it's not just about a particular example, instead I get a general feeling of inaccessibility when reading Rust code. It's not enough to know C or C++, one has to understand the Rust way. This might seem unfair to complain about, but I didn't have any such issues when learning e.g. Swift or Objective-C, to give an example of non-GC languages.
src/first.rs:1:1: 4:2 error: illegal recursive enum type; wrap the inner value in a box to make it representable [E0072]
src/first.rs:1 pub enum List {
src/first.rs:2 Empty,
src/first.rs:3 Elem(T, List),
src/first.rs:4 }
error: aborting due to previous error
is now error[E0072]: recursive type `List` has infinite size
--> src/main.rs:1:1
|
1 | pub enum List {
| ^^^^^^^^^^^^^ recursive type has infinite size
2 | Empty,
3 | Elem(i32, List),
| ----- recursive without indirection
|
= help: insert indirection (e.g., a `Box`, `Rc`, or `&`) at some point to make `List` representable
Users have reported a lot less frustration after these changes, and we're always trying to make them better. For more on this effort https://blog.rust-lang.org/2016/08/10/Shape-of-errors-to-com...You haven't seen template metaprogramming, have you?
Most new languages are far safer than C. For some that's the main selling point, for others it's just a side-effect of having GC. I wonder if you could get a lot of traction with a language as expressive as Ruby, but with footguns galore. I'm not sure that space has been adequately explored.
Whereas with some other languages, if you ever need to get more down-and-dirty than the language is designed to allow, even for just a little bit of your program, well, too bad - you're not allowed to.
But could you write directly to a hardware address? I don't know, but I doubt it - Turbo Pascal ran on the PC, which didn't have memory-mapped hardware. You could do an outp. But in C, you could do
*(unsigned long*)0xFFFE0004 = 0x8000FF2C;
which was either absolutely necessary or absolutely stupid, depending on whether or not you had custom memory-mapped hardware at the specified address. Could you do that in Turbo Pascal?What's more, Turbo Pascal was PC-only, at least initially. If you had to work on a 68000-based embedded system running PDOS (not PC-DOS), and you had to use Pascal, it wasn't Turbo Pascal. It was just bog-standard Pascal. Inline assembler? No way. Pointer to a variable-sized array? Can't do it (you cannot give it a type). It was painful to work in that environment. And that environment was Pascal as specified by the language standard. Turbo Pascal fixed most of the problems, but it trampled all over the standard to do it. (At least it did have separate compilation.)
Why was it painful to work in that environment? First, we couldn't access our custom hardware without having to link to an assembly-language subroutine. In doing so, we lost type safety, which led to at least one hard-to-find crash. It also was just much more difficult and error-prone to write those routines. Second, we wanted to have a user-specified variable-sized array. We wound up having to create the largest array we could given the memory the machine had, and only using the part that the user specified, which was a pretty ugly kludge. Third, Pascal was just clumsier to use than C. It was more verbose and more finicky. (One part I remember in particular was the semicolons. You couldn't have a semicolon on the last statement in a block. As you added or removed statements, you kept having to fiddle with the semicolons, including on lines other than the ones you were changing.)
myMemoryAddr : Longint absolute $FFFE0004;
There were Turbo Pascal compatible compilers for Amiga, and it was the most common dialect, to the point ISO Extended Pascal compatibility was mostly ignored.There are still companies selling Turbo Pascal compatible compilers for embedded systems.
https://www.mikroe.com/mikropascal/
Inline Assemly is not part of C, it is a common compiler extension, whose semantics are not even portable across compilers.
Back in the day when Pascal compilers were more widespread, C was hardly much more portable, with each compiler having its own little world between K&R C and ANSI C89.
Old computers (and by definition, old computer software) never dies. The users do.
Logo is 50 :)
Sadly these AHL's are unrelated.
For C++ I remember spending years trying to carefully track compiler versions, precise build setup, etc. or your library would probably fail to link. I have never had a binary compatibility problem with a plain C library.
I somehow doubt git's popularity has a whole lot to do with it having been written in C.
What they are not seeing is the simple and beautiful part, which is a shame because I think that's why C is still around and still popular.
It's a little like looking at your spouse. You can fixate on the flaws and you'll hate your spouse. You can fixate on the positives and you have a much better chance at loving your spouse. Having a spouse that is artificially constrained to doing only the things you approve of is kind of what the Rust people are trying to do. Feels sort of weird to me but if that's their thing, go for it.
Or maybe that's a messed up analogy?