[1] http://animats.com/papers/languages/safearraysforc43.pdf
[1] http://animats.com/papers/languages/safearraysforc43.pdf
That's the second time in two days I've come to read offensive ranting about a language + "let's move to Rust" as the proposed silver bullet solution here on HN.
If this is the way the Rust community tries to evangelize their language I'm inclined to keep myself as far away as possible from it.
Now don't get me wrong: I think Rust is fine. But that "C/C++ is dumb and everyone who continues to use it is also dumb and/or ignorant" sentiment won't win you any friends. There are valid reasons why people keep using C++ and can't switch to any other language in the near future.
I wouldn't work on Rust if I didn't think it was important and a step forward, but even technologies with flaws (read: all of them, Rust has flaws too) are worth using, depending on your situation. Trade-offs are the heart of engineering.
I'm convinced that in order to make this work, we have to be more generous in how we perceive other people's comments.
https://news.ycombinator.com/item?id=10208786
So, there's an example where I had to step back, assume other person has intelligent reasons, and let the person show them with good results coming out of it. I still have a use of "C/C++" in these discussions but now do it less often due to superior safety over C of modern C++ development. New standards mean I'll probably be hit with a similar conversation again lol.
For decades, it has been known that C and C++ are bad languages due to their unsafety, and the use of C/C++ is probably the top cause of bugs and security holes in software; however, switching required giving up performance or features.
Now that finally a solution to this problem has become available after so much time, it's natural that people are enthusiastic about it.
> As usual, mainstream acts like everything else ever
> happened.
Not Rust. The reference has a long list of languages which have been influential (https://doc.rust-lang.org/reference.html#appendix:-influence...) and the official book has a list of research papers which have helped to guide the design of the language (http://doc.rust-lang.org/book/academic-research.html).Note: Strange that they draw no inspiration from Ada despite it countering many issues by design. I figure at least its tactics if not syntax would be a nice start on a new language. Least they borrowed from and built on a lot of good ones, though.
Also watching Julia as it has a host of good features with potential to take on Python and R at same time. Especially macros and painless C calls.
Note: Btw, good call on integral types. That's quite beneficial. Also, existential types would prevent all sorts of issues that happen when two things are same size but should be treated differently. Like a mi-to-km mismatched that was a downer for one program in particular. ;)
> Also, existential types would prevent all
> sorts of issues that happen when two things
> are same size but should be treated differently.
I don't know what existential types are, but we use phantom types for that: https://blog.mozilla.org/research/2014/06/23/static-checking..."I don't know what existential types are"
Unfortunately, I couldn't find a non-dense explanation of it that had relevant detail. Good news is the Ada book I've been referencing has a great explanation with examples. See Ch 2 "Safe Typing" for the answer in the first, two pages. They call them "Distinct Types" in there (shrugs). Included link to whole book in case you see safety tactics worth copying.
http://www.adacore.com/knowledge/technical-papers/safe-secur...
So, they make sure the types and implementation are separate so unique types ("existential types") can catch interface issues that can be subtle and nasty. Casts work around them where necessary while making the risk more obvious. Aim to try to keep overhead down compared to heavy OOP + regular types. Seeing it made me wonder, "Why don't the rest have this!?"
"but we use phantom types for that"
Thanks for the link! Adding it to my list of things to read and review. :)
As a consequence, it seems that if you use Ada's GC, you might as well use Java instead, and if you don't you might as well use C++ instead, which along with its unattractively verbose and not C-like syntax is probably why Ada is not very popular despite having been around for 30 years.
And, as I often say, the empirical studies all showed C doing far worse in terms of defects introduced. Ada benefiting safety & productivity over C is a proven fact rather than an opinion. So, questions are: (a) should we use it for this project? (b) should we build on a proven solution and make it better?, and especially (c) should we borrow anything in this that made it effective when designing a new language? I especially push for (c) to be applied any time new tech is built. Ada is my reference for safe, systems programming because it did what it promised and is the oldest one still updated/used. Also turned out almost as future-proof as COBOL with a lot safer, easier maintenance. Such longevity, along with safety, benefits enterprises writing mission-critical software.
Details on various safety mechanisms:
http://www.adacore.com/knowledge/technical-papers/safe-secur...
Far as replacements, Rust is looking promising. Needs to mature a bit and better tools for various situations. A great design, though, with a lot of potential. Not sure how it would fare for embedded, though: Ada was used in microcontrollers and stuff.
Nobody's saying that C/C++ or its users are dumb and/or ignorant. The cost of switching languages is high, and Rust is not very mature.
If you have to write 'C', learn 'C'. I'd say the same of C++, but by the time you learn it, it's become something different :)
OK, but the fact remains that undefined behavior due to bounds checks is orders of magnitude less frequent in, say, Ruby than in C. I have a hard time seeing this as anything other than a good result.
There are issues other than safety that influence these decisions. Rust & Lua are also closer to 'C' than is Ruby, so ....
And I'll add that most of the older, safe languages let you turn off safety for performance if you have to. You just then have to use you brain like you would with C or whatever. Just for that module, though, along with calls into it.
This was in 1987 or so. Perhaps that was unusual; I don't know. It was part of the culture where I worked. I'd programmed professionally in Pascal before, so I'm familiar wit the compare/contrast here. Tactically, it was safer, because 'C' covered more than did Pascal, and the balance of the code when you used pascal was assembly. Perhaps the assembly was sufficiently dangerous that we didn't make as many mistakes, so who knows? We felt assembler was more of a liability.
The choices were more limited then. Performance, shmerformance - darn near everything runs fast enough, now, with the exception of a few grotty corners.
Rust/Python/what have you are all very nice. I just sympathize with the poor people who get drafted into working on 'C' code bases without seeming to be happy about it. I personally would find somebody to mentor me a bit and review the code if I were in those shoes. Perhaps that sort of thing has passed; shame if it has.
It's anecdotical, but most of actual C++ programmers I know won't produce more defects than Java programmers for instance. In the last 5 years I have seen buffers overflows just once and was in a legacy code (and C++ has been my main language in that period). Java has its own sets of defects (such as resources leaking).
In my opinion are good practices like code reviews the ones that help to produce good code and not the choice of the language.
You're also mischaracterizing Animats' comments as "offensive ranting". That's John Nagle (this guy: https://en.wikipedia.org/wiki/Nagle's_algorithm) and he knows his way around C.
Friends are a social, not technical, matter. Far as C, using it outside of necessity (eg legacy or critical libraries) is dumb for projects aiming for robustness. It's not even opinion so much as empirically-backed fact: every study the military and researchers did on C vs other programmers back in the day showed the C people had highest defect rate and among lowest productivity [often due to debugging]. Many were an investigation into Ada vs other languages w/ Ada cutting defects in half not being uncommon. Later Java studies showed same problems with C and benefits by ditching it albeit with much higher overhead. ;) The CVE's continue to reflect the same with Java and Ada being interesting as many of their apps flaws occurred when they leveraged a C library.
Like Ada's style or not, it kept defects and overhead low by systematically [1] looking at where problems happened and implementing solutions to them at language/compiler level. Modula-2 did the same thing to a lesser degree. I believe PL/S, a systems variant of PL/1, also had design choices that boosted safety partly inherited from PL/1. Burrough's ALGOL (esp NEWP) did quite a bit in language to reduce risks. These were used in reliable software ranging from mainframe OS's to desktops to embedded. They also predate the C++ language which also attempted to solve C's issues albeit with great C compatibility.
So, there are numerous languages whose coders can build robust software at a similar or better pace to C coders building buggy software. These languages, and their descendants, have been around for a while with tool support. The apps built with them are more reliable & had fewer coding vulnerabilities. Clearly, C is just badly designed in terms of robustness, productivity, and maintenance. No secret why: its lineage dumped most of the features of better languages to run on hardware [2] that couldn't support them with acceptable efficiency. Over time, we ditched that hardware for a series of others better suited to our needs. It runs so fast that the CPU is often idle 70-90+% of the time our apps are running. We should likewise ditch C and raw efficiency for something better, safer, and still efficient. Rust is a promising option among others.
Note: One trick in the past where the GUI, a legacy library, etc absolutely needed C or C++ was to split the app. Ada, particularly, made cross-language design easier to support this use case. So, the critical part of the logic might be written in a safer langauge, less critical in C or C++, and a careful interface glue them together. This was done for Mondex Certificate Authority, IIRC. I did it myself in numerous designs for fault-isolation or security reasons.
[1] http://www.adacore.com/knowledge/technical-papers/safe-secur...
First, it's probably not useful to arbitrarily claim what the biggest flaw in the design of C is without context or data.
Second, I don't think it's reasonable in this case to place the blame on the design of a language (certainly, the choice of C has long been a pragmatic decision). As Linus alludes to directly, the tools exist with which to write more correct, idiomatic C code for the kernel (e.g., ARRAY_SIZE and a useful gcc warning).
Finally, as the historical record should demonstrate, 'getting rid of X by replacing it with Y' is not particularly actionable; it certainly isn't the most efficient solution in terms of resources and likely not in terms of correctness over the short- to medium-term.
If I'm to hazard a guess, the point of his message was to remind everyone the biggest source of flawed code lies in our own hands and in the biases individuals and groups bring to large-scale development.
The choice, here, isn't between using a good language and a bad language; it's between using a language and its idioms correctly, and using those things lazily with foul consequences.
"Move to Rust!" is not a productive or helpful suggestion, unless you have an massive army of programmers ready to port Linux to Rust.
EDIT: Downvotes eh? Then I suppose I'll state outright what I left implicit: Linux is not inextricably tied to any single language, and it's a strawman to suppose that any attempt to integrate Rust with Linux would first have to reinvent the universe from scratch. Rust has excellent interoperability with C (though, notably, unions are iffy), and it's common to write Rust code that gets called from C with C being none the wiser.
>Linux is not inextricably tied to any single language
It is pretty tightly coupled to GCC, which, as you know, is a C compiler. Some people are working on supporting Clang, but last I heard they haven't succeeded.
It would certainly be interesting to see someone link Linux to Rust code though.
The LLVM build patchset is being actively maintained, here hare slides from February: https://events.linuxfoundation.org/sites/events/files/slides...
That's good though. Compiler freedom is better for everyone.
Even in the 1970s, this was kind of lame. Pascal could pass array sizes through a call.
Other systems programming languages older than C did it properly.
Was that "actually better" due to C, or to Unix? Well, maybe some of both, but Unix was written in C, so C was useful for writing the operating system that you blame for C taking over the world...
If you don't understand why C was better in practice, you're going to fail in attempts to create something better, and never know why. You're just going to be ignored, while you whine about how your way was better.
Back in those days we used to pay for our compilers, usually in values of thousands, remember those?
Only languages delivered as part of was then the OS vendor's tooling got used at most work locations.
It always required lots of persuasion to buy compilers for languages not delivered as a standard part of the OS.
So as UNIX managed to gain a foothold into the enterprise and universities, pushing mainframes and other OSes, C gained mind-share.
UNIX took over the world by accident, by having AT&T initially providing the code for free (which they repented later on) to universities, which happened to have persons like Bill Joy and Scott McNealy that got successful with their startups using the OS they enjoyed at their universities.
Had AT&T never given the code for free or those startups floundered, UNIX will be another footnote alongside MULTICS and friends.
Just like nowadays no one sane would use JavaScript, if the browsers had first class support for other programming languages.
Those OSes written in other system programming languages: Why didn't their computers take over the world? I mean, sure, AT&T was pumping money into Unix, but other companies were pumping money into the competition. Why did Unix win? It's not just because of evil AT&T. It's because Unix could deliver working features, and others struggled to keep up.
Why could Unix deliver working features? Yes, partly because of AT&T's money. But partly because C turned out to be really useful as a systems programming language.
These other languages you mentioned a few posts ago: Sure, they looked good on paper, only where's the beef? When it came time to deliver, what was written in them? More functionality was written in C, which is why Unix won.
I'll repeat my previous statement: If you don't understand why C was better in practice, you're going to fail in attempts to create something better, and never know why. You're just going to be ignored, while you whine about how your way was better.
History ignored your "better" languages. The world moved on from them, for good reasons. You think they were better, but in real life, they weren't.
I think the situation is not far off from what pjmlp is saying: there were several more or less equally good languages available at the time, and C won by virtue of being in the right place at the right time. There are lots of really nice things about C from a systems point of view, but most languages of the 70s had those too. C was a nice language for the late 70s (not so much today, IMO), but not hugely nicer than its contemporaries.
But remember, the original claim by pjmlp was "It is a big flaw to ignore what other language communities were doing. Other systems programming languages older than C did it properly." That doesn't apply to the Pascal you speak of, because at the time C arose, Pascal wasn't "once you get to a recent enough version supporting dynamic arrays/first-class pointers". It was Pascal with the size of an array being part of the type of an array. That's type safety, sure. But it also means (to use an example that I have personal experience with), if you're writing a numerical simulation on a 2D array, and you want to let the user specify how big the mesh is, and then allocate memory to hold the mesh, you have a problem. You can't define a variable-sized array at all in the Pascal of the time.
Now try to think about how you'd write a memory allocator in that. Good luck.
Another example: We were on an embedded system. To write to hardware registers, we had to call an assembly-language subroutine. In C, we would have simply used a pointer to an absolute memory address.
This is what C gave you - you could just do things without the language getting in the way. Yes, you could cut yourself on C's sharp edges, whereas Pascal protected you. But when you needed the sharp edges to cut something, Pascal didn't have them, and you were stuck.
Again, I don't know enough about Algol to meaningfully compare it to C. The Pascal of the day wasn't the answer, though.
Pascal was originally used for teaching it doesn't count. The first time it was used for writing an OS, was with the Object Pascal dialect for the Mac OS.
What counts were Algol 60, Algol 68, Algol W, PL/I, PL/M, Mesa among many others.
As for the statement of only C being usable without Assembly, check the B5000 from the Burroughs Corporation, developed in 1961.
EDIT: where => were
I'm not assuming it. The environment I was in had memory-mapped IO.
> Many machines used only IO ports. Where are the C features for that?
C compilers for machines that only had IO ports typically had a library function to do it. At least, by the era of Turbo C, they did. (The x86 architecture is the first one I was familiar with that did IO that way, and Turbo C was the first C compiler I had on it. I presume that, if earlier architectures did IO that way, C compilers for those architectures had similar capabilities, but I do not know that first-hand.
> Pascal was originally used for teaching it doesn't count.
I was replying to pcwalton, who asked what you can do in C that you can't do in Pascal. So it may not count to you (four comments upthread from here), but it counted to the comment I was replying to (two comments upthread from here). Perhaps your comments would be better addressed to him/her.
> As for the statement of only C being usable without Assembly, check the B5000 from the Burroughs Corporation, developed in 1961.
I didn't say that only C was usable without assembly. I said that using Pascal, in order to write to memory-mapped IO registers, you had to call a function written in assembly, and that in C you didn't have to do that. I never said only C was usable that way. Please stop putting words in my mouth that I didn't say.
In what language do you think the Turbo C library functions for port IO and all those handy BIOS and MS-DOS calls were written on?
Assembly, of course.
As for doing memory mapped IO with Pascal, you could something like this in Turbo Pascal. Other dialects had similar extensions.
var
videoMem : array [0..255] of byte absolute $A0000;
begin
videoMem [0] := 12;
end;Of course. But they still looked like just another function call in C, and they came with the compiler, so the programmer writing in C didn't care.
In my Pascal example, there was no such function, so we had to write it, so we had to care.
You seem to be consistently trying to make my words say things that I am not saying, and then arguing against positions that are only in your own mind. It's getting quite tedious.
"As for the statement of only C being usable without Assembly, check the B5000 from the Burroughs Corporation, developed in 1961."
Good call. First great, overall system as I see it. Might also remind them that Wirth's and Jurg's Lilith workstation ran on Modula-2 and assembly. Most comparable to UNIX in terms of hardware and personnel constraints. Shows even constrained teams could do better than C or UNIX. The Oberon systems used Oberon and assembly, as well, with a GC'd OS and software. That helps in another recurring debate about OS's in "managed code." ;)
So, yes, that C is necessary for systems programming is a long-running myth refuted by examples which were better because they didn't use it.
Note: Nicklaus Wirth's solution to same problem was much better: P-code. He made idealized assembler that anyone could port to any machine. His compiler and standard library targeted it. Kept all design advantages of Pascal with even more implementation simplicity than C. Got ported to something like 70 architectures/machines.
Now, for OS's. Let's start with Burroughs MCP. The Burroughs OS was written in a high-level language (ALGOL variant), supported interface checks for all function calls, bounds-checked arrays, protected the stack, had code vs data checking, used virtual memory, and so on. That's awesome and might have given hackers a fight!
Later on, MULTICS tried to make a computer as reliable as a utility with a microkernel, implementation in PL/0 to reduce language-related defects, a reverse stack to prevent overflows, no support for null-terminated strings (C's favorite), and more. It was indeed very reliable, easy to use, and seemed easy to maintain. You'd have to ask a Multician to be sure.
So, the OS's were comprehensible, used languages that made reliability/security easier, had interface/array/stack protections of various sorts, consistent design, and all kinds of features. Problem? Mainframes were expensive. The minicomputers Thompson and Ritchie had were affordable but their proprietary OS's were along lines of DOS. You can't do great language or OS architecture on a PDP-11 because it's barely a computer. It would still be useful, they thought, if it had just enough of a real language and OS to do useful work.
So, they designed a language and OS where simplicity dominated everything. They took out almost all the features that improved safety, security, and maintenance along with using a monolithic style for kernel. Even the inefficient way UNIX apps share data was influenced by hardware constraints. The hardware limitations are also why users had to look for executables in /bin or /sbin for decades: original machine ran out of space on one HD & so they mounted another for rest of executables. All that crap is still there because fixing it might break apps & require fixing them. Curious, did you think they were clever design decisions rather than "we can't do something better without running out of memory or buying a real computer so let's just (insert long-term, design, problem here)?"
The overall philosophy is described in Gabriel's Worse is Better essay:
https://www.dreamsongs.com/RiseOfWorseIsBetter.html
As Gabriel noted, UNIX's simplicity, source availability, and ability to run on cheap hardware made it spread like a virus. At some point, network effects took off where there's so many people and software using it that sheer momentum just adds to that. Proprietary UNIX's, GNU, and Linus added more momentum. After much turd polishing, it's gotten pretty usable and reliable in practice while getting into all sorts of things. One look underneath it shows what it really is, though, with not much hope of it getting better in any fundamental way:
https://queue.acm.org/detail.cfm?id=2349257
So, aside from not knowing history, there seems like there's not even a reason to debate the reason behind bad design in C and UNIX at this point aside from merits of overall UNIX architecture vs others. The weaknesses of C and UNIX were deliberately inserted into those by the authors to work around hardware limitations of their PDP-11. As those limitations disappeared, these weaknesses stayed in the system because FOSS typically won't fix apps to eliminate OS crud any quicker than IBM or Microsoft will. Countless productivity, money, and peace of mind were lost over the decades to these bad design decisions in the form of crashes or hacks.
Using a UNIX is fine if you've determined it's the best in cost-benefit analysis but let's be honest where the costs are and why they're there. For the Why, it started on a hunk of garbage. That's it. Over time, when it could be fixed, developers were just too lazy to fix it plus the apps depending on such bad decisions. They still are. So, band-aids everywhere it is! :)
That is a flaw.
In fairness, C++ containers know their size. And Rust's solution to buffer overflows is the same as every other language — run time bounds checking.
Can you show us the benchmarks for that?
As for actual benchmarks, I think they would actually be quite easy to produce. There's exactly one line in the Rust stdlib that provides bounds checking for indexing, right here: https://github.com/rust-lang/rust/blob/master/src/libcore/sl... . All you would need to do is take out that line and compile Servo with your modified stdlib and compare the results of running the built-in benchmarks. I may just do this myself as a blog post. :)
For the record, there seem to be a lot of bounds checking independent of indexing:
https://github.com/rust-lang/rust/blob/master/src/libcollect...
https://github.com/rust-lang/rust/blob/master/src/libcollect...
https://github.com/rust-lang/rust/blob/master/src/libcollect...
https://github.com/rust-lang/rust/blob/master/src/libcollect...
Right, because that's the only way to do it? What's wrong with that?
The point is a language like Rust has built-in bounds checking...
As does C++. See e.g. vector.at and libstdc++ debug mode:
https://gcc.gnu.org/onlinedocs/libstdc++/manual/debug_mode.h...
The difference is only opt-in vs. opt-out. To its credit, Rust provides additional protections against iterator invalidation and dangling pointers, but its approach to buffer overflows is the same.
Also, in C++ containers, ".at()" is usually checked, but "[]" is not. So C++ code still regularly has buffer overflow problems.
LLVM does optimize these out, and in more cases than Go does.
> C++ can't do that because the compiler doesn't know that a template-implemented bounds check is a bounds check.
It doesn't need to. SCCP is a very general optimization. GCC and LLVM optimize out bounds checks in more cases than Go can do.
Maybe the fact that it's wrong? Compile-time bounds checking is such a classic "hello world" example for type checkers that it's become a cliche (eg. https://www.quora.com/What-are-some-good-examples-of-practic... )
Dependent types make zero-overhead bound checking trivial. More recently, less powerful type systems like liquid types have been proposed (because reasons). Zero-overhead bound checking is usually on their feature list, since a) it's ubiquitous b) it's easy.
But until then, Rust's approach is runtime bounds checking.
http://research.microsoft.com/en-us/um/people/marron/selectp...
So, out of curiosity, what's the odds of something like this proposal getting adopted in a standard, GCC, or anything else?
Maybe if Amit Yoran was still at Homeland Security...
It's an axiomatic choice inherent to the language. You add a parameter representing the extent of the array, and respect that.
I have no heartburn with Rust, but there's sure a lot of legacy 'C' out there.