Some Were Meant for C – The Endurance of an Unmanageable Language (2017) [pdf]
humprog.org
humprog.org
Also several statements about how C must be used because somehow it's closer to the real world/hardware than other languages. Which is easily shown to be false given that hardware designers have had bend backwards and into contorted shapes to emulate the hardware environment that C was originally created to work against. This great article is a nice rebuttal of that: C Is Not a Low-level Language: Your computer is not a fast PDP-11. https://queue.acm.org/detail.cfm?id=3212479
These types of arguments feel like they come from people who don't realize how much the compiler reworks your code to make it act like it does what you told it to do.
If there's ever a radical new ISA which supplants x86 where C's constructs prove to be a bottleneck, and that ottleneck only applies to C, you will finally see it fall to the wayside.
>These types of arguments feel like they come from people who don't realize how much the compiler reworks your code to make it act like it does what you told it to do.
I guess it's hard to argue against a vague undefined term like "rework" but i can assure you that the output of gcc largely resembles what i expect it to. If the compiler does something bizarre i will know about it because i actually do check the assembly.
Which in the very initial paragraphs points out it is not about having a hidden set of registers to work with - it's about exploiting parallelism and the tension between a sequential programing language and massively parallel hardware architecture.
The ISA is not the hardware it's a 40 decade old abstraction that the hardware bends itself into to make your code run.
Precisely! And because the weirdness happens below the ISA, all the mismatches Dave Chisnall describes apply to any other language just as much including actual machine language. So raw machine language is not a low-level language? <head scratch>.
It'd be interesting to see what interface we'd choose today with multiple decades of hardware and software development.
I have a sneaking suspicion that a more dataflow-oriented interface might be useful. The CPUs do dataflow analysis, the compiler does dataflow analysis, and higher levels are also often dataflow, but all communicate using a non-dataflow instruction stream. On the other hand, the only commercial dataflow CPU I am aware of, the NEC 7281 (http://www.merlintec.com/download/nec7281.pdf) was less than successful.
A worse one? Not a couple days ago there were a couple stories about Itanium here on HN (a designed-from-scratch ISA that turned out to be practically worse than the much more ad-hoc x86_64). Even e.g. RISC-V (which is not entirely free of legacy) is practically 1:1 with ancient ISAs like ARM.
We'd be in a damn sight better place if it was, though.
It may be a valid point, but value-less, since this basically implies that most hardware you are going to be able to buy today (ARM, x86, whatever) is going to be a fast PDP-11 (or at least it's going to present itself as one).
So it may be that your "clever" C algorithm which in your head translates into just six CPU operations, unfortunately on a real modern CPU is six hefty macro-ops that will take dozens of cycles to execute and repeatedly go to sleep waiting for main memory, whereas the algorithm in a modern language that looked ludicrous to your C programmer eyes compiles to sixteen tiny ops the CPU can consume two at a time with no waits and it's done in eight cycles while the C code was still waiting for a main memory read.
I'm not disagreeing with that. It's a valid remark. What I'm arguing is that it's a useless remark in practice.
> So it may be that your "clever" C algorithm which in your head translates into just six CPU operations, unfortunately on a real modern CPU is six hefty macro-ops that will take dozens of cycles to execute
Note that this was already true even for the original PDP-11, and more so for the original 8086, both being ridiculously CISC and microprogrammed. C was born in this context.
> whereas the algorithm in a modern language that looked ludicrous to your C programmer eyes compiles to sixteen tiny ops the CPU
Even if true, it's irrelevant, since even a "modern language" will be forced to target that machine with its PDP-11 -emulating ISA, and is therefore subject to the same limitations than "modern C" is. The very point the article is making leads to the conclusion that most machines today are actually presenting themselves as a PDP-11-like, and therefore any language which does not target a PDP-11-like is doomed to a niche.
For example have you noticed your CPU is multi-core? The PDP-11 wasn't multi-core, and C doesn't really do well for writing concurrent software. Practically though, despite Amdahl's law it makes sense to provide a multi-core CPU. Most C software won't be able to take advantage, but some of your software is written by people who either will put the extra effort in despite C or use a different language where it wasn't so hard.
Take another example, the 0-terminated string. This terrible C design is more or less omnipresent. And sure enough CPUs have features to enable this terrible data structure, because it's so omnipresent. But those features don't make a better string worse they just pull the C string closer to parity than it might be otherwise. Rust's str::len() is still much faster than C's strlen() even if your CPU vendor focuses entirely on the C market.
The other two examples are highly questionable. 0-terminated strings are not an artifact of the computer architecture, but a conscious trade-off, and in fact other languages already used Pascal strings (even on the PDP-11). Same as if you want your lists with a constant time size function or not. Sure your size() is now "much faster", but something else now becomes slower. In fact, you could argue x86 bends to supports _both_ string implementations, since the rep prefix does support (and update) a count.
And the final example, "isdigit" is arguably not even a language issue, and definitely not a legacy of the PDP-11 nor x86 architecture. I'm rather sure no one in his right mind would even consider wasting 0.25KiB just for such function, not until well into the 90s at least. These types of code-bloat optimizations are more of the 2000s and even 2010s "performance" craze. You yourself even say on the other thread that glibc is doing this to facilitate locales, and not for performance reasons.
There are a bunch of other things computers could do that would in principle go real fast but humans can't successfully reason about them. It seems like a stretch to blame C. I reckon that even the most precocious humans have experience with causality a long time before they write any C and it's this experience which makes it hard to handle a world which cannot be understood in terms of cause and effect for example.
Can you give an example of that?
Rust's u8::is_ascii_hexdigit is a predicate which decides whether the 8-bit unsigned integer is the ASCII code for a hexadecimal digit, that is 0 through 9, A through F or a through f.
Inside the implementation is exactly what a modern programmer would expect, a pattern match, it'll do a whole bunch of arithmetic operations on the 8-bit unsigned integer and get a true or false result. Your modern CPU can do that real fast, because it's all just arithmetic on registers.
If you go look at a C library isxdigit() function, sometimes (not so often these days) you will find it has a mask and a LUT, it's assuming that "just" looking up the answer in memory will be faster. Your modern CPU can't do that quickly. It must calculate the index into the LUT, and then read the appropriate memory, hopefully that's in L1 cache or something, if it needs to go to main memory that takes an eternity, and either way it means something else can't be in the cache because this LUT is taking up space.
Are you sure you're not just used to shitty code?
Its excuse is that this simplifies locale handling, you know for if your locale has hexadecimal digits in a different place from everybody else, presumably because you're from some alternate universe where Z is a hex digit.
Other C libraries don't have this problem.
http://git.musl-libc.org/cgit/musl/plain/src/ctype/isxdigit....
C is indeed a low level language.
If you wanted a language that's close to the actual hardware you have, you'd have a different language for different hardware. (That's what assembly is)
Meanwhile, C is simple and easy to write compilers for. Therefore, it gained popularity because of its portability
Please don’t start a C flame war either HN. I know I’m nerd sniping you all on this one
(But, I also totally concede that a language is inseparable from its std lib in reality)
In my experience it's only the language lawyers who actually know what to look out for, and there are damned few of them.
Which is stupid since undefined and unspecified behaviors are hugely important to how C and C++ are used in the field.
Is it really that fundamental if you can setup your project to throw warnings/errors, and even use static code analysis tools to flag those?
I'd argue that it's far worse to work on a project that's not setup to detect those errors than to expect that sort of issue to be handled at the recruiting level.
Why are they though? Should they be?
It is a feature introduced by compiler writers who prize esoteric optimisations more than simplicity.
And it is not forbidden by the standard, though discouraged.
Use-after-free or double-free is undefined behavior.
Therefore a C implementation that provides malloc(3) should provide an implementation of free(3) that is a no-op and provide a garbage collector that actually frees the memory.
It's not an argument, it's an observation and statement of fact.
> Use-after-free or double-free is undefined behavior.
Yes.
> Therefore
Why therefore?
> a C implementation that provides malloc(3) should provide an implementation of free(3) that is a no-op and provide a garbage collector that actually frees the memory.
If you think so. It certainly doesn't follow from what I wrote. My point is that if you use a pointer after you freed the memory, you get whatever is in memory that the pointer points at on the machine the code is running on.
That is not a "weird and bizarre" implication, it is a plain and straightforward implication given how C is implemented on real machines. It is (obviously, I hope!) outside of the scope of the C standard to say what will be at that memory location at that point, and therefore it is UB.
And it being UB, "what you get" may also, for example, be a segfault because the library/OS have decided to unmap that memory. And again, that is also not "weird and bizarre", but perfectly straightforward.
¯\_(ツ)_/¯
It's an error even if it's undetectable by the compiler. I understand that nomenclature from few decades ago makes a distinction but nowadays we just call such thing an error.
How so? It is behavior left undefined in the C standard.
> It is a feature introduced by compiler writers who prize esoteric optimisations more than simplicity.
It's not the compiler writers who left the behaviour undefined, but the standard writers. If it was so clear what should be done (so all compiler writers should do it in one way) then why did they not define it?
> And it is not forbidden by the standard, though discouraged.
What do you mean it's not forbidden in the standard, it's undefined in the standard hence the name. Because of this you really should not rely on it, because the behavior could change.
C _looks_ simple, deceivingly so. As soon as you need to deal with a large project in a corporate environment on a code base aged in double digit years you don't think it's simple anymore. C is incredibly complex, unmanageably so.
But, I’m not arguing with you. I honestly think that any corporate codebase after double digit years (and seven figure LoC) turns into a completely unmanageable mess. If it isn’t, it’s because of the team culture and pure rigor. Not the language
I don't limit my statement to just C. The topic was C however so I mentioned C.
I’ve worked in a lot of massive and unmaintainable codebases, and I’ve always had the most fun untangling C ones. I don’t know exactly why, but I suspect it’s because at their core they are always plain ol’ simple C.
Your comment reads like a red herring though. You're not actually trying to argue that C makes thing easier or hard. Your red herring only means that hard problems affecting all languages aren't magically turned into non-issues in C. You can still write a C program that uses threading and is as safe as any other language and end up being a far simpler project, but that means nothing because the subject is still hard. And what would that actually say about C, let alone refute?
Writing C is simple.
Writing multithreaded programs is difficult and C doesn't help you with it.
Let's not argue about religions online. The mullahs will arrest us for improper use of void.
In a corporate environment, just use Java or C# or whatever. It's a corporate environment, probably not that interesting. C is not necessary.
Even modern drivers for Apple, Google and Microsoft platforms are written in C++.
Mind expanding on that?
I've done this before, never felt like C wasn't simple.
I think your main argument here is that large C code bases can easily become unwieldy, and I agree with you on that. However, simple multithreading in C is simple. It's not a terribly difficult problem to solve.
Sharpened rock is a simple tool. Versatile too, but try using it to make a sculpture and it's going to be hard, very hard. Especially if you never used it extensively.
C is a simple language. No doubt about it. But doing multi-threading and memory safety in C is like trying to build a 100m tall structure using nothing but a sharpened rock. Possible but, why would you.
Java is a complex language. Lots of complexity. Classes, annotations, async, System.out and so on. But multi-threaded and memory safe (memory leaks allowed ofc) programs is as simple as just not using Java's array of unsafe tooling.
> But multi-threaded and memory safe (memory leaks allowed ofc) programs is as simple as just not using Java's array of unsafe tooling.
as _easy_ as just not...
"Complex" is the opposite of "simple". "Easy" is the opposite of "hard".
I could port most of the multiprocessing based python scripts I have lying around to C without a problem. Multithreading can be easy when you have a few rules in place what threads can read/modify data at a specific point in time.
C has near universal portability. It's everywhere. That's the most important property.
I first learned C when I was 11 years old; admittedly I learned BASIC and Z80 assembly before then, and I'm an outlier.
That’s what I mean by simple. And that’s the reason why it also has universal portability. Cause it’s simple
I have a project still at the planning stage. I have many Rust crates lined up. So far I really like the bits of Rust I have learned. C plus generics? Sign me up! But damn, after coming from really awesome IDEs like Visual Studio, it just seems like it's taking me forever to make progress.
Right now I'm asking myself, does Rust save me time in the long run from chasing down bugs? I'm not really sure, because I think I'm already decent at avoiding C foot guns.
In general, C is a better option if you find you are escaping to macro defined Assembly or other memory ops (glib back-ported data structures for C highly recommended). Rust likes to keep unsafe operations organized, but is a persistent pain if a problem scope requires many "unsafe" operations.
I'd wager Rust will end up like Boost... really cool.. but too chaotic to trust past 10 months. =)
https://en.wikipedia.org/wiki/Second-system_effect
i.e. the attempt to "improve" ultimately adds additional complexity, and defeats its stated purpose due to conflicting use-cases.
Rust is replacing C++ in several well documented places, however C++ itself was a replacement for C in many places. So by definition the use-case for Rust overlaps the use-case for C.
Arguments like this sound like old arguments that were re-told to me by my father where people argued that C couldn't replace the use case for Algol at the University of Michigan MTS system. (Assuming I'm remembering the second-hand argument right. I'm a second generation software engineer.)
Also, there's nothing of note in the last 10 months about Rust that's different from the 10 months before that.
Actually, Cpp originally was a C template pre-compiler if I recall... just like Rust was. Note, people argue for years about various language merits.
"there's nothing of note in the last 10 months"
Indeed, still a pain to build, port, and package on some systems due to the poor project design required dependencies. There are papers pointing out the common fallacy of increased memory safety with the current Rust builds as well.
Of course, Rust may work fine for whatever you are building. =)
So when he started at Bell Labs and all it had was C, he quickly got busy getting his high level tooling back.
Being a C preprocessor was the quickest way to achieve that goal.
The problem with those comments is that they are intellectually lazy, and use age as a proxy for "better". Language X can be designed with specific features in mind that offer interesting features that Language Y fails to offer, but that is not ensured by the release date, nor might that be an acceptable tradeoff when considering other factors.
Pointers and pointer math were a 10 billion dollar mistake.
Unchecked array access is a several billion dollar mistake.
Goto has its place: consolidated resource error unwind cleanup. That's basically its only valid use except for mechanically-generated finite state machines. Beyond that, don't bother. Other programming languages use reference counting and lifetimes to manage resources.
int
foo() {
void *b0 = malloc(1000);
if (!b0) goto err0;
/* do something */
void *b1 = malloc(1000);
if (!b1) goto err1;
/* do something else */
void *b2 = malloc(1000);
if (!b2) goto err1;
/* keep going */
/* ... */
free(b2);
free(b1);
free(b0);
return 0;
err3:
free(b2);
err2:
free(b1);
err1:
free(b0);
err0:
return 1;
}I don't agree. Null is just a tombstone value. If you're dealing with specifying memory addresses, you either specify the address or point a value that tells you the variable is not pointing to an address.
If anything, conflating arrays with pointers was a bad idea. The language would be far better equiped to deal with memory safety issues if an array data type included info like total size, used size, and perhaps also stride length. Relying on tombstone values to delimit strings or arrays was just prone to blow up on everyone's face.
I for one really hate destructors. They mean you can never tell what `delete` does. I remember a nasty double-free bug I had because of that.
> They mean you can never tell what `delete` does.
Can you elaborate on this? A destructor is a function that's built into an object. It's really not more complicated than that. If malloc/free are functions that should be used in pairs, then the constructor/destructor pair tries to provide a convenient and structured way to know where malloc/free go. You would call new in the constructor and delete in the destructor.
I'm curious to know what you prefer?
Somewhere down the line these other objects are freed again, causing double free.
It therefore still calls the destructor.
So this has literally no bearing on the problem. It doesn't help at all.
“It can scarcely be denied that the supreme goal of all theory is to make the irreducible basic elements as simple and as few as possible without having to surrender the adequate representation of a single datum of experience.” ( Albert Einstein )
There are lots of scenarios where C++ is the only available alternative to C, with similar tooling level and libraries.
Had it not been for UNIX freebie, it would had been a footnote like so many others.
Outside UNIX clones and the few surviving embedded workloads without alternative, there is hardly a reason to reach for it instead of C++, which is also a UNIX child.