Wrong. Counter-example: locales [1]
Also: String handling, which is responsible for so, so, so many security vulnerabilities. The underlying cause is zero-terminated strings (instead of using start/end pairs or start/length pairs). You can't even tokenize a string without either copying or modifying it!
Also also: just a single, apologetic mention of undefined behavior. Responsible – in cooperation with over-zealous compiler writers – for so many more bugs not already caused by improper string handling.
[1] https://github.com/mpv-player/mpv/commit/1e70e82baa9193f6f02...
I don't think it has really been proposed anywhere though, unfortunately.
In a localized world searching for delimiters also starts to make less sense. Eg. I have been told it doesn't make sense to break on whitespace for Chinese text.
To be clear, I am not saying these functions are bad or evil or to blame for their limitations (a bunch of the problem space wasn't invented yet when they were introduced), just noting they have limits. A bunch of more recent languages and libraries have the same or similar issues, too.
UTF-8 continuation characters are limited to the range \200 through \300 so there's basically zero chance that if you choose something like comma as your delimiter that it's going to tokenize the middle of a multibyte sequence.
Also take into consideration that, under the hood, functions like strpbrk() are typically accelerated by CPU instructions such as PCMPISTRI which doesn't support UTF-8 natively but it does support UCS-2.
Not just "basically;" there is no possible collision between ASCII characters and any valid multibyte encoding. This can be seen somewhat visually in this table[1] and is an intentional aspect of the UTF-8 design.
On joiners / combining characters: I'd encourage using composed normalization (NFC) rather than decomposed normalization (NFD).
Just curiosity: are there any glyphs that lack a single codepoint representation, where one of the joined codepoints is an ASCII character? (That only helps after normalization, of course.)
The ISO C library string handling stuff is for systems programming, not for scanners and parsers for natural written language.
> I have been told it doesn't make sense to break on whitespace for Chinese text.
It could make sense to break on whitespace in some programming language or data format that allows Chinese (and other) identifiers.
A command interpreter that allows Chinese arguments (such as file names) wants to break on spaces, as usual.
> if a delimiter appears as part of a multi byte sequence, you may see strange results.
UTF-8 was designed by a dyed-in-the-wool C-and-Unix engineer, who ensured that such a thing can't happen. No character in the 0x00-0x7F range can occur in a multi-byte character.
The byte which starts a UTF-8 character cannot occur anywhere other than at the start, which is why we can use strstr to look for it.
I quit using it myself because I could never remember just what the exact protocol was for 0.
But one of the best features of C is that you can mostly ignore the standard library and still enjoy "C the language", e.g. nobody ever choose C for its standard library ;)
(also re UB etc...: use the mighty trio UBSAN, ASAN and TSAN!)
What is the standard library for the C which ends all the other standards libs?
Same for string handling/processing, there are specialized libraries out there which do a specific job better than the rather generic standard library functions (this is also true for the C++ stdlib), one just has to find those libraries.
PS: I don't think that a "batteries included" standard library (like python has) would even make sense for C, such a library would need to be opinionated by definition, better to keep (too many) opinions out of the standard, as this road just leads to another C++ ;)
I actually can't think of a lot of these that are super common across projects.
ICU is probably the most notable one, with its particular focus on unicode correctness.
A lot of large projects end up writing their own string library of sorts.
There isn't a universally recognized "standard library replacement" (and people argue about what parts of the standard library they don't like), but here:
https://stackoverflow.com/questions/486383/safer-alternative...
are a few options:
* Glib (the basis for Gtk): https://developer.gnome.org/glib/stable/glib.html
* The Apache Portable Runtime (APR): http://apr.apache.org/
Caveat: I haven't used them.
C is one of the hardest languages to do this with given its anaemic dependency management.
> (also re UB etc...: use the mighty trio UBSAN, ASAN and TSAN!)
All of which will miss some cases, even in combination.
Because C is so bare-bones, it usually leans on POSIX as its extended standard library, and that is also full of old cruft.
https://en.cppreference.com/w/c/string/byte/strtok
About getenv: This sort-of fixed with `getenv_s()` in C11
> Microsoft Visual Studio implements an early version of the APIs. However, the implementation is incomplete and conforms neither to C11 nor to the original TR 24731-1. For example, it doesn't provide the set_constraint_handler_s function but instead defines a _invalid_parameter_handler _set_invalid_parameter_handler(_invalid_parameter_handler) function with similar behavior but a slightly different and incompatible signature. It also doesn't define the abort_handler_s and ignore_handler_s functions, the memset_s function (which isn't part of the TR), or the RSIZE_MAX macro.The Microsoft implementation also doesn't treat overlapping source and destination sequences as runtime-constraint violations and instead has undefined behavior in such cases.
[0]: http://www.open-std.org/jtc1/sc22/wg14/www/docs/n1967.htm
When I learned C (early 90's), multi-threading wasn't even something you considered. It was either multiple processes or an event loop with select/poll.
[1] https://github.com/coreutils/coreutils/blob/master/src/yes.c
What do people that need higher resolution use? And the people that don't care about that amount, do they pay the performance penalty?
Every physical or humane value is hard to some extent.
For reference, the universe is about 13.787 billion years old [0]. That's about 13.787 * 10^9 * 365 * 24 * 3600 = 4.348 * 10^17 seconds, which (I think?) is a 59-bit number [1]. 10ths of a second will require 62 bits, which is right about at the edge of what a 64-bit signed integer will allow.
If you want milliseconds, you'll need at least 69 bits. For nanoseconds, you'll need at least 89 bits.
So you'll either need an integer type that's wider than what's natively supported in most hardware (thus potentially sacrificing performance), or you'll have to sacrifice precision.
[0]: https://en.wikipedia.org/wiki/Age_of_the_universe
[1]: https://www.wolframalpha.com/input/?i=13.787+*+10%5E9+*+365+...
I'm mostly kidding. But if you're thinking of representing nanoseconds since the big bang, the wait for 128-bit CPUs is not very long...
[0] https://gcc.gnu.org/onlinedocs/gcc/_005f_005fint128.html
Since larger than 64 bit ints are a disaster for portability, the reasonable solution is to go with a 64 bit signed seconds, 32 bit nano offset field. A lot of language std libs have adopted something along these lines.
DJB was advocating for everything to be in a format like this, referenced to TAI (UTC without leap seconds, basically). Sadly that didn't get any traction.
I was horrified even more when i learned that future leap seconds are undefined, and we literally can't tell what is the time on the clock lik3w a million seconds from now.
But of course GNU is kind of notorious in this regard. Compare their `yes` to OpenBSD's. It's night and day.
https://github.com/openbsd/src/blob/master/usr.bin/yes/yes.c
If you want to know what's broken? Most real time clock modules. They almost all want to store time as HH:MM:SS MM:DD:YY and sometimes 1/256 of a second but sometimes not.
rust-lang saves the day, yet again https://www.brandonsmith.ninja/blog/favorite-rust-function
What type (and how big) should length be?
I can understand that this might have been too much implementation complexity/risk to contemplate 40 years ago, but this kind of pattern is very well established at this point, especially in scripting languages with loose typing.
size_t.
Perhaps a committee was involved.
It takes more code to use it correctly than to hand-code what you think it is supposed to be doing for you. strtok is Cursed.
strlcpy is similar. If you don't write that much more code, you are not using it correctly, and it is not giving you the value that is the reason you thought was why you were using it.
That's not really a counter-example. Sure the string handling was an unfortunately choice but it is the standard. The standard library must implement it as defined. Being standard-compliant is not lack of polish.
Ha ha ha ha ha ha ha ha ha ha.
Locales are a massive clusterfuck, basically too simple to handle localization if you actually care about it, but supports enough of it to screw you over if you don't care about it. The "wide character" support is also a nightmare. The time library support is also quite a bit wonky (years are measured as years since 1900 because Y2K is definitely not a pressing issue in 1989!).
> inline Assembly
Fun fact, here is the C specification's entire mention of inline assembly:
> The asm keyword may be used to insert assembly language directly into the translator output (6.8). The most common implementation is via a statement of the form: asm (character-string-literal);
There is no discussion of what inline assembly can and cannot do, how it interacts with the rest of the code in term of semantics, how to pass arguments to and form inline assembly, etc. You might get some of this information from the manuals of compiler implementations, but even that can be surprisingly free of necessary information. Compare this to Rust's inline assembly documentation: https://rust-lang.github.io/rfcs/2873-inline-asm.html (which is more detailed than even gcc's or LLVM's inline assembly documentation).
Everyone of these language brought new ideas, but they don't stand a chance because their designers don't understand the point of C. The C language didn't win because it was the "best" language or had the best set of features. Far from it. Even in the mid 70s it was a backward language compared to other cool languages of the day like Algol and Lisp.
C won the competition because it just gives programmers the bare minimum functionality to put an operating system and a compiler in place! It is flexible, you can provide your own library if you want, and therefore gives your easy portability. OS writers will chose C any hour of the day or night because it makes their job easier.
By comparison, other languages will require a huge library to be available, and sometimes a complex runtime system, just for you to write a simple "hello world"! Imagine if you need to write a new OS, a compiler, a linker, or a shell interpreter... you get the idea.
My conclusion is that language designers still didn't get what made C so successful and therefore keep coming up with shiny complex things that don't stand a chance to become the next C.
People proposing every other language that tried to replace C thought the same. Only time will tell, of course, but I wouldn't bet on it. Nowadays C++ is seen by many as a "garbage pile", the same can happen to rust.
C++ was seen as a garbage pile from its inception by a large number of people.
A significant number of people kept using C because the only alternative was C++ and that wasn't acceptable.
Did you mean API? The C++ ABI situation is famously unstable.
http://www.open-std.org/jtc1/sc22/wg21/docs/papers/2020/p186...
"there is a non-trivial amount of performance that we cannot recoup because of ABI concerns. We cannot remove runtime overhead involved in passing unique_ptr by value, nor can we change std::hash or class layout for unordered_map , without forcing a recompile everywhere etc. etc."
I thought passing template instantiations over an ABI boundary was generally discouraged and thought to be asking for trouble. I guess this isn't really at odds with what the paper is saying though - it could still be that a lot of people are doing so.
edit Thinking about Hyrum's Law, [0] mentioned in the article, makes me think perhaps there was an upside to Java firmly refusing to support any kind of ahead-of-time compilation for so long. It fully closed the door on any funny business distributing Java packages as brittle precompiled native-code blobs, ensuring the bytecode format remained the way that Java packages were distributed, presumably avoiding some fraction of the issues C++ now faces.
Of course, Java still has backward-compatibility obligations, but unlike in C++ they align pretty well with API compatibility, if I understand things correctly.
is a couple ABI breaks for std::string between C++98 and C++20 that unstable ? You can write code that uses std::string today and links against a .so built a long time ago (modulo compiler bugs of course, thus the various versions here: https://gcc.gnu.org/onlinedocs/gcc/C_002b_002b-Dialect-Optio...). GCC & libstdc++ go to great lengths to preserve ABI compatibility.
Just _look_ at all these short keywords and special symbols. It's legitimately hard to read without focusing on each character.
#[bla(foo)]
fn print_refs<'a, 'b>(x: &'a i32, y: &'b i32) {
println!("x is {} and y is {}", x, y);
}
impl<'a> Default for Borrowed<'a> {
fn default() -> Self {
Self {
x: &10,
}
}
}
This old fart thinks Rust is the new Perl.Here's a representative example of the code I wrote. This bit is outputting stuff to a file in some binary format: https://github.com/ValveSoftware/Proton/blob/proton_5.13/med...
E: If you want to take another stab, the way I learned Rust was the book: https://doc.rust-lang.org/book/ Actually type out every code example, it will help your fingers learn the "feel" of the language, and give you an opportunity to break things on purpose to test your understanding. Learning something new is always a challenge, but I really do like Rust and think it's worth the trouble.
[1] https://doc.rust-lang.org/book/ch10-03-lifetime-syntax.html
E: There are only three instances of using explicit lifetimes in the entirety of the project I linked, if you want to see some real-world examples of it. In all of these, it is used to indicate that the struct being declared references another struct which must outlive that struct. That way we don't end up with a dangling reference if the referenced struct failed to outlive this struct.
https://github.com/ValveSoftware/Proton/blob/proton_5.13/med...
https://github.com/ValveSoftware/Proton/blob/proton_5.13/med...
https://github.com/ValveSoftware/Proton/blob/proton_5.13/med...
#[] - why the [] if # already makes that line different from the usual code? seems superfluous
'a - is that thing next to 'a' a smudge on my display? Did I forget a quote? better wipe the display with my finger
'&a - "mut" and "ref" but ' ?
foo! - yelling out function calls. "print!" "exit!" "macro!". Angry Codes!
foo? - when you're done yelling, make sure to ask existential questions of the return result. This one I have the least problems with since it actually makes me question the return value ("hm, something's weird here, it could be null, pay attention"), but when coupled with the yelling, just makes the whole thing look dramatic. "do_it_now!(); did_we()?"
fn - by itself not a huge deal, but the list of truncated words that are used frequently is "impl", "mut" and "pub". I save some characters (am I really in such a rush?) at the cost of reading this broken English "f-n impl moot pahb". At least C doesn't have that. Pascal had "interface" "implementation", "begin", "end", etc. Java has "public", "interface", "extends", "class".
Default for Borrowed - suddenly English! No time to type 'function' or 'mutable' but "Default for Borrowed" is a-ok?
underscores all over the place in the standard library. Even C doesn't have that many, mostly in the _r variants that were added.
unwrap() - what does that word have to do with errors? (I know what it does) It wouldn't be my first choice.
I have a whole list of awesome stuff, meh-stuff and wtf-stuff I noted down about Rust while trying to learn it (on multiple occasions), and there are a lot of excellent things about Rust, but the aesthetics of the language are important, otherwise we'd all have no issues coding in Brainf*ck.
Yes it's possible to go nuts with syntax, but this just feels like the "programmer art" of language syntax. I think it's usable (clearly), it's just not elegant to me.
I don't even know what #[bla(foo)] does, or why all that punctuation is needed. maybe nesting is allowed?
> 'a - is that thing next to 'a' a smudge on my display? Did I forget a quote? better wipe the display with my finger
this is a lifetime annotation, which is a genuinely noisy bit of syntax.
> foo! - yelling out function calls. "print!" "exit!" "macro!". Angry Codes!
I guess they really want to make sure you know when a macro is being used. probably a result of ptsd from debugging c and c++ code :)
> fn - by itself not a huge deal, but the list of truncated words that are used frequently is "impl", "mut" and "pub". I save some characters (am I really in such a rush?) at the cost of reading this broken English "f-n impl moot pahb". At least C doesn't have that.
the c keywords aren't too bad, but the standard library is full of this kind of thing. stdio.h and string.h immediately come to mind.
I'm not so much defending rust as I am pointing out that c's syntax isn't that great to begin with. I've been writing c and c++ code every day for several years now, so it's usually pretty easy for me to skim and understand what is going on. but if I try and place myself in the shoes of a newcomer, I don't think the syntax is much better than rust. remember the first time you tried to parse the type of a nontrivial function pointer?
int and char aren't exactly words. And let's not forget that a type like "double" makes no sense whatsoever by itself. (It's two of something, but two of what? Oh, it's "double-precision" floating point! How could I miss that?)
Now, let's look at C's standard library:
strcmp, strpbrk, isalnum, ispunct, setjmp, SIGSEGV, SIGFPE (that's the divide-by-zero exception, isn't it obvious?), SIGABRT
Sure seems like C has its own massive issue with "we can't let a name be long"
> underscores all over the place in the standard library. Even C doesn't have that many, mostly in the _r variants that were added.
va_arg, FE_DFL_ENV, int32_t, etc. Continue on into most C libraries, and underscores are pretty common because C has no other namespacing mechanism.
Syntax is pretty strange in any language if you're not used to it. If you're comfortable with C, C's syntax and spelling quirks don't stand out to you.
Ahh yes, I believe that came about due to the great "Vowel Bowl" of the 1970s. There was a serious shortage of vowels.
Since the first edition of ANSI/ISO C was trying to codify the already-existing common practices for maximum portability, it reflects those existing limits (5.2.4.1 "Translation limits"):
"The implementation shall be able to translate and execute at least one program that contains at least one instance of every one of the following limits:
...
31 significant initial characters in an internal identifier or a macro name
6 significant initial characters in an external identifier"
If you look at many abbreviated function names from the C stdlib, they are specifically 6 characters long - strcpy etc. I wouldn't be surprised if the 6-char external identifier limit goes all the way back to the first K&R C compilers.
I personally have higher expectations of a language designed in the 2000s by people standing on the shoulders of 40 years of computer science and language research.
Here's a contrived analogy:
Imagine if you bought the newest Tesla truck meant to replace old Ford model trucks and you had to vigorously shake the steering wheel in order to lower the driver's window.
You're saying: "Ford trucks have always used a hand crank and that feels strange too if you're not used to it!"
I'm saying: "Why did they choose to make it equally strange in the first place?"
> #[] - why the [] if # already makes that line different from the usual code? seems superfluous
# does not "make that line different from the usual code." The whole #[] construct is it, it has nothing to do with lines. You could put "#[foo] #[bar]" on one line if you wanted, you could write "#[foo] fn lol()" if you wanted...
> foo! - yelling out function calls. "print!" "exit!" "macro!". Angry Codes!
This helps both humans and computers parse; macro invocations don't have to follow regular Rust syntax, and the ! helps indicate that that's true.
> Default for Borrowed -
This is not language syntax, this is the name of two types. You can name your types however you'd like.
Over the long term, I don't think symbols are harder to read than keywords. C itself provides some evidence for that: imagine reading C code if experience with another language had conditioned you to look past * and &. That would be terribly confusing, but for an experienced C programmer, those symbols leap out of the code at you because you know that they carry a lot of meaning.
fn print_refs(x: &i32, y: &i32) {
println!("x is {} and y is {}", x, y);
}
In general you only need to specify lifetime parameters when there's an ambiguous situation. For instance, the following builds and runs just fine. fn make_substring_of(input: &str) -> &str {
&input[1 .. ]
}
I'll be the first to admit it takes a while to ramp up into Rust. Part of that is learning to let go and trust the compiler. Unlike many other languages, it won't hurt you haha. It's on your side. Sometimes you need to you know, give it a little more info so it can do it's job.The Rust team has made huge strides in improving writability of the language, especially with non-lexical lifetimes.
Soon, `fn` and `impl` and `Vec` fade into the background, like `char` and `short` and `long`.
I am sure it has its place but I think it's just too ugly to be attractive. There is something pleasing about writing C, Python or even Javascript which you will never get with Rust. It will never be a language a lot of people enjoy writing imo.
Not really. C won because it was the standard compiler for Unix, and was a free compiler for a free operating system in a time when both were highly unusual.
But in the past few decades, C has lost a lot of its market share to other programming languages. In the realm of desktop (and mobile) applications, C is basically unused for new projects--its standard library is truly anemic here, and major support libraries (e.g., GUI toolkits) are often in C++ and not C. Where C is still the dominant language is the land of embedded applications, and it's not had a lot of competition here since most languages don't bother trying to define a freestanding implementation.
Rust is really the first language to try to contend this space. And there are signs that it may supplant C: Intel is apparently looking to move its firmware to Rust; Linux is allowing Rust for device drivers and kernel modules. Hell, even some OS programming courses (e.g., Stanford, Georgia Tech) have moved their curriculum to use Rust instead of C.
Oh, did this happen? I remember some discussion some months ago, has it actually been merged?
Dunno how often this class runs, and if it was just for one semester or more than that.
That is a patently false statement. To understand low-level programming concepts, you need to understand fundamental notions about how machines represent state in registers and memory, how memory is organized (including primarily the concept of function calls and the stack), and the indirect referencing of memory via pointers. Note that nowhere in that list did I describe a concept that is unique to C.
In fact, one of the more common approaches to introducing developers to low-level programming is to introduce them to these concepts via assembly (say, Nand2tetris). In my own experience TA'ing such a course, I am more than willing to translate code into whatever language the student is most comfortable with to express the concepts as necessary. You can absolutely learn these concepts in other languages, and my own suspicion is that unsafe Rust does a slightly better job of it than regular C does.
C does not have a monopoly on understanding the low-level organization of code, and quite frankly, C's lack of coverage here can be frustrating. C has no concept of multiple return values, functions with multiple entry points, unwinding the stack, computed goto, SIMD vector types, nested functions, discontinuous structures, or tail calls, and these are all concepts that are present in other languages that cannot be expressed in standard C or often even in vendor-extended C.
> C has no concept of multiple return values, functions with multiple entry points, unwinding the stack, computed goto, SIMD vector types, nested functions, discontinuous structures, or tail calls, and these are all concepts that are present in other languages that cannot be expressed in standard C or often even in vendor-extended C.
ATS does a better job than safe Rust and you can think of it as a vendor-extended C (http://ats-lang.sourceforge.net/DOCUMENT/INT2PROGINATS/HTML/...).
To put it in short, C programmers hate needless abstraction and work with raw memory and hand-crafted data structures whenever possible. More data, less code.
Embedded developers use C because they have to, not typically because they want to. And they have to use C because of proprietary toolchains and existing libraries, among other similar reasons. Good reasons, sure, but not because C the language is so great.
> hand-crafted data structures [...] More data, less code
You're paying lip service to the principle of datastructures-over-algorithms, yet you're advocating for a language which has neither algebraic data types nor tuples? Come on. That's like praising functional programming but using Fortran.
I couldn't afford abstractions, memory safety guarantees, algebraic data types, generics, pattern matching, thread data safety (at least in the context of interrupts). The languages that had these things were hulking languages with giant runtimes and exactly zero support for embedded development. Not to mention no vendor toolchains.
Rust supports all these things, with no allocator, no standard library, and often with zero additional cost -- in terms of compute and in terms of memory. Then, being built on LLVM means that the vendor toolchain support is quickly becoming a non-issue. I suspect we'll start to see more and more of Rust in embedded, but only time will tell.
And as others have called out, Rust has very few OOP features like an optional notion of a `self` on a function bound to a structure. There's no classes, no subclassing, no message passing, no inheritance at all (structure, interface or implementation), limited dynamic dispatch, no polymorphism.
As an embedded developer for a living, if you put it that way, then the first thing it comes to my mind is "so why should I bother learning Rust" for embedded.
Other than the usual "please consider UB and network/memory handling/security issues" (reasons that don't really affect me), so far nobody could provide a convincing answer.
It's not that the usual answer I get is wrong or invalid or doesn't have a point. But if I ask "ok, what else?" there is really little motivation for me to move on.
Someone once told me I'll become an outdated curmudgeon here on HN, but I think I'll be long gone before something truly deserves to replace C for embedded (and low level).
> Someone once told me I'll become an outdated curmudgeon here on HN, but I think I'll be long gone before something truly deserves to replace C for embedded (and low level).
You (and I) may become an outdated curmudgeon before long, but it won't be because you refused to learn Rust haha.
Generics alone are a huge time saver compared to C, even if you use hacks like macros in the latter to do something similar already.
This. Even if it is less safe.
I wrote Monocypher in C for one reason: portability.
From a systems point of view, crypto libraries are trivial: you don't need any dependency, code is pathologically straight-line, there is no almost data structure to speak of beyond arrays of bytes and arrays of words.
Yet I can tell you that if not for portability, Rust would have been a better fit. So I could group buffers and their size in a single argument. So I could provide genuinely high-level interface. So I could use types to avoid silly mistakes and enforce some invariants. Portability won over all that goodness: worst case I can have a Rust wrapper. Heck, someone else already wrote one for me.
In addition operating systems gives you bunch of safety measures if you design around them. (Obvious being separate processes)
One has to also remember that even if your program would be memory safe it still does not guarantee its flawless nor secure. One very good recent personal experience I can give you is certain popular swift SQL library. It was written in swift which also boasts the "memory safe by default". However, that library due to its design was subject to the very basic SQL injection attacks. Needless to say I ended up writing my own SQL bindings. I hope that library is now fixed or something replaced it, as since at the time I had to write swift (and no I didn't like it), it was the only and most popular (scary) option on github.
I could also write long rant about threads and how the implementations and the threading primitives often have subtle bugs or are broken in different ways. And how using threads in general makes your application basically undefined unless you only absolutely know every piece of code that runs on that thread.
Because Rust is redundant, its safety guarantees are a subset of safety guarantees of ATS[1]. A project in plain C integrates more naturally with it at any point of its development cycle[2], and it doesn't require giving up safe pointer arithmetics[3]
[1] http://ats-lang.sourceforge.net/DOCUMENT/ATS2FUNCRASH/HTML/c...
[2] http://ats-lang.sourceforge.net/DOCUMENT/INT2PROGINATS/HTML/...
[3] http://ats-lang.sourceforge.net/DOCUMENT/INT2PROGINATS/HTML/...
EDIT: those who downove, let's discuss the topic in substance and let's avoid zealotry. If you promote and pitch Rust to the audience of C by advertising its safety guarantees, zero cost abstractions, and how well it integrates with C ABI, at least be consistent when it turns out that there's another $TECH that does it safer and more consistently with C programmers' reliance on certain useful features of C.
What's the point of waiting for 1.0.0? It's just a tag that doesn't save you from bugs and breaking changes.
What is industry support? What does it have to do with a team where everyone can read the documentation of the tool that is already built on top of a mature GCC ecosystem and that adheres existing approaches to debugging, profiling, releasing and maintaining C codebases?
> and doesn't use a modern build system or package
Why do I need a separate solution such as Cargo if I can build, package, and distribute everything with Nix and get reproducibility, distributed builds, transparent caching, and environment isolation along the way for free?
So I'd amend this to: C won because it was the only cross-platform low level language that was competently implemented for DOS.
Objective-C, Java, Swift, and C# have become massively successful as application programming languages because C is/was terrible at it. They learned a lot about how painful it was to do basic higher level programming tasks when you are restricted to C's semantics and memory model.
C is great but I don't think it's worth romanticizing since history has shown that C isn't that great for writing anything but systems code. Which is a restricted domain to begin with, and isn't even that attractive for it anymore.
The one thing C has over anything else is interop. The language of FFI is C. There's no inherent reason for that other than history, and it's not super broken so we're not going to fix it.
C sucks when you need convenience of big standard library or safety above performance. There is nothing like C when you care about speed, memory footprint and efficient memory management. It's great other languages took over in areas C is terrible at but it's not like areas when it's the best and often they only option disappeared.
You can absolutely beat C in everything you list. And you also can't. I don't agree that C is the end all be all of performance or code size, except in a handful of cases where there's nothing else available.
C started to be popular in the microcomputer world at a time when systems programming was being done in assembly languages. For instance, most arcade games for 8 bit microcomputers, were written in assembler. Some applications for the IBM PC were written in assembler, such as WordPerfect.
The freedom with pointers thanks to arithmetic would instantly make sense and appeal to assembly people, who would find a systems language without pointer flexibility to be too strait-jacketed.
C++ too to a lesser extent. I work on spacecraft flight software, and there's a significant push to move from pure C to modern C++.
No single language is going to (or wants to IMO) replace C in every single use case, but replacing C in specific use cases has been a huge boon for productivity.
Otherwise, I can't think of any circumstance Rust isn't trying to muscle in on C and C++'s territory.
New languages mostly don't replace existing ones. Rather, they supplant them for some uses, and open up new kinds of software which are easier to implement or to conceive with the new language. Now, you referred to C++ specifically, and since I'm somewhat familiar with it I'll address some of the points you made with respect to just that one:
C++ was not intended to "replace C", but rather to combine features of BCPL (later C) and Simula. See: https://www.youtube.com/watch?v=69edOm889V4
Bjarne Strousup said: "If you want to create a new language, a new system - it's quite useful not to try to invent every wheel." For a long while now, C++ teachers/trainers encourage their audience _not_ to think of C++ as an "augmented C" or "C with feature X Y and Z", and to avoid most "C-style" code in favor of idioms appropriate to what the language offers today.
Also, C didn't "win the competition" because there isn't a "bestest language for everything" competition. It has been, and is, a popular language with many uses. Writing operating system kernels is one kind of programming task, where C is the most popular. Even at this level (and lower still), other languages are potentially interesting and often used. See, for example:
* IncludeOS - Running C++ without an operating system: https://www.youtube.com/watch?v=cQPrtTsM7Zg * Generating optimal assembly in an embedded setting at (C++) compile time: https://www.youtube.com/watch?v=CNw6Cz8Cb68
Finally, C++ doesn't require a huge library nor a complex runtime system because it has a "freestanding mode" in which the requirements are very limited (although more is required than for C). See: https://en.cppreference.com/w/cpp/freestanding
Operating systems are currently written in C, and therefore have a C interface. Which language is best at interfacing with C? (No trick here, just a rhetorical question.) C of course. Other languages would need some sort of FFI, which is generally unwieldy enough that the designers hid it behind a comprehensive standard library.
C doesn't need an huge library to be available, you say? Oh but it does. It's called the kernel. Comes with a freaking huge runtime too.
> Imagine if you need to write a new OS, a compiler, a linker, or a shell interpreter... you get the idea.
I think I do, but I'm afraid you don't. Writing an OS (and the rest) in Pascal (I'm thinking of Oberon specifically) is no harder than to write it in C. If you write your whole OS in Blub, interfacing with Blub will be easiest in Blub, you won't need an extensive Blub standard library because you already have the kernel, all the tools (debuggers, editors…) will be Blub friendly…
Lisp machines used to be a thing, you know.
> My conclusion is that language designers still didn't get what made C so successful […]
Language designers can't even address what makes C so successful: network effects.
Not true. Operating systems are written in C with some assembly, and the interface to userspace is universally written in assembly. E.g., https://git.kernel.org/pub/scm/linux/kernel/git/stable/linux...
Which by some eery coincidence happens to conform to the C ABI of the platform. Come on.
I stand by that claim: the only reason operating systems have a C interface is because they are (at least originally) written in C. The fact that the kernel has some parts in assembly is immaterial, even if those parts happen to comprise the userland interface.
And it's not just the kernel: when UNIX was re-written in C, everything was written in C. The compiler, the core utilities, the editor, the shell… the whole OS, not just the kernel. Of course it was easier to interface to that in the same language everything else was written in.
Likewise, the Oberon operating system, which was written almost entirely in the Oberon language (which should have been named Pascal-3) has an… Oberon interface. Want to use C to interface with it? My, you'd have to write a whole compiler, perhaps a non-trivial runtime, debugging tools, and of course an FFI to interface to Oberon, the de-facto linga franca.
C looks much worse when it's not already the king of the hill. Its strength lies in its ubiquity more than in any quality of the language itself.
PS. I appreciate you reaching out and making an effort to understand why the comment seems snarky to me.
You can't because it is impossible to escape the type system. Without the ability to cast pointers you can't write a memory allocator. You can't write a function like dlopen()...
Not sure about the memory allocator, Oberon may have avoided the problem by having a GC.
http://www.cs.virginia.edu/~evans/cs655/readings/bwk-on-pasc...
Here's what low-level programming in Pascal looked like on DOS: http://www.retroarchive.org/swag/TSR/0022.PAS.html
it did well enough that C stdlibs (at least MSVC's, LLVM's) and compilers (... pretty much all the big ones) are implemented in C++ and just export C symbols nowadays, likewise for newer OSes like Fuschia.
SerenityOS (https://github.com/SerenityOS/serenity) was written from scratch in two years in C++ and goes as far as having a custom web browser & JS engine. Where is the equivalent in C ? Where are the C web browsers, C office suites, C Godot/Unity/Unreal-like game engines ? Why is Arduino being programmed in C++ and not C ?
Imagine a world where most the languages follow Lisp's syntactic conventions, or Pascal's. What a nightmare.
The default is not the best, the default has just beaten the world so many times on the head that anything else became foreign and weird and got laughed out of the room before it even had the chance to say anything.
One programming language history book/blog/paper a day keeps the nonsense notions away, C "Won" like people win the lottery or the roulette.
var x: *int;
more often than: int* x;
i.e. type name follows variable name, and type modifiers work more like unary prefix operators on types. And this is because it's less ambiguous to parse, and makes more complex types a lot easier to read, since you simply go left-to-right, instead of C's "spiraling" declarations.The rest of Pascal's syntax is not particularly problematic, either. I'd say that the two biggest problems with it were begin..end for blocks, and having all local declarations in a block separate from code. But Modula-2 already dropped "begin", and various Pascal dialects added inline variable declarations eventually. So, on the whole, I think we'd actually be better off in terms of code readability if Modula-2 rather than C became the standard systems programming language.
Say you add fn and var as keywords. Exactly how hard would that be to fix? You could probably write a tool to do that.
import core.stdc.stdio;
int main() {
printf("hello world!\n");
return 0;
}
and yes, it is calling C's printf. To build it: dmd -betterC hello.dProgrammers have been show, time after time, not to be particularly trustworthy. Have we not learned the lesson that it's really easy to make mistakes, and we should trust tools instead of people to check our work?
> Don’t prevent the programmer from doing what needs to be done.
Ditto the above.
> Keep the language small and simple.
It is in some ways, but it's "smallness" leads to a serious lack of simplicity as seemingly simple things are incredibly hard to do right consistently. For instance, avoiding indexing past the end of an array or rolling over an integer.
> Provide only one way to do an operation.
That's nice, I'll grant you, although there are of course exceptions that prove the rule, like:
a[b]
is synonymous with *(a + b)
*((uint8_t *)a + (b * sizeof(*a))
> Make it fast, even if it is not guaranteed to be portable.Well... I mean...
if (((a > 0) && (b > 0) && (a + b) < 0) ||
((a < 0) && (b < 0) && (a + b) > 0)) { / * Overflow */ }
Now you may be saying, well, isn't there a branch-if-arithmetic-overflow instruction in practically every single architecture ever? To which I would say simplicity matters.[edit] </sarcasm>
T const highOrderBitMask = (1 << ((sizeof(T) * 8) - 1));
T const hobA = (a & highOrderBitMask);
T const hobB = (b & highOrderBitMask);
T const hobR = ((a + b) & highOrderBitMask);
if ((hobA == hobB) && (hobA != hobR)) { /* Overflow */ }
Or is it undefined behavior to assume a 2's complement representation also.[edit] Yep, 2's is only required for C++20, not C, and worse, it's permitted for an implementation to trap on overflow. Back to the drawing board.
The representation of signed numbers is implementation-defined, not undefined.
if (((a > 0) && (a > INT_MAX - b)) ||
((a < 0) && (a < INT_MIN - b))) { /* Would Overflow */ }
Not standard, but not undefined either are the checked intrinsics: __builtin_add_overflow(a, b, &x) == false
Thanks, that was a fun exercise :)When I've learned that RISC-V had no carry bit, I couldn't help but think it might have been designed for C to begin with. Sure, they give reasons for this choice, none of them linked to C. Still, it hurts multi-precision arithmetic any language with BigInt would have benefited from (I recall Python, Haskell, and Scheme at the very least).
I'm no hardware designer, though.
I thought this was funny because these two are the same:
a[i]
*(a + i)
Likewise all of these are the same: (*s).x
s[0].x
s->x
As are these: if (x)
if (x != 0)
But I guess the point can be applied elsewhere. 0[x]
x[0]* It's everywhere
* There's a standard ABI
there's one ? the "C" ABI is just the ABI of whatever platform it's running on, which may or may not have funky behaviour that vendor-provided compilers kindly hide for you - functions being prepended with '_' on macOS, the two-dozen calling conventions on windows with i386, sysv and itanium ABIs...
Do you think you can tell what's the ABI of
struct foo some_function(struct bar);
?will bar be passed in a register, on the stack ? who knows, that depends on your platform, your compiler, etc etc. Things going cleanly on the stack is just a convenient lie that your first year comp. sci teachers tell you because it's too early to talk about how the real world works yet.
But typically the C ABI is the only stable ABI those platforms provide. That's a huge benefit.
https://developer.r-project.org/Blog/public/2020/11/02/will-...
https://developer.apple.com/documentation/xcode/writing_arm6...
https://docs.microsoft.com/en-us/cpp/cpp/argument-passing-an...
As far as the ABI goes: the important thing is that there is a standard ABI on a specific platform that all compilers on that platform agree on. Sounds kinda obvious, but it's not common in other languages.
it definitely is - see e.g. https://github.com/llvm-mirror/lld/blob/master/lib/ReaderWri...
What is the nit of it is it almost works. You have a good shot at getting it to compile in a short amount of time. The rest of the work will be lots of time in ye old debugger and going over the docs for your platform. The fun part is you will find bugs that were there already, or are they just part of the platform, or were you using it wrong?
In effect it is yet another CRT and the idea is sound. But many times what I found was you may have one that works on say linux and windows and bsd. All the same 'code' but you dig under the covers a bit and it is a maze of ifdefs so each platform has its own quirks. For example threading between fork and createthread is on the surface not too different and you can wrap createthread with it (several libs did). But you dig into it a bit and you find portions that just do not map at all between the systems (usually with IPC and locks). At best they do not compile, slightly worse they return error codes, at worst they act like they work.
A real good example of what I am talking about is the pthread library. It works up to a point but it is a very linux/bsd orientated library. There are some gaps in there from windows that just do not map and the other way around. What is worse is the docs on some of these do not talk about cross platform issues. Luckily you can see the source code of most of them and can tell what is going on. Annoying but one of the things I learned moving code between platforms is that each one has its own way of doing things. You can try to work against it or sit down and unwind what is going on, which takes time. I have even seen this sort of issue in python and java. Where you get down to some low level thing and it just is different on different platforms.
The moment you start doing things like threads or shared libraries, yeah, it all breaks down very quickly. But that isn't standard C.
But I can't think of any implementation that doesn't conform to C90 in that regard. So long as you don't venture into implementation defined / UB category...
Sure, there are compiler switches and language extensions that can break the ABI if you use them. But, well, you don't have to use them (at the interop boundary), and neither do your API clients.
But the C ABI is awful to work with. The C language itself offers no help to guarantee ABI compatibility. What ABI it compiles for depends on headers, which may depend on a jungle of ifdefs and typedefs.
[1] But you need to write a bit of C glue code for OCaml so it's not quite so seamless.
Gamedev tends to use its own patterns, particularly arena allocation for long lived fixed size tables, or their own internal object / entity / component model . It's not really c with classes or c with templates, just kind of its own dialect.
It's less template metaprogramming and just templates to generate efficient code - stuff used to be done by abusing the proprocessor can now be handled by a (slightly) more elegant templating engine rather than a string pasting engine.
Classes are used for resource management/RAII, like we have a AQUIRE_MUTEX_IN_SCOPE() macro which will release mutex when scope is exited, this is supremely useful and generalized to many resources.
Lastly namespaces are huge. In big C codebases you have to be super pendantic about naming modules and APIs consistently because otherwise it becomes a nightmare.
C++ does get in the way still sometimes, like when you want to do something slightly dirty for perf reasons, say aliasing between structs. You first write it in a way that makes sense, basically how you would write it in C, but it's UB in C++, so you rewrite it with virtual calls or memcpys such that the compiler should be smart enough to arrive at the same result as C would have with the straightforward implementation. This works great until it doesn't, your last option is to try to solve it with templates, and that hole is very deep.
Actually, only half of the universe is using C++ to write games, the other half is using Unity and writes their game code in C#.
If language interopability would be dramatically better, this "lock-in" into a specific programming language wouldn't be half as bad as it currently is, and it would be much easier and less risky to use "fringe languages" for game development.
Check out the Doom 3 source code ;-)
[1] http://ats-lang.sourceforge.net/DOCUMENT/INT2PROGINATS/HTML/...
[2] http://ats-lang.sourceforge.net/DOCUMENT/INT2PROGINATS/HTML/...
Maybe consult https://en.wikibooks.org/wiki/C_Programming for any syntax you don't recognize.
This seems promising, would be nice to hear what HN thinks of it
But I would not recommend it to someone, who has no experience and wants to learn C as a first language.
I would also recommend using the Steve Summit's "C Programming Notes" as you work your way through K&R. Pay attention to the "deep sentences". https://www.eskimo.com/~scs/cclass/krnotes/top.html
And don't skip the exercises at the end of each chapter. The discover of solving the exercises on your own is a revelation that far surpasses any benefit obtained by being told how to do it! :-)
http://man.postnix.pw/9front/2/intro
Also might help if I add a paper on porting plan 9 which illustrates some of the niceties of the C library and portability ease. http://doc.cat-v.org/plan_9/4th_edition/papers/libmach
Of course I'm not saying C is not useful anymore (it still is, for example if you are doing system/kernel programming, very likely you'll deal with C). But this is not the typical case for beginners.
I'm not entirely sure what you mean here, but this is true for neither of the interpretations I have. Most compilers are not written in the C language (indeed, even the major C compilers are all written in C++ now!). Alternatively, most languages compile down into a bytecode language, or into native assembly, without converting into anything that is or looks like C in the process. Indeed, for languages targeting native assembly, C is an unwelcome intermediate step precisely because its semantics can be too constraining.
I mean that most of the popular languages I'm familiar with are written in C.
Even looking at major languages [1], what do we have:
* Self-hosting languages: C#, Go, Haskell, Java, most LISP implementations, OCaml, Pascal, Rust, Swift
* C++ implementation: C/C++/Objective-C compilers, Fortran compilers (using the same toolchains as the former), JavaScript, probably Visual BASIC (although that may be C# instead)
* C implementation: Perl, PHP, Prolog, Python, R, Ruby, Shell (although note that many of these languages have their libraries largely written in their own language).
* Not sure: ALGOL, APL, Cobol, Erlang, Forth, Kotlin, SQL, Simula, Smalltalk. Although I suspect that many of these are self-hosted.
7 of 32 is a far cry from "most languages".
[1] Using the list of programming languages in Wikipedia's category box at the bottom here.
That aside, you need to account for multiple implementations. E.g. for Python, CPython is written in C, but PyPy is self-hosted (kinda; it's complicated: https://doc.pypy.org/en/latest/architecture.html#layers).
Fun fact: even major C implementations are mostly in C++ these days.
Ruby, Python...PHP. What else?
Two examples: