C for Rust programmers
bd103.dev
bd103.dev
Though the author doesn't look to be trying to write a great tutorial, rather prioritizing sharing what and how they learned something from their perspective.
Please stop smashing all the nuance out of compiler diagnostics this way.
One of the biggest successes of Rust has been its great diagnostics and I can assure you that it would not help to smash my Clippy lint suggesting that what I wrote looks a lot like an implementation of the addition operator† into a fatal error like the one I get for forgetting to initialize a variable.
Rust even has a (begins empty) diagnostic category [named "expect"] for "This warning should be here" which will flag cases where a notable thing not only might happen here and if it does we can ignore that, but if it's no longer detected that is itself suspicious and should be diagnosed.
† Yes it does Clippy, and I considered implementing Add but I had a good reason not to, so here is a suppression annotation.
Tribalism is sooo ... sadly now.
Please don't.
C got lint in 1979, and is incredible how much advocacy is still required for folks to use the tooling their compilers offer out of the box.
This problem isn't even C specific, other ecosystems suffer from similar adoption issues.
Or if you turn your argument around: if people depend on workflows and automation (I do of course too), then this is a good reason to standardize the diagnostics and errors of the checking tools to the point where the workflows and automation can depend on them. Otherwise they will remain tools which are completely optional to run, or even worse, are not part of your normal work flow.
However, like in any ISO languages, that only happens if enough people vote for it, and someone is willing to submit a paper in first place.
In some circles having the compiler dictate the way, and only true way, is seen as straightjacket programming.
But this illustrates how absurd the discussion is nowadays. Incrementally making things stricter and fixing warnings along the way is one of the most cost efficient way to improve quality and safety. Rewriting the software one of the most expensive. So if the former is rejected because the burden is too high, then obviously the overall aim to make things safer can not be too important and the push for rewrites must partially be motivated by other interests.
"The Power of Ten: Rules for Safety Critical Coding by Gerard Holzmann"
https://youtu.be/GRJtYwneG2Q?is=M0z7hnpQ78C8fLLP
At a moment he mentions a trick he played on JPL folks, his static analysis tool, instead of showing all the issues found, would only display the top 10, without telling them there were actually more than 10.
So they felt motivated to fix them, it were only 10 after all.
Naturally eventually they became aware of the trick.
Requiring all of them to adopt common linting rules at the same time is infeasible.
Also, it’s not true that linting rules never become part of the compiler. For example, the original lint (https://wolfram.schneider.org/bsd/7thEdManVol2/lint/lint.pdf) states
“The type-checking rules also require that, in structure references, the left operand of the -> be a pointer to structure”
and mentions warnings on such statements as
int i 1;
i =- 1;
Modern C compilers also warn way more often about unreachable code, pointer incompatibilities, etc. than those from the 1970s- zero-initialization via `= {0};` will still leave padding and any data not covered by the smallest elements of unions uninitialized, so serializing data structures zeroed via that method is unsafe
- A pointer that increments any more than 1 past the end of a valid memory region (e.g. the end of a buffer) is instant undefined behavior even if the pointer is never dereferenced
- strict aliasing is on by default for pointers of differing types, but not for void/char/uchar. This means that having `struct sockaddr` and `struct sockaddr_in` pointers pointing to the same struct is UB.
The more I learn, the more I run from C.
You really should move to C23 and use ={}.
One of my favorite things to show just how bare this is in C is to show array access commutativity.
char c = {1,2,3}
c[1] == *(c + 1)
*(c + 1) == *(1 + c)
c[1] == 1[c]
C is wonderfully simple at times.Also arrays aren't pointers in C, they decay into pointers. You can see this since sizeof will work differently in the function that instantiates the array versus one that takes in the pointer as a parameter.
At end of day minimalism is a neat but not decisive feature. If minimalism was decisive we'd all be writing Brainfuck.
In C the minimalism comes from providing only the minimal set of things that are available in most (maybe all) architectures (the Von Neumann paradigm), like linear memory, a simple function calling convention, close mapping to Assembly operations etc. But C syntax is not very minimalist compared to Lisp , let alone Forth. The fact that C syntax became prevalent in the programming world seems to be mostly an accident to me, it’s not objectively better than those minimalist languages’ or Pascal’s, Prolog, ML families.
By that logic Brainfuck is even more minimal. Again. I'm saying minimalism isn't the goal. It's a good quality but not most important one.
The rise of C in the 1980s induced CPU makers to design instructions that cater to C semantics.
Edit: And by "worked everywhere" I mean C did the heavy lifting of being a programming language that was even a little bit portable across architectures.
In contrast, Intel introduced a lot of instruction for PASCAL, e.g. enter / leave etc. which are basically unused.
This is beautifully illustrated in the old classic The C Companion by Allen Holub which overviews a simple abstract demo architecture and code generated for it. This is a small book (in the vein of K&R C) but contains enough illuminating info. for a programmer.
Zoom out a little: C is a lot like Unix: simple probably isn't the right word; `underengineered` comes to mind. Which leads to complexity, as you need to make things work in the real world. And so Unix syscalls being designed in the early 1970's for a PDP-11 don't really map to modern needs. And so every unanalyzed complex C app has memory leaks.
To be clear, I like C, I like C++, I like Objective-C, I like Swift; I'm comfortable in all of them. But, in 2026, I don't see why for native code everyone shouldn't be programming in Rust / Swift / other modern memory-safe flavor for any new production use. Zig if you want faster compile times, I guess.
Because C is very simple (unlike Rust) and works everywhere. Zig is not yet stable. Go and Swift are under the governance of tech companies.
The true appeal of C for me is the standard and how it applies only to the language. You can easily take a project from 2 decades ago and port it to a current platform. Lot of current ecosystem is way too fussy about tooling to do this.
That's definitely true. C is a "Worse is better" language and as such it's everywhere. You can knock together a halfway usable C for some crap hardware in a few weeks and then the hardware is saleable because there's a C implementation.
> You can easily take a project from 2 decades ago and port it to a current platform.
Much of the software I wrote two decades ago in C won't even build today. Good luck figuring out why, periodically I try to figure out which weird 2000s era Mac OS hacks don't like a 2026 Linux system and eventually I give up and write modern software instead.
In theory C written in 2006 definitely "just works" on a 2026 Linux machine but in practice real world C is broken and even though I (co)wrote it I don't know why. Also in practice the first serious Rust project I wrote in 2021 I just found it, checked out the oldest working version ("first rough working code" says the log) from git, cargo run, works as expected.
Maybe it's because the C was four times older, but I think it's because in the real world you don't end up writing that mythical portable C code too often.
'C' seems to be a language that makes it easy to write incomprehensible statements, but it's not that hard (or it was not that hard) to write code that is trivial bring up to date. Perhaps it was easier for me because I started with Fortran.
Do you also put Java in quotes? Rust?
I am quite sure my MS-DOS C code is going to have some hard time being compiled on GCC 16.
It's because C, over time, went from a bunch of almost compatible implementations, to a standard that differed a little from every existing implementation, to a standard that is updated every decade or so, which is then implemented by a bunch of different products, with enough ambiguity in the standard (UB) that there will be differences between compilers and versions of compilers.
Rust is a single implementation. Always has been.
It's a single product, so you code to the product. The product has no external forcing function, like a language standard, that produces changes every decade. When you write Rust, you aren't thinking "Wait, lets make sure this works on the Watcom compiler too".
If Rust code cannot compile 30 years later, that's a failure on the language, because the single implementation is in full control of compatibility. If C code cannot compile 30 years later, that's not a failure of the language, because the implementation in 30 years has had external pressure forcing changes.
People often forget that Rust is a product, C is a specification.
In practice the cadence is exactly half that of WG21's breakneck C++ pace. Every six years. C11, C17, C23 so far and C2y is expected to land before the end of this decade.
But anyway, all I can do is report the observable fact. If you want to say it's because GCC 15 is a different product from GCC 4 whereas rustc 1.99 is the same product as rustc 1.49 then I guess you can re-phrase my position as "In C you will keep needing new products and that sucks" if you like.
Then they get irritated that not all C compiler vendors have the same language extensions or implementation defined behaviours they got used to rely on.
It's also easy to write C that won't build on the system it was originally developed on, a few years down the line. But that is true for almost every other language ecosystem. I would say, C makes it _possible_ to write code in such a way that it works on many different platforms and compilers, and/or can continue to work in decades to come with relatively little maintenance.
The latter isn't true for most other ecosystems, because to get anything done in them, you need to rely more on that language and its ecosystem -- specific language features, specific libraries. C isn't a glue code language. It's very good at letting you create your own thing.
To do it well though, requires a lot of care and expertise.
Is that the reason? I thought it was primarily because C stupidly* makes a ton of platform-specific things more convenient than the platform-independent equivalent - the width of integer types, locales, etc., all vary by platform, and it's easy to accidentally depend on them.
* Maybe it wasn't stupid at the time when C was developed, but it's a bad choice now.
Much bigger problem is understanding the scope of dependencies. Dependencies can go bad and you need to update. A dependency might not be available on some new platform so you might have to replace with something else on that platform. To make the software portable, good modularity is required. This is to a large extent an aspect of software architecture. Even Rust can't magically make this happen, if you depend on 300 crates that's probably not a great place to be in either if you want to be portable and maintainable.
There is probably a point that C's flat namespace is a major contributor to badly designed software, because people aren't aware of the dependencies they're mixing all the time. Also C encourages transitive includes, leaking implementation details to the user instead of just to the compiler. (But btw. I find C++ and Rust namespace to be unergonomic syntax-wise, and what's needed is not actually namespaces but control over visibility).
* Programmers are lazy, they're going to type int more often than they should
* All these fancy modern "sized" types are actually just aliases for C's built-in types like int chosen to match the sizes people want...
* ...but C's integer promotion and silent rules about conversions work only on "real" C types you were trying hard not to think about and all the rules for those types, including types you never mentioned but were using because of promotion, apply to your program. You can write code which looks like it's all unsigned, but oops, under the hood a promotion means a signed integer came into existence, then it overflowed, now you have Undefined Behaviour.
I am interested to hear what you don't like about the namespace rules, in Rust in particular.
I also do not see why the simplicity of C should lead to complexity. I have build complex systems in C. I did this using C++ in the past, but switched to C because I found out that all the complexity of C++ was a huge distraction and not something that helps me design better systems. C++ also came with a lot of claims that all the features are absolutely needed to build good abstractions and that this is not possible in C, something I found to be completely false. (And then, a lot of complex and still extremely reliable software I use daily is build in C. My practical experience completely contradicts the myth propagated nowadays by some that all C is inherently bad and unreliable . In my experience it is the exact opposite: The C software I use is in fact the most reliable.)
My view is that Linux and free software is still going strong, because it managed to largely keep out the unnecessary complexity out of the core infrastructure. With Rust and AI, I worry that now get a lot of overly complex infrastructure and that it will be hard to ever get this complexity under control again.
The lack of OpenMP for Rust etc. seems like the most obvious reason for "why use C instead of Rust in 2026" to me.
In most languages, the indexing operator is not commutative. In Rust, pointer offseting is expressed as a function call or a method, also not commutative. I have never seen a single complaint about either. It's not something people want or care about, it's just a tedious detail.
It’s not something idiomatic, but indexing in C is syntactic sugar. Not sure why they allow it in the syntax, but forgetting that arrays are pointers and not special type is just asking for bugs.
Yesterday, I watch a quick video[0] where Matthew Butterick was comparing book sizes and their appeal. “The C Programming Language” was my second programming book (after one about JavaScript 1.x) and I still remember it fondly. Easy to start with (with CodeBlocks on Windows and gcc on Linux) and the concepts were nicely explained. The book were also very nice.
I don't mind adding new features in committees but if they are going to remove compatibility, they should rename their language to nuC or CantbelieveitsnotRust(yet).
No offense intended to Rust which is a fine language.
> C11 is last C as far as I'm concerned.
Are you sure? I think C23 brought so many enhancements it's a pity to ignore it.
"C++", perhaps? :)
> if they are going to remove compatibility
C as a language has one job, and it is to not do this.
> If your machine runs out of memory, reversed[i] will cause a segfault upon being accessed. In order to avoid unhelpful segfaults, the program should check for a null pointer and gracefully exit if one is found
If overcommit is enabled (often by default), malloc will not return null when machine runs out of memory. malloc will still return null if input argument is invalid (e.g. too large).
Isn’t this implementation details, and an information already available caller side. It would be like returning the filename for a file handle. The more extraneous details in an API, the less flexible it is.
Obviously this is untrue. Under GNU we have malloc_usable_size(3). What's true is that it never became standardized, but standards for just about anything in C are very bare, so it's not too surprising.
https://www.open-std.org/jtc1/sc22/wg14/www/docs/n3899.pdf
Many fast allocators do not store the size next to the allocation, and may not store the requested size at all. This has the advantage that freeing a large allocation does not fault in cold pages with a write, as well as reducing allocator memory overhead.
> nobody in decades thought to make this information accessible to the programmer
> otherwise free could not work
I'm inclined to say that you just answered your own question as to why.
I was always taught that "There are parts of a binary that are there for you, the lucky high-level language programmer. And then there are other parts of the binary which are there for the compiler to clean up after the spoiled little high-level language programmer who doesn't have to write their own assembly :)"
and
"If you want to argue with the compiler, that's what `gcc -S -fverbose-asm` and $YOUR_FAVORITE_TEXT_EDITOR` is for".
But I'm curious if this advice is unrealistic for said exotic usecase.
Do tell?
I’d recommend anyone (including the author) wanting to understand C to read K&R’s “The C Programming Language”, which among other things will illustrate how iterating over pointers is idiomatic in C (though not quite in the way the author’s example does it).
https://docs.gtk.org/glib/data-structures.html#doubly-linked...
GSL GNU Scientific Library for C:
https://www.gnu.org/software/gsl/doc/html/intro.html
In general, no one should be custom building most basic structures in C these days. =3
That boolean is actually mostly an extension of the integer system whereby we now have an integer type that stores only one bit.
Whereby the bit stored either results in a 'true' or 'false' value.
Anyways, I know booleans are useful in systems development when you have strict memory / storage constrains / bandwidth (networks).
Yeah, that works for me.
Also for most code it will be premature optimization to worry about that.
std::vector<bool> is a perfectly nice growable bit array type, and if the exact same code were in the C++ standard library named std::growable_bit_array nobody would be annoyed about this type, some people would use it, others would ignore it, nobody would write epic rants about it or name it the singe worst thing about C++.
The problem is that C++ popularized generics, and std::vector<T> is a generic growable array type, you ask for a std::vector<Goose> you get a growable array of your custom Goose type, great idea, very popular these days -- yet std::vector<bool> is not a generic growable array of bool, it's this other thing instead that's similar but not quite similar enough to be a drop-in replacement.
In C++ there is no way for the specialization to be bit-based without it being apparent that you are not in fact getting a growable array of the bool type.
Because that is the problem? That's what C++ programmmers are unhappy about.
> ideally we'd have a better API for vector that might require C++ language changes
Ideally after almost thirty years "maybe we could re-design the whole programming language to paper over this bug?" would be such an obviously bad idea that I wouldn't see it in a response.
Meaning, if you have a struct and you want to bit-pack your booleans, you can declare each one as, say, uint8_t some_bool : 1;
You may then do `x.some_bool = true;` etc.
It's a small nicety to avoid bitwise operators, anyway.
I suspect this is because they are usually introduced like "here are bitfields, but DO NOT USE THEM for anything that might become ABI because their order is implementation defined". I think that causes people to mentally file it under <sketchy language features to avoid> when actually there are many contexts where the ABI risk is not real.
CPUs don't have types, everything is integers or floats. You do your work on registers which have fixed sizes. CPUs have built in instructions for "is this register not zero" which leaks into C. 1 is true in C, but so is 2.
Also, single bits are rarely used for booleans because it requires more CPU power to extract a single bit. Everything is byte aligned at a minimum.
When doing something like network code, if you want to store a bunch of booleans you are typically going to either pack them into a byte, or you'll burn the extra bits and send a single byte for the boolean value. Typically this was flags and masks.
SQL Server still has no boolean type and groups bits in the same row into a byte if possible.
Not necessarily, on modern CPUs memory access is often the bottleneck. If you have many bits storing these in a bitarray can be quite beneficial for performance.
> single bits are rarely used for booleans because it requires more CPU power to extract a single bit
I don't believe this to be true; (on x86) `cmp X, 0` is exactly the same amount of operations as `test X, pow2`. One of them is a `sub` and one of them is an `and`. I believe this is how it works on all important CPU architectures of today back to the early microcomputers (although I don't know how the minicomputers handle this, so it's possible that booleans being integers makes more sense on a PDP-11 or a honeywell).
If malloc() fails, there is no need to exit the program completely with exit(), only return from the current function with an error.
So it is a good advice for beginners.
In fact I've written a whole interpreter that can recover from OOM by raising a recoverable exception to the user. It's really only because Zig made recovering idiomatic, and I'm not sure I could've done it in another language (maybe Rust but I'd have to rewrite large parts of stdlib to both return an error and take a custom allocator).
I’m not sure when those are becoming stable, but Rust for Linux has been driving a bunch of this work, in my understanding, so that’s helped a lot.
What Zig does is great, but it is not actually comprehensively checking that all allocations are fallible and handled. That's not possible on Linux.
I think it's because the new_uninit_slice call Rust will trigger a panic? Or abort? With little-to-no chance for recovery? (I know little about Rust.)
If so, I can see why someone that someone coming from Rust might consider exit() to be the appropriate solution for C, even for library code which should never be in charge of deciding how a program should exit.
I think the essay could be improved by highlighting the different worldviews.
I'm also old enough that
// Allocate enough room for the string and its null terminator.
char* reversed = malloc(len + 1);
makes me nervous. Even for char -- I've never been on a system where sizeof(char) != 1 -- I want to see the sizeof included in the calculation, like: char* reversed = malloc((len + 1) * sizeof(*reversed))
so I don't have to think about sizeof(char) being special.As long as I'm here, I'm a bit confused about the purpose of the "char* error_message" in the proposed Result. Why a char* vs a const char * or even better, an int with an error code? Who sees the message? Do we expect they know English, or will they be localized? Will the error message text be frozen forever, or might it change in the future?
And you never will, since sizeof(char) is guaranteed to always be 1.
I'm guessing you were thinking of CHAR_BIT != 8, but even then I'm not sure it would make a difference since malloc takes its argument size in bytes and a char more or less is a byte in C.
(Consider that char*s are also how you access the byte-level representation of objects in C. If chars were not the minimum addressable unit then that use wouldn't work)
https://smd.hu/Data/Analog/DSP/SHARC/C&C++%20Compiler%20&%20... says the cc21k compiler for ADSP-21xxx DSP systems has char as 32 bits signed, and that the compiler handles ANSI/ISO standard C.
So I don't believe your statement "sizeof(char) is guaranteed to always be 1" is correct.
> chars are also how you access the byte-level representation of objects in C
Where does the spec say that a char can be used to address any point in an object?
There's all sorts of oddities like tagged architectures which the C spec handles which I know essentially nothing about, but which break common expectations about how C works. I believe this is one of them.
I believe the following is undefined behavior in C, even though your compiler may let you do it, at least sometimes, and on modern desktop hardware:
int i = 12345;
char *s = ((char *)&i) + 1;
char c = *s;
I believe the following is the correct (or less incorrect) way to do it: char tmp[sizeof(int)];
memcpy(tmp, &i, sizeof(int));
char c = tmp[1];Yes, but in C standardese a "byte" is not necessarily the 8 bits that it's normally thought to be these days. From C89 Section 2.2.4.2 Numerical Limits [-1]:
> maximum number of bits for smallest object that is not a bit-field (byte) CHAR_BIT 8
i.e., CHAR_BIT is the number of bits in a byte. C23 has a similar definition, and further defines CHAR_WIDTH that is defined to expand to the same value as CHAR_BIT.
> So I don't believe your statement "sizeof(char) is guaranteed to always be 1" is correct.
From C89 section 3.3.3.4 The sizeof operator [0]:
> When applied to an operand that has type char, unsigned char, or signed char, (or a qualified version thereof) the result is 1.
This wording remains basically identical through C23 [1].
> Where does the spec say that a char* can be used to address any point in an object?
From C89 section 3.3 Expressions:
> An object shall have its stored value accessed only by an lvalue that has one of the following types:
> <snip>
> * a character type.
This also remains the case up through C23.
I think you're thinking of the strict aliasing rule with your example, but character types are one of the exceptions to said rule so I think your example is actually fully defined. It'd be UB if you casted to an incompatible type like a float, I believe.
[-1]: https://port70.net/%7Ensz/c/c89/c89-draft.html#2.2.4.2
[0]: https://port70.net/%7Ensz/c/c89/c89-draft.html#3.3.3.4
[1]: https://port70.net/%7Ensz/c/c23/n3220.html#6.5.4.4
So that scaling factor is placed into CHAR_BIT, and would be 64 on that DSP compiler.
> An object shall have its stored value accessed only by an lvalue that has one of the following types:
Yes, my confusion comes down to my confusion of what "byte" means in the C spec.
Thank you for your time in pointing this out.
Bit of a (not so?) fun fact: char* being a universal alias can lead to some potentially unexpected slowdowns [0], especially if the char* bit is behind a typedef.
[0]: https://travisdowns.github.io/blog/2019/08/26/vector-inc.htm...
It’s not any of that. It’s because of overcommit being the default for basically every Linux system. With that, malloc will never fail, and it’s the later access of that memory that will. In practice, you’ll virtually never see malloc actually return a failure, and so most software, no matter the language, is generally not robust to this condition.
$ uname -a
Linux boxcar 7.0.0-31-generic #31~24.04.1-Ubuntu SMP PREEMPT_DYNAMIC Mon Aug 10 09:38:02 UTC 2 x86_64 x86_64 x86_64 GNU/Linux
$ cat tmp.c
#include <stdio.h>
#include <stdlib.h>
int main() {
char *s = malloc(50000000000ULL);
if (s == NULL) {
printf("Boo, hoo!\n");
} else {
printf("Look at all that memory!\n");
}
return 0;
}
$ cc tmp.c
$ ./a.out
Boo, hoo!
As to the lack of robustness of most software, that's a fact. But, for example, Daniel Stenberg of curl fame is a developer of robust software who does not appreciate how Rust handles out-of-memory errors makes it impossible to implement libcurl in the way he expects a library to work. See https://www.youtube.com/watch?v=HFH2vZRTKrA&t=2080s from 4.5 years ago as an example.To get back to the essay, I can easily understand why someone who has Rust as their first-and-only systems language, and therefore expects abort-when-out-of-memory, will not immediately consider how C allows a different approach to how to handle that condition, but instead will try to replicate Rust's behavior in C.
"I can program FORTRAN in any language." :)
I can't speak to the specific implementations, however servicing an allocation request has multiple fallible steps: reserve address space, then confirm/acquire/defer the physical memory backing the allocation.
My understanding is that overcommit allows deferring assigning physical memory to the allocation, however it could still fail to find a chunk of address space to service the allocation request.
For context, steveklabnik suggested that malloc will virtually never fail on basically every Linux system, because of overcommit being the default.
I showed a trivial example of malloc failing, when trying to allocate more space than on my machine.
Another example is when using resource limits. I've modified my code to malloc only 500M bytes, the set a virtual limit of 500000KiB, which works, then 40000KiB, which doesn't
$ grep malloc tmp.c
char *s = malloc(500000000ULL);
$ cc tmp.c
$ ulimit -S -v 500000
$ ./a.out
Look at all that memory!
$ ulimit -S -v 400000
$ ./a.out
Boo, hoo!
I clearly disagree with steveklabnik's because it's easy to demonstrate cases where malloc fails on Linux-based machines."It's hard to handle correctly for beginners" is a reason for beginners to ignore a return value. It's not a reason for the system to (by default) not do the bookkeeping to know this will (likely) fail.
Sparse mapping & reserving address space then deliberately faulting in pages is a good & advanced technique, but it shouldn't share interface with the beginner/default API.
(I say this as an aspiring C major pianist.)
... but that ship has already sailed
It's becoming increasingly common. Rust is my first systems programming language too. I tried to learn C a few years ago but I found it too austere and prickly, which put me off.
Rust is taking steps or has been taking steps in this direction too.
It's not surprising that a lot of experienced systems and desktop development engineers looking to getting their hands on the newest tool already have experience or at least familiarity with C and C++.
For someone getting a CS degree, that just seems so, so odd. Along with that, he's been studying edge detection in images. In his second year. Not to run off topic but, again, I find that so, so odd.
Very reasonable imo as universitys educated people that get positions with responsibilitys.
[obligatory] The original "C for ... programmers" article that convinced the world to make C as pervasive as it is today is K&R's "The C Programming Language".
I know it's in book form (egad!) and if you download a PDF, it's like 80 pages (double egad!). But it makes for surprisingly light reading.
Fluent C: Principles, Practices and Patterns by Christopher Preschern - https://www.oreilly.com/library/view/fluent-c/9781492097273/
Nowadays I actually believe I'll likely see a world without memory corruption within my lifetime.
Yeah I was thinking of prefixing my "memory corruption" with "software-bug induced"! I don't see a credible solution to Rowhammer. ("DDR[n+1] fixes it" - lol)
> an unsafe block, an FFI boundary
Honestly these feel solvable to me at this point! I think we'll see:
- unsafe code shrink as languages get more powerful
- amount of analysis we can apply to each unsafe line shoot up exponentially as AI gets cheaper
- amount of FFI we actually need shrink as it gets easier to just click "rewrite it in $lang" on the decision card when your coding agent says "I found a library for that but it's in a different language"
(Having said all of that, people seem to be adopting Zig for some bizarre reason... So maybe I'm naive to expect unsafe lines to shrink)
(But also, maybe AI gets so good and so cheap that we can just type "go fidn all the bugs andfi xthenm" into an LLM, between sips of a Piña Colada)
The biggest reason I use Zig is it's a very explicit language. The creators made a very intentional decision to avoid too many "high level" designs. This doesn't mean there's no capabilities for abstraction (comptime is great for that), but when you see array indexing, you can think "ptr + index * size with bounds check". There's lots of other things like that where the language does exactly one thing, and that thing is a low level operation.
This is terrible when you want to create high level abstractions that hide details from the programmer, but it's what I need when doing realtime audio synthesis or what I'm doing now which is writing an interpreter. I know exactly what allocates, I know what calls IO and can block (both operations explicitly take in an allocator or IO parameter), no data structures have private fields so I can always poke around at the insides. I know what types of errors a function returns, and creating errors is cheap with Zig's error union design.
So I don't use Zig because I think it's the safer language, I know it has sharp edges because of the number of times I've caused a panic on a poisoned pointer or use-after-free. But because it gives me such a transparent view into what is happening I find it liberating.
I don't see rust a "memory safe" language, or a niche one. I have complaints about it etc, but it's overall a fair baseline of reasonable decisions. When I look at C or other languages rust has learned from, I have more "Yikes, that's rough" takes. So... rust as the language of least "fucked up", to use your phrase? Ownership/safety are one part of the picture, but not what defines it for me.
The real value of Rust is likely pattern matching and tagged unions. Apart from modules/packages that are not a disaster, like Gabriel Dos Reis the saboteur's disaster with modules in C++.
There are two ways of constructing a software design: One way is to make it so simple that there are obviously no deficiencies and the other way is to make it so complicated that there are no obvious deficiencies.
— C.A.R. Hoare, The 1980 ACM Turing Award LectureFor what it's worth, in that specific instance there's no memory safety issue since Rust is guaranteed to crash on stack overflow on Ubuntu (and probably other supported Linux distros)
The first is that it is "C is fucked up" instead of the poor defaults of the C implementation your are using.
The second one is the idea that C all has to be fragile low-level pointer fiddling instead of much safer high-level code build around abstractions, which good C would usually have.
I do actually think that with modern C++ it's possible to get to a style of programming where memory unsafety is mostly a theoretical concern but not with C.
For me, memory safety is also mostly a theoretical concern in C. This is achieved by having proper abstractions instead of low-level pointer fiddling and by having a clear strategy for managing lifetimes.
It's also a shame the article uses the self-delusional C++ style of pointer declarators.
Otherwise pretty okay.
On a typical PC that doesn't seem like a problem, and it will (at least kinda) work which might give you the false impression it's required to work, which it very much is not in Rust. On CHERI it's obvious why this can't work. CHERI's pointers are 128-bit. Rust does have 128-bit integers, but Rust's isize and usize on CHERI will be 64 bits. Because only half of CHERI's pointer bits are address bits, and Rust told you that isize and usize were big enough for the address not the whole pointer.
Many clever pointer tricks only want to fiddle with the address. For example hiding bit flags in an aligned pointer works, as does hiding the entire value inline in today's enormous pointers (64 bits! Luxury) and using a single bit to mark "not a real pointer". In Rust we do these with the actual raw pointer types, they have methods like any other type, but in C or C++ you need to convert to a pointer-sized integer and then do tricks with the integer or you will write UB.
As I said, these types are the same width as an address on the target but a CHERI pointer isn't just an address, that's why they are so wide.
Maybe start here: https://doc.rust-lang.org/std/ptr/index.html#strict-provenan...
The strict provenance APIs or the "Rust has provenance" RFC are not relevant here (more precisely, they help clarify the options but do not solve the problem).
The problematic statement is still in The Reference (https://doc.rust-lang.org/reference/type-layout.html):
> Pointers to sized types have the same size and alignment as `usize`.
Which just cannot be guaranteed on CHERI with 64-bit `usize`.
There are discussions like there were before, and it's pretty clear that we'll have to give up either performance or possibly quite a lot of compatibility, but no decision.
When I was building my kernel in my late teens, I first wrote the bootloader by hand on paper in assembly language. I then referenced the x86 manual for the instruction set and converted my assembly code into the equivalent hexadecimal machine code values of the x86 machine instructions, which I also wrote by hand on paper. Then I used a hex editor on the desktop to manually write the hex values into a file and used it as the bootloader i.e as the first 512 bytes.
The whole exercise gave me a sense of hard grounded zero magic, raw, unfiltered experience. This is an experience that is hard to replicate in any other way. I deliberately did that so as to peel away as much magic/abstraction layers as I possibly can.
Later on when I started using C, I never had to learn C but merely just had to reference the equivalents of the assembly language. Like how the primitive "if" doesn't exist in the hardware but is a composition of cmp and jmp instructions. Seen this way, C becomes a glorified portable syntactic sugar over assembly language.
Then higher up the ladder to C++ for object oriented problem solving while retaining the spirit of functional programming. Rust was a breath of fresh air, where correctness across a myriad of use cases was a first class primitive.
Each abstraction layer can thus be evaluated for its utility in problem solving while its underlying mechanics remain understandable down to the hardware level.
More recently, our own arcc compiler extends correctness to our architecture and not just the types.
So if you are young and have time to spare, I suggest a little bit of Assembly => C => C++ => Rust.
This essentially makes you immune to hype train bullshit.
With Rust, you are front loaded with a myriad of compiler gymnastics you need to think through.
But once you get comfortable enough, it becomes natural. You also have to get accustomed to writing code that is more verbose than C which might look ugly at first but later you start to accommodate it as the necessary cost for the utility you are handed in return by the compiler.
For example, multiple variables in Rust doesn't necessarily mean multiple memory allocated variables at runtime like in C. The rust compiler will usually keep track of and ensure multiple variables (non Copy types such as String with Move semantics) map to one memory allocated variable at runtime (in normal single threaded use cases under normal circumstances without using RC, ARC, etc.). Eg: let a = String::from("hello"); let b = a; ... Note: The example is for illustration purposes only and not always true. In summary source code variables are abstractions and may not belong to distinct runtime memory location. Yes it is true even for C. But Rust's ownership model makes that distinction aggressively visible.
There's a _lot_ of utility in things that aren't represented by a literal difference in the compiled code. It's not unusual to make a type in Rust that's just a single field; you're not making a struct to gather related data together, you're giving it a different name because the same data means something different.
It's a little like the difference between a uintptr_t and a uint16_t*. They're represented the same in the machine, the difference only exists in the source language.
Rust derives a lot of value from things that only mean something at compile time. I think you'd "unlearn" what C (commonly) teaches when you don't find a C struct with one member "weird".
Why would a Rust programmer learn C? Isn't that basically a regression?
That also reflects in that many embedded systems offer C support and do not offer Rust support.
Write, say, a boot loader, for enlightement.