What memory model should Rust use?
paulmck.livejournal.com
paulmck.livejournal.com
The purpose of the unsafe keyword is to mark code that could cause safety issues (undefined behavior) if its expectations about the surrounding code are not met. (e.g., an unchecked array access expects its index to be in bounds).
With unlimited time and resources, it should be possible to give a mathematical proof about those unsafe blocks, showing that their expectations are met and that undefined behavior will not actually occur (by for example proving that the array index supplied will always be in bounds).
But since relaxed and consume have unclear semantics, it hard to formally reason about the correctness of any code that uses them (and it is also hard to build tools to do that checking, as mentioned).
So in an ideal world, we would have a 100% clear semantics for all parts of the language, _including_ the unsafe fragment. Unsafe should not be a place to put poorly specified languages features.
Of course, we don't live in this ideal world, but I would rather hope for fixing the memory model so that it be made completely formal and precise (if necessary, giving up some currently allowed compiler/hardware optimizations if they're more trouble than they are worth). As an extreme example, sequential consistency is easy to understand as long as you don't mix it with other types of atomic accesses and avoid non-atomic data races.
That is one use of the `unsafe` keyword. You see this version in the qualifiers for unsafe functions, where safety depends on the inputs and other context. It means that the marked function cannot be called from safe code. The `unsafe` keyword is also used to mark code blocks where the content of the block must be manually proved to be safe, even for valid inputs—this includes calls to unsafe functions, dereferencing raw pointers, etc. (Unsafe functions use this mode by default.)
The same keyword is also used to mark unsafe traits and their implementations, where an incorrect implementation of the trait—even the members which are not themselves marked as `unsafe` and thus can be called from safe code—can impact the safety of other parts of the program (via unsafe code blocks or unsafe functions). An example of this would be a trait with a member which returns a raw pointer (itself a safe operation) which is then called from within an `unsafe` block and dereferenced.
Perhaps openHellPortal is complicated enough that you feel it should be a separate function. On the other hand, you want to be clear that nobody should expect openHellPortal to check the necessary precaution are in place. So, mark openHellPortal unsafe and provide a safe wrapper SafeHellPortal which internally calls openHellPortal after checking some necessary precautions are in place and ensures closeHellPortal is called when appropriate too.
As far as I know nobody builds hell portals with Rust (You have to assume there'd be a HN post about it right?) but they do plenty of embedded systems development, and those people use unsafe this way. Twiddling the SOC's interrupt controller is nothing to Rust itself, but the people writing the interrupt controller twiddling code can mark it "unsafe" and explain what callers must do to avoid chaos. Then callers can obey those instructions and make a safe wrapper for a particular purpose e.g. "SerialPort", and the average developer needn't even know that under the hood randomly twiddling the interrupt controller causes everything to catch fire, because they're using a safe SerialPort abstraction which can't cause that even if they hold it wrong.
My issue is with using "unsafe" to mean "watch out, you might be using this incorrectly, but no one even knows the exact rules that you should stick to to make sure your code is correct."
Of course, having it be in the "unsafe" part of Rust is better than having it be in safe Rust. But ideally you would be able to look at any part of a Rust program and be able to tell what it does, without having to think in terms of specific compiler optimizations or hardware details. Otherwise formal verification of concurrent algorithms and data structures becomes hopeless.
* Rust threads run on CPU cores, but also GPU cores, FPGA cores, and a wider range of nastier hardware than Linux kernel threads run on.
* The author implies throughout the text that safe Rust operations might not be allowed in unsafe Rust. This hints at a deep misunderstanding of how Rust works.
* C, C++, and similar languages, do not have as a goal to be "sound". Soundness is, however, core to Rust's value proposition.
* C, C++, and similar languages, do not have the goal that undefined behavior must always trap in their abstract machines. But Rust does.
* C++ and similar languages do not have the goal of forever-backwards-compatibility (e.g. C++ breaks backwardward compatibility on almost every new standard release). Rust, however, guarantees that all programs written for Rust 1.0 in 2015 that have no undefined behavior will build and run correctly _forever_.
* Rust has many compiler backends, e.g., LLVM-IR, GCC, Cranelift, SPIR-V, etc. Rust code is translated to the IR of those toolchains. That is, those IRs would need to be extended with the new semantics as well to be useful.
> But again, this decision rests not with me, but with the Rust communty.
The Rust community needs people working on formalizing the memory model further. So if they are interested, they should jump right in.
They should probably start by understanding Rust as is today, learning about unsafe code, and well the blog post cites https://plv.mpi-sws.org/rustbelt/rbrlx/ , but it is clear from the exposition that the author has not deeply understood that paper, so maybe they should also take a deeper look at that. All the proofs are open-source. So if they think they can extend the soundness proofs with new features that are also sound, that could be a place to start.
What do you mean ?
However to be 100% sure of what can be considered safe, it needs to be the same regardless of what hardware is being targeted.
That is how C got its UB in first place, WG14 didn't want to rule out any kind of hardware that was possible to target so UB it was.
Other memory safe systems programming languages have gone through other path and were willing to add abstractions to their runtimes instead, ensuring that the memory model remains consistent to the developers.
Naturally in languages that want to be zero cost abstraction, that solution is not welcomed.
https://doc.rust-lang.org/reference/behavior-considered-unde...
C, C++, etc. don't even have implementations of their abstract machines. They don't have one existing as a goal. And they are happy to make certain operations exhibit undefined behavior even if that implies that it would make an implementation of the abstract machine that traps impossible.
This is why even if you were to combine valgrind with address sanitizer, memory sanitizer, thread sanitizer, undefined-behavior-sanitizer, and other existing C and C++ tools, there is still a lot of classes of undefined behavior that these tools can't detect.
That's fine in C and C++, but not fine in Rust. In Rust, if we add a new type of undefined behavior, the constraint is that it should be (demonstrably) possible to extend miri to detect it, such that if a user doesn't know whether some program exhibits undefined behavior for some inputs, they can just run it under miri, and miri will precisely pinpoint which part of their code exhibited undefined behavior and why, and how their program execution got there.
> This is why even if you were to combine valgrind with address sanitizer, memory sanitizer, thread sanitizer, undefined-behavior-sanitizer, and other existing C and C++ tools, there is still a lot of classes of undefined behavior that these tools can't detect.
Do you have specific examples of such UB?
Many of them are impossible in theory to check, and many are impossible to efficiently check in practice.
Furthermore, because they're standard library constructs, they are not straightforward to change in an edition either.
1. Add new unsafe relaxed atomics features, by whatever means. Being new there's no backward compatibility worry.
2. Modify the existing safe atomics to treat Relaxed as Acquire-release. This doesn't break anything since Acquire-release is necessarily not breaking Relaxed guarantees, although of course some code now has worse performance than before.
Edited to add: Also, what do you mean you can use Consume today in Rust? Where? My std::sync::atomic::Ordering doesn't even list Consume as an option.
How do you use the consume ordering in Rust today?
Not only "std library" (which is only available when an OS is used), but also "core library" constructs. That one is available completely independent of the OS. Which means also in the current Rust for Linux project, or any other bare-metal platform.
"Inspired by the apparent success of Java's new memory model, many of the same people set out to define a similar memory model for C++, eventually adopted in C++11." - https://research.swtch.com/plmm
"Rust's growing connection to the deep embedded world rules out the memory models of dynamic languages such as Java and Javascript."
So what is Rust using now, C++11 or C++20? "memory model for atomics from C++20"?!
The embedded community in this context is people who were able to:
#include <embedded_vendor/mmio_stuff.h>
... a file full of dodgy C pre-processor macros into their C++ project and are unhappy that J. F. Bastien's volatile deprecation means now their compiler points out what a terrible idea some of those macros actually are.
Having the compiler shut up and silently allow this might make them feel better, but the situation they're in isn't actually improved by such a reversion at all.
Removing the batteries and disabling the smoke alarm does mean you get a peaceful night's sleep, but sometimes you don't wake up after that peaceful night's sleep because meanwhile your home burned down and there was no alarm.
The claim made in the proposal paper is that every C++ programmer who is including such macros definitely ensures by construction that their use of the macros is actually safe. You've probably read a bunch of real world C++ code written by people who are very confident in their ability to write correct programs, how do you feel about the likelihood that all these programs are in fact correct by construction despite using the unsafe macros?
It seems that in later revisions of the proposal, they considerably reduced the scope of the un-deprecation, realising that even if you earnestly believe (interrupt_register |= ENABLE_IRQC) isn't a footgun, it's much harder to rationalise (clock_rate /= 3) and either way nobody wanted half the random places C++ allowed you to write "volatile" anyway, so deprecating those isn't actually a problem and "reverting" the deprecation just highlights yet another mysterious black box in C++ with no clear purpose.
Rust simply omits Consume altogether. It can always be added if there's consensus that we now understand what's going on here and can meaningfully deliver it on hardware people actually have in the real world.
"While I'm on the topic of concurrency I should mention my far too brief chat with Doug Lea. He commented that multi-threaded Java these days far outperforms C, due to the memory management and a garbage collector. If I recall correctly he said "only 12 times faster than C means you haven't started optimizing"." - Martin Fowler https://martinfowler.com/bliki/OOPSLA2005.html
I get the feeling this quote is possible due to the memory management rewrite of the entire JVM in version 1.5 to allow Doug to make the concurrency package?
But I cannot explain it, even if I have experienced it!
The Java program isn't doing anything that is literally impossible in C, but powerful abstractions can make it exponentially easier for everybody to do it in Java.
I can just about keep in my head how Quick Sort works. So, I figure I could write one, from scratch, in maybe a day or two in a language like C and have it work. But obviously I shouldn't do that, your standard C library comes with a perfectly nice fast sort. However, Java is supplied with an array of equally good features, and C is quite sparse on that front. You could program every one of those features in C, but, you probably didn't.
Memory management is an example like that, but spread through most programs. If you're meticulously allocating and freeing things in C, that takes time. Is it worth it? Maybe not but can you keep track well enough to do anything faster without just leaking unlimited memory? Java comes with a good Garbage Collector which can avoid all that explicit tracking you need to do, yet free all the same memory. Clever. You can do this in C, some people do, but it's hard, you probably don't do it.
I think the real technical reason, like you allude to, is closer to this quote:
"Many lock-free structures offer atomic-free read paths, notably concurrent containers in garbage collected languages, such as ConcurrentHashMap in Java. Languages without garbage collection have fewer straightforward options, mostly because safe memory reclamation is a hard problem..." - Travis Downs https://travisdowns.github.io/blog/2020/07/06/concurrency-co...
But I have yet to meet someone (4 years into this realization) that can clearly explain why in a way that anyone could understand and more importantly accept!
Importantly, the x86 ISA provides much stronger guarantees than C/C++. AFAIK the only memory reordering allowed is that a write followed by a read can be swapped around. This is extremely strict (we can’t even reorder adjacent reads) and literally any other processor architecture will provide much weaker guarantees. And because C and C++ formally defined a weak memory model in 2011, modern compilers are free (even when targeting x86 chips) to perform optimizations that reorder memory accesses in ways programmers familiar with x86’s semantics might not expect.
So an algorithm that worked on x86 processors 10-15 years ago may no longer be correct when compiled with a modern optimizing compiler, or when running on ARM. (If in doubt, you should be OK replacing anything volatile with an atomic and specifying sequentially-consistent ordering semantics for all operations — which you can do in Rust!)
[0]: https://systemtbe.blogspot.com/2014/05/volatile-memory-barri...
Java volatile is different of course.
Although, having to go all-volatile is maybe too much of a constraint, when many algorithms (like lock-free queues etc.) only require like 3 non-reordered accesses (2 to begin, 1 to end?) to safeguard a transaction, and then the remaining reads and writes can be reordered arbitrarily (inside the safe bounds).
No, you really won't, neither in C nor in Rust. There are (at least) two reasons for that:
1) You can't do atomic read-modify-write operations (like fetch_add) with volatile.
2) "volatile" is meant for the compiler, not for the processor. Volatile accesses are compiled down to regular read and write machine instructions, so you won't get any ordering guarantees for code running concurrently on different cores.
The reason "volatile" is useful for memory-mapped IO is that such memory is mapped with different flags for cacheability than regular memory. For example, given the right flags, the processor won't cache stores in a store-buffer, or won't coalesce smaller stores into larger stores before sending them up the memory hierarchy.
Are those operations so important? All I've ever done to be honest is implement some lock-free SRSW queues, and those can be easily implemented with just acquire/release semantics. No atomic read-modify-write needed. Could it be that those are only needed when more than 2 threads are involved?
Yes, absolutely. You can't even implement thread-safe reference counting (std::sync::Arc in Rust) without atomic RMW – or even a thread-safe counter.
A lock-free single-produce-single-consumer (SPSC) queue is pretty simple. As soon as you have more than one producer or more than one consumer, you need atomic RMW.
On x86 if everything that can be potentially shared is volatile, you may get away with it if you only need acquire/release, but in practice not everything being shared are not going to be volatile. For example if you implement a spsc queue only with volatile operations and use it to send a pointer to a non volatile object to another thread, the consumer might see a non initialized object.
It's not "sufficient" for MMIO, it's literally the sole use case for it, and the sole situation where it's correct, because the hardware will not fuck with MMIO accesses either (and not eliding reads is important because reads might have side effects). Although even then you may need to use atomics anyway if the driver is concurrent, because volatiles don't guarantee any sort of atomicity and the reaction of MMIO devices to concurrent access is very likely to be funky (then again these days MMIO is usually a thing on uCs where there is no uncontrolled and arbitrary concurrency).
> Although, having to go all-volatile is maybe too much of a constraint, when many algorithms (like lock-free queues etc.) only require like 3 non-reordered accesses (2 to begin, 1 to end?) to safeguard a transaction, and then the remaining reads and writes can be reordered arbitrarily (inside the safe bounds).
Except your accesses will be reordered arbitrarily:
* volatile instructions are only of concern for the compiler, everything below it has no idea about them, and if you're not in an MMIO context they'll absolutely try to reorder and remove them (and even when they won't since volatiles are not synchronisation point you have no idea how they'll propagate)
* volatile instructions are not fences, even for the compiler are no "safe bounds" with volatiles unless all the operations are volatile
That's exactly what I said right? But good to know that volatile won't emit fences.
That “(inside the safe bounds)” disclaimer is the reason we can’t use volatile for anything meaningful, because volatile does not provide that guarantee. The compiler & hardware will happily let the contents of your critical section float right past the volatile reads/writes “protecting” it.
If I write a lock-free queue and use volatile for synchronization, then I’m just protecting the queue itself and not the contents. For instance, if I write to some memory and then push a reference to it onto the queue, other cores might see the reference but not the data. Maybe the volatile writes got rushed through the memory hierarchy & to the other cores, but all the non-volatile writes are still stuck in a cache somewhere until next week.
And AFAIK volatile still doesn’t guarantee atomicity in the case of concurrent access. If one thread writes to a volatile variable, other threads can easily see torn writes/“pink elephants”/out-of-nowhere values. Again this is somewhere that x86 provides unnecessarily and somewhat ridiculously strong guarantees that nobody else does, so code that worked on dumb compilers in the 90’s may be very, very, broken today.
So please just use atomics. They’re very nice, and they solve all of these problems. If you don’t understand memory ordering semantics just mark everything as sequentially consistent and don’t think about it, I don’t care, just please use atomics.
The reads can always be reordered, unless they access uncacheable memory, which can happen only in device drivers, not in normal applications.
x86 provides 3 memory fences (LFENCE, SFENCE, MFENCE) which must be used when you need a certain ordering of the reads or of the reads around writes.
On x86, reads can be reordered between themselves only for non-temporal transfers.
For the normal case, when the accesses are made to cacheable read-write memory, the reads are not reordered between themselves.
Meaning they’re completely different and way incompatible with Java’s.
But the semantics of "volatile" are completely different in C++ and Java! And I've seen lots of bugs from people assuming Java semantics for volatile in C++.
In C++ land, volatile is mostly useless and sometimes harmful (see [0]). It's dangerous to use it in concurrent code and hardly ever useful for memory mapped i/o without appropriate memory barriers (which make the volatile qualifier redundant).
In Java land, volatile can actually provide some useful guarantees when it comes to multithreaded uses but afaik they will usually just be turned into atomic load/store instructions so you might as well be writing atomics explicitly.
And it can be debated whether "volatile" should be a type qualifier in the first place. I would argue that it's better to make load/store operations on volatile memory explicit (for those very rare cases when you need it).
There is experimental support in Rust for the latter, `volatile_load` and `volatile_store` in `std::intrinsics` [1].
But still, doing volatile loads and stores in multithreaded code is hardly ever the right thing to do, at least without accompanying them with the relevant memory barriers (which are especially important on ARM CPUs).
[0] https://github.com/torvalds/linux/blob/master/Documentation/... [1] https://doc.rust-lang.org/std/intrinsics/fn.volatile_store.h...
It's not really experimental. You've linked the underlying core intrinsic, but the stabilised interface, which is what programmers should actually use, is std::ptr::write_volatile / std::ptr::read_volatile
Essentially you linked an implementation detail and that's why it's labelled unstable and you can't use it in stable Rust.
But yeah, you can have your volatile in Rust then. It's not a type level annotation, but an explicit (unsafe) operation.
That's not entirely true e.g. `read;read` can be very relevant in an MMIO context (a read can have side-effect), but atomic reads can technically be optimised out (though it might not currently be done by most compilers).
The memory model should explicitly forbid this optimization, especially if you use an explicit memory ordering with your loads and stores. I'm not 100% sure what the default (sequential consistency) does in this case.
And if you do `read(); barrier(); read(); barrier();` like you should on most CPU architectures, the compiler definitely is not allowed to reorder the reads (barriers affect the compiler AND the CPU). Which barrier you need depends on the MMIO semantics and CPU architecture you're on. Adding volatile does not change this but may inhibit other, useful optimization.
In your example of `volatile_read(); volatile_read()`, the compiler is not allowed to optimize or reorder the loads but the CPU is still allowed to reorder them with each other or other instructions. Adding `volatile` does not make this example work.
It has no reason to. Reading the same location twice, even atomically, has no “reason” to change the value so reusing the same read is perfectly valid.
If there is an other operation between the two and that has a sequencing implication then yes, otherwise no.
> In your example of `volatile_read(); volatile_read()`, the compiler is not allowed to optimize or reorder the loads but the CPU is still allowed to reorder them with each other
If there are memory reads then sure but i’m assuming actual mmio here (since volatile is otherwise useless) and i’d think an mmio implementation which reorders mmio reads is completely broken.
No, this is definitely not how atomics with memory orderings work. If you do two atomic reads, the CPU will do two load instructions. Neither the compiler or the CPU is allowed to optimize this.
> and i’d think an mmio implementation which reorders mmio reads is completely broken.
This is exactly how ARM SoCs work. The CPU core does not know if a load address is mmio or not and may reorder the loads as it sees fit. It's up to the memory subsystem to deal with this, the CPU core does not care.
I've done lots of mmio code on ARM at $work in the past few years, and we've spent a lot of time to make sure we have all our memory barriers right. They are quite tricky on ARMv8, as mmio can be either on the system bus or the main memory bus which need distinct kinds of memory barriers.
Using volatile in mmio code on modern CPUs is almost certainly a bug. Some microcontrollers may be an exception.
Of course it is.
> If you do two atomic reads, the CPU will do two load instructions. Neither the compiler or the CPU is allowed to optimize this.
If you perform two atomic loads on the same location in sequence, then it is completely feasible and normal that the two loads would return exactly the same value, thus the second load can be optimised away. Even under sequential consistency, this is perfectly, well, consistent. It's also perfectly valid to fold atomic writes. See http://www.open-std.org/jtc1/sc22/wg21/docs/papers/2015/n445... for more discussion on the subject
No compiler currently bothers doing that (that I know), but it's not because they can't, it's because the effort is probably not worth it, at least at the moment.
> I've done lots of mmio code on ARM at $work in the past few years, and we've spent a lot of time to make sure we have all our memory barriers right. They are quite tricky on ARMv8, as mmio can be either on the system bus or the main memory bus which need distinct kinds of memory barriers.
I can believe that.
> Using volatile in mmio code on modern CPUs is almost certainly a bug. Some microcontrollers may be an exception.
Not using volatile in mmio code on modern compilers is almost certainly a bug as well.
Yes they are. And other optimizations are allowed too. If you do an 32-bit atomic load and then mask the result with `& 0xff`, the compiler can theoretically narrow it into an 8-bit atomic load, as long as the architecture guarantees that such a load still provides the necessary ordering guarantees – which I think is true on most architectures. This won't necessarily work for MMIO, where incorrectly-sized accesses may trap or return the wrong data. But the compiler is not responsible for preserving MMIO semantics if you only asked for atomic.
That said, I'm not aware of any major compilers performing /any/ significant optimizations on atomics, so you won't get in trouble for this in practice.
> This is exactly how ARM SoCs work. The CPU core does not know if a load address is mmio or not and may reorder the loads as it sees fit. It's up to the memory subsystem to deal with this, the CPU core does not care.
You can configure in memory attributes whether to allow reordering or not.
> They are quite tricky on ARMv8, as mmio can be either on the system bus or the main memory bus which need distinct kinds of memory barriers.
Which SoC?
Java's "volatile" is sequential consistency on steroids, a promise that changes to a volatile variable "happen-before" anything else that thread does, and "happen-after" anything the thread already did. Rust's std::sync::atomic has a hardware fence which you could use to deliver this same outcome (throw away performance, gain sequential consistency).
For concurrency C++ volatile does not supply any semantics, except on MSVC. If you're looking at papers or reference implementations which assume C++ volatile has semantics that are relevant to concurrency you're probably looking at papers that only needed Acquire-release and it worked on the x86 CPU they tried it with, or they forgot to mention it only works on MSVC (which promises Acquire-release even on platforms other than x86).
I mean, the box says: https://doc.rust-lang.org/stable/std/ptr/fn.read_volatile.ht...
> Rust does not currently have a rigorously and formally defined memory model, so the precise semantics of what “volatile” means here is subject to change over time. That being said, the semantics will almost always end up pretty similar to C11’s definition of volatile.
Which isn't exactly "exact", you know?
If so I feel that risk is small and continues to shrink.
The reason people ask for this is because they haven't thought this trough.
Even if the compiler "preserves" all volatile operations, the HW does not.
A trivial example is WASM. Rust can generate WASM code that, when using volatile, has many duplicated loads and stores in some order. Then the browser runs a WASM optimizer on the WASM code, which discovers that these instructions are redundant, and there you go, your program now doesn't do what you wanted anymore (formally, it never did).
x86, ARM, PPC, FPGAs, GPUs, etc. all do this in HW as well. So while you could argue that, e.g., WASM should have volatile loads and stores, the same argument can be made at the next level down below, down to hardware.
You'd need to ask every hardware vendor for some sort of serializing loads and stores, and people have done that, and vendors have implemented that, but then when users finally got what they wanted, they complained that they are too slow, so at that point they just stop using volatile and use atomics which is what they should have done in the first place.
Since now I know that you knew that the feature you were requesting is impossible to implement, and you did not mention that in your feature request, now... i have no idea of what the goal of your comment was.
Since when is Java a dynamic language?
Object being a valid write destination for any reference means the VM has to maintain type information in all objects (small 'o'), which is a distinctively dynamic thing to do. I mean, simply declare all variables to be Object and you can see a faint shadow of a python-like class system lurking underneath the static types. The fact that nobody does this (I hope!) is of no use to the VM, the semantics say it must do this.
All object allocations are on the heap (and only ever moved to stack if escape analysis proves they could), all methods are dynamically dispatched. Hell, even class loading (i.e. linking libraries) is done at startup, which means you can recompile libraries without re-linking the client code which uses them.
Sure C++ offers you both stack and heap allocations, virtual and ordinary methods, static types and std::any for values of unknown/irrelevant runtime type. But the default zero-effort thing to do is to use statically-typed stack-allocated objects with non-virtual methods. Java only offers two of these things and one of them needs to be explicitly stated upfront. And, furthermore, Java's runtime still maintains the full machinery needed to do dynamic typing and dynamic method dispatch (the final modifier on methods is purely a promise to javac, the JVM has no concept of a method call that doesn't involve looking up the class it refers to and potentially all of its superclasses), it doesn't matter that you don't use it, your calling code or called code might, and the specification says that it can.
C++ objects compile to dumb bit buckets that contain exactly what's asked of them and nothing more. No matter what you do, you can never have all class types in C++ share a common ancestor, and the ones you can will know nothing about their parents unless specifically asked to with 'virtual', they can only remember this when allocated on the heap.
I don't get the snark, these things are only allowed for virtual-method-containing classes, something I explicitly mentioned and said it doesn't count because it's not on by default, not idiomatic to be on by default, doesn't exist in binaries or at runtime when not in the source, and is extremly limited compared to the rich metadata every java object hold no matter what you do or say (so even objects of final classes have it).
>loading C++ classes via shared objects
A capability of the underlying operating system, which _is_ a "runtime", just not a standard or required part of the language, and shared objects or dynamic libraries know a tiny drop of water about their source compared to the ocean of metadata in every .class file and the runtime representation of them.
>C++ language runtime
What does that even mean in this context ? I have heard it said to refer to the standard library, which is ridiculous in our context because it's neither required (let alone mandatory) part of the language nor carries information or controls execution of the calling code like the java runtime does. I have heard it being said to refer the variant of C++ that Microsoft runs on the CLR, which might as well be C# in this context. The closest thing would be the Exception Handling Mechanism and the runtime info it requires, but that's again extremely limited compared to what java does and a qualitatively different thing. Dumb tables of pointers or return addresses can never be comparable to the very rich representation of types and other metadata about code that the JVM or the CLR does.
Secondly a C++ runtime is responsible for running all global variable contructors and destructors, handling exceptions, soft floating operations, loading shared libraries on demand for dynamic linking (depending on OS capabilities or bare metal), how new/delete translation into the underlying OS/MMU.
As for the metadata, so when I use static reflection in C++, either via template metaprogramming, or with the features that sadly will only come in C++26, then what?
And really I just picked C++ as first example that came into my mind.
We can also start discussing Eiffel, Delphi, Swift or Go for example, with their rich type information, dynamic dispatch, rich runtimes, and native compilation model, so they are also dynamic?
Ok? so This point contradicts you, the language defines no runtime support, so it's very difficult to do the things that a language with a runtime in its specification considers for-granted.
>a C++ runtime is responsible for running all global variable contructors and destructors, handling exceptions, soft floating operations, loading shared libraries on demand for dynamic linking (depending on OS capabilities or bare metal), how new/delete translation into the underlying OS/MMU.
This is a misuse of the term as used in our context, almost all of the above can be considered as simply linked/injected code. The global variable initialization is done by startup code injected by the linker, the construction and destruction code is routine boilerplate inserted by the compiler at specific fixed locations, the new/delete operators simply wrap calls to the underlying memory allocators to call constructors and destructors properly, etc... Nothing of these features knows anything of value about the source code it's injected into, they are just boilerplate function calls inserted for housekeeping. Only loading dynamic libraries can be considered as remotely "runtimey", and I have never heard of them done on anything without an OS.
>so when I use static reflection in C++, either via template metaprogramming,
This is compile-time reflection, it's neat but has nothing to do with dynamacity or runtime support. The original point was 'C++ is drastically less dynamic than Java because it's binaries are very thin and it's philosophy is that the compiler shaves as much of the high-level semantics as it can and leaves nothing but raw, dumb, optimized bits'. I don't see how compile-time metaprogramming contradicts this.
>or with the features that sadly will only come in C++26,
What are those features exactly? do they include, for example, that the compiler injects the name, fields and method signatures of every class into an inspectable object? If yes, then those features, when they come, _will_ make C++ significantly more dynamic. (and its binaries significantly fatter)
>really I just picked C++ as first example that came into my mind
A very bad example to say the least.
>We can also start discussing Eiffel, Delphi, Swift or Go
Let's do it indeed, I love discussing programming languages! What exact features or abstractions of each of those language that you think will classify it as dynamic? If they are similar in quantity or quality to what the JVM does it should indeed qualify as more dynamic than C/C++ (almost everything is). Dynamic-Static is a spectrum, not a binary category.
More correctely : C/C++ + Windows, because Windows is the runtime implementing the dynamicity, the language standards alone say nothing about generating executables or operating on them, for all they care you can generate .class files from them and obtain the same benfits of java classes, it still would be the java runtime doing all the work. Java, in contrast, defines the executable generation and runtime loading to be part of the language semantics, anything you write that can't do what they say is by defintion not a "java runtime". It's like the difference between the process abstraction in Erlang and the process abstraction in linux-hosted C.
It only matters if your program has concurrency (more than one apparently simultaneous execution context, today often because there is physically more than one CPU core) and the concurrent elements of the program are able to read or write the same memory. Notice that at the speeds CPUs run, the actual main "memory" of the computer may take a relatively long time to read from or write to and so we use multiple layers of caches further complicating the problem.
We want: A model that's fast yet also a model that's understandable for merely human programmers. This turns out to be a difficult problem.
It turns out to be much more complicated than "the data is just written", because there are other threads that could be reading and writing at the same time, multiple layers of inconsistent caches, and clever compiler optimizations or hardware that can change execution order of the code.
Much of its purported virtue seems to be mostly lore and myth, and doesn't keep its promises under close inspection and scrutiny.
For example the "Rustonomicon" is only half finished and doesn't have that much real insight that couldn't be gained from a lot of basic-ish rust tutorials.
The memory model is not propery defined and only kinda sorta borrowed from C++ as the article mentions.
The borrow checker has only been formally specified and veryfied in a much simplified form and only very recently [2] ,with the actual model "just being what the borrow checker implementation does".
Having a fancy datalog based borrow checker is cool, but I'd much rather have a formal specifiation of the semantics.
Rust took the big bang approach of building everything at once, implementing stuff first without any formal semantics and just floating on an iceberg of implementation definedness, with endless discussions on how things COULD be formalized.
That's pretty much the antithesis to rusts safety philosophy; it's teaching water but drinking barrels of wine. It wants your code to be well defined and follow its rules, but its rules are completely arbitrary, not properly defined nor written down.
I wish instead they had just started with a small borrow checked core language, and grown it from there, like described in Guy Steeles talk on "growing a language" [1].
I feel that Zig is doing a much better job in this regard, they have a tiny well thoughought language with a rigorous semantic description, and a good reference document.
I'm sure it would be much easier to add a borrow checker to the well-defined-ness and correctness of Zig, than to make Rust well defined and correct.
Super friendly community though. Who knows, maybe someday the language that people think rust is, will crystalize from that community through their collective effort, like a self fulfilling prophecy.
While Rust is well thought out it’s killer feature is it’s the somewhat unique ability to allow people to get things done relatively quickly (sometimes/often with the expressiveness you’d expect from Ruby or Python) whilst also offering gigantic correctness guarantees and a (initially) low barrier to entry.
But hey, Zig has a reference document. Watch out World!
Perl was successful for a time but its messiness eventually caught up to it. OCaml is of the same vintage and grew much more slowly but I dare say it has a brighter future.
The language was grown. It just wasn't grown with an initial central feature, it was grown with an initial problem that it was trying to solve: a concurrent, safe systems language. As Rust explored the problem space, tradeoffs were identified, some were weighed differently than previously, some new strategies for solving these problems were discovered, and things were refined.
Verification is actively underway. It's totally understandable to want to wait until it's here, but that doesn't mean it's impossible, just nontrivial.
(Also, there are more papers that are older, like https://people.mpi-sws.org/~dreyer/papers/rustbelt/paper.pdf from 2018, which this paper cites.)
Normally I'm a huge fan of a philosophy of backwards compatibility. Clojure has shown that it is possible with enough dedication. But I feel like it's a liability with the lineage you describe for Rust, the problem space was just too unexplored and vast, and the language has grown too large.
Joe Armstrong has a joke in one of his keynotes [1], that after you've written the first version of something you should throw it away completely, because you didn't understand the problem at all when you wrote your solution, and now you understand it a bit better.
A bit like how the web would be a much less ghastly ecosystem if we hadn't abandoned versioning for living documents which force every behaviour ever to interact, and god forbid be compatible, with one another, in a combinatorial hairball.
Modern browsers would be less unwieldy if they just shipped 30 different rendering engines, one for each major revision, than ship a blob that purportedly is able to still render everyones Dogs website from the 90s (which is not true of course, it's just that nobody from the 90s is here to complain).
It's a bit ironic that Rust fell into the same trap that caused the complexity that inspired its inception in the first place.
Maybe once the major "features" for safe system programming were identified a new language (revision) should have started with only those features at its core. I fear that if such a language were to crop up now, Rust would just eat its children.
We didn't have backwards compatibility in those days. Radical change happened. I personally think of Rust as being three or four different languages before 1.0 happened. The problem is the exact same as others have said; doing that degree of breaking change all at once is effectively starting a new language project, which means that you also throw away all of the community and things that they've built so far.
Also, Rust does have a kind of "core" semantics, that is, MIR. You can't code in it directly, but it is a rich target for verification and analysis tools.
The sad truth is probably, that Rusts complexity is just a mirror of the entire industry.
If we, as a profession, placed a lot of value onto tiny, simple, yet well thought out (and therefore powerful) things, rewriting everything once a year wouldn't be as much of a problem.
No codebase would have more than 5kLOCs, Browsers would only do layouts with the thing we build after learning from Flexbox, Grid, and Column, and Rusts language spec would fit into a white-paper like Scheme did.
At least Rusts borrow checker uses such a tiny, well thought out, piece of code with its datafrog engine.
[Anybody stumbling upon this comment, and wondering what Alan Kay is/was up to, this might be a good philosophical starter: https://www.youtube.com/watch?v=FvmTSpJU-Xc]
Zig doesn't have a "rigorous semantic description" that I can find. The reference manual is nowhere close. A "rigorous semantic description" would either be something like that of Java or, ideally, the operational semantics of Standard ML.
I glanced through the "undefined behavior" section in the Zig reference and didn't see use-after-free anywhere. If it's that easy to find missing pieces then it isn't even close to complete. Wait until you start trying to define things like pointer provenance; it's not even clear what semantics LLVM is upholding.
> Having a fancy datalog based borrow checker is cool, but I'd much rather have a formal specifiation of the semantics.
There's a formal specification of the semantics of a subset of Rust via the RustBelt project. These things almost always start with a subset of the language; very few languages have full formal specifications, including type safety proofs and so forth. I should note that I'm not aware of any such work being done for Zig.
> The memory model is not propery defined and only kinda sorta borrowed from C++ as the article mentions.
Where's Zig's memory model? I didn't find it. I did find a GitHub issue [1], which has had far less discussion than that surrounding Rust's memory model.
> I'm sure it would be much easier to add a borrow checker to the well-defined-ness and correctness of Zig, than to make Rust well defined and correct.
No. This is just like the "just add a borrow checker to C++" meme that we used to see until it was shown to be infeasible. The problem is simple: if you can have mutable aliasing today in Zig, which you can, then the ecosystem has surely come to depend on it. So by forbidding mutable aliasing, you will break tons of code. It is not a reasonable subset of code to break; in C++ the set of code you break is "all methods" (because "this" has unrestricted aliasing), and I expect Zig to be no different. The borrow checker is not some magic dust that you sprinkle on a language to make it memory safe; the entire language, including all libraries, must be designed around it. Are you really going to break the doubly linked lists everyone has written?
It is impossible to overstate how much memory safety in Rust is balanced on a delicate precipice; if the language has already fallen off the precipice I'm confident in saying it will never go back. (Well, unless you add a GC, which is what I'd do if I were in charge of Zig. There's nothing wrong with GC; it's the tool people have created for the very purpose of ensuring memory safety with unrestricted aliasing.)
That's a good description of the issue. It's a bit frustrating to see this notion that Zig is "safe" or comparable to Rust in regards to memory safety at all. I don't prefer to use Rust as it's verbose, but at least I realize that if you don't do the lifetime analysis dance in the language you essentially need a GC to manage memory.
Yep. And that's totally valid! I like using GC'd languages; garbage collectors are some of the most successful systems ever developed in computer science.
However as Rust has shown, creating non-cyclic async is really hard, and Nim isn't an exception. That means async requires you run ARC with a cycle collector (ORC), and it works on embedded with tweaks. Overall the non-locking ARC + move semantics runs really fast even on MCUs so it's worth avoiding cyclic strctures.
I've been working with Rust since pre 1.0 and I've felt fearless running on bare metal coming from a Higher Level Language for most of my career. It's almost embarrasing how care free it is to never have to think about memory management. But I did take a look at Zig earlier this year and noped out because of it's lack for memory safety. Although it does look nicer than C, for me it feels like I would be going backwards
That is probably not feasible either, for the same reasons the borrow checker isn't feasible. Zig has this pattern where they manually pass in an allocator into functions that allocate memory.
TL;DR: Yes! But Zig is a Stand-in for simpler languages, and while in its infancy, I can foresee it to be "done" and "fit into my head". Rust requires a suspension of disbelief that all this complexity can be tamed and "gotten right" with a leap of faith for the language, compiler and everybody in-between including myself, that I'm not confident enough in the "proofs" it provides.
To quote from Dijkstras essay "On the cruelty of really teaching computing science".
> Right from the beginning, and all through the course, we stress that the programmer's task is not just to write down a program, but that his main task is to give a formal proof that the program he proposes meets the equally formal functional specification. While designing proofs and programs hand in hand, the student gets ample opportunity to perfect his manipulative agility with the predicate calculus. Finally, in order to drive home the message that this introductory programming course is primarily a course in formal mathematics, we see to it that the programming language in question has not been implemented on campus so that students are protected from the temptation to test their programs. And this concludes the sketch of my proposal for an introductory programming course for freshmen.
There are more ways to proof correctness than through the type system alone. Clojure programs are surprisingly high quality despite being dynamically typed.
I feel almost sorry to not write a lengthy counter rebuttal, but sadly I agree with everything you said! Would I prefer Zig to have a strong type system? Absolutely! But the reason why I choose Zig as a comparator is not because I feel that it does a significantly better Job at correctness than Rust, neither do.
If I had to build something (computer)-verifiably correct I'd probably Idris with linear types or something, but that's not a real contender, and neither is Carp or Linear Haskell.
But I feel that with Zig I'd at least have the chance to proof the correctness of my program manually (with pen & paper). Although Zig is still heavily in flux it's simple and focused enough that it "being done", and me thus being able to understand it in its entirety, is at least visible on the horizon (especially if one compares the number of people working on Zig with the number of People working on Rust, although Carp might be an even better contender in that metric). With Rusts complexity that date is much further out into the future, if it is ever achievable. And even then I'm not sure that I could fit enough of the language into my head to be satisfied with the correctness of my proofs, ignoring the fact that it also has to fit into the compiler devs heads sufficiently to ensure that the compiler is behaving correctly.
My critique is that little of this complexity actually comes from the parts that make the language unique, non of it is essential, it's all incidental complexity. It's in the macro system, the standard library, the hand-wavy semantics of "box" and compiler intrinsics, of magic low level traits, an overzealous module system, way to much fondness of abstractions and their removal by compiler magic, and a general nonchalantness towards C++'s complexity.
I don't understand how something that isn't type-safe (memory safety is a subset of type safety) can be more "correct" than a language that is type-safe. It is certainly hard to prove Rust type-safe. But that's infinitely easier than proving something type-safe that trivially isn't type-safe!
> way to much fondness of abstractions and their removal by compiler magic
Heavy reliance on abstractions is what makes lifetimes and the borrow check practical. For example, Arc, Mutex, Cell, and RefCell are all abstractions that serve as library-defined extensions of the core lifetime rules and are generally optimized away by the compiler. I'm extremely skeptical that lifetimes/borrow checking can be made usable without this sort of thing. If a goal of Zig is to minimize abstractions, then that's yet another reason why adding a borrow check to it after the fact will not work.
When Zig is "done", it will have to either (a) become memory safe by addition of Rust's lifetimes/borrow check (which is probably impossible, but if it were done it would bring about a lot of the things you say you don't like about Rust); (b) become memory safe by addition of a GC (which is the best option, but I doubt it will do so); (c) not be memory safe (which is probably what will happen). There is no free lunch. And then, assuming it chooses option (c), you run into this problem: Try as hard as you want, you will never be able to prove a language type-safe that isn't type-safe, no matter how simple that language is.
Unless rust has some kind of de brujn criterion (and properly formalising and bootstrapping it from a minimal core would be a step into that direction) the proof it produces is only as good as the entirety of the codebase.
Correctness proofs are an inherently human thing, it's all about convincing your opponent of the proofs validity.
One could of course build a language on inuitionistic math a a la coq, but that's neither practical nor realistic, so human checked proofs will have to do, and simplicity helps a lot with those.
Edit: It feels like we're talking about subtly different definitions of correctness. I mainly care about correctness of specific program instances, whereas you are more concerned about the correctness of every instance expressable by the language.
In other words your argument was incoherent and you have no rebuttal. What a nice way to tap out.
Saying that Rust isn’t verified enough is fine. But the post became incoherent when you brought up another language.
Obviously the spec isn't meant to be a necessary read for end users.
Experimentation is great for determining what the implementation you have right in front of you does, but not what is actually guaranteed. Trusted reference web pages are great (in fact, cppreference is AWESOME particularly for highlighting text that applies version-by-version of the C++ standard) but don't always cover everything. "I googled my error message" hits are so-so (but they may help you solve your problem)
But the standard is the document that really lets you get into "how SHOULD it work", even while doing battle with 2 or 3 different compilers each of which implement a subset of the language you are theoretically programming in. (It's getting better, but finding a compatible subset of C++ supported by MSVC, gcc and clang used to be an interesting puzzle. And what a bummer to "switch to C++14" -- only with a list of caveats longer than most people will even read through, because of differential language & library support)
I'd also just recommend becoming comfortable finding & extracting data from a daunting document like the C++ standard! It's a skill that will stand you in good stead, and I feel like being able to do this is one of the things that led to having my technical expertise well appreciated at work.
Also, completely unrelated, I think it's really sad for Zig that it has become (against its creator will) the champion of the rust-hating crowd. Zig is cool and Andy's work on cross-compilation is a piece of art, but why is Zig in this discussion? Zig is definitely a “growing” language, not a language built on top of a formal spec. Zig doesn't have a defined memory model either!
Rust has a lot of great ideas, but it feels more like a prototype that got out of hand, than a focused safety oriented tool. Combine that with a healthy dose of over-optimistic marketing and wishful thinking and you got yourself a bubble.
Rust is way too much cathedral than bazaar. Does it really have the need for a Macro system yet? Except for C++ template envy, JS has shown that transpilers are a workable alternative. Is the complex module system with access control necessary yet? Does the Std library need to be this comprehensive?
> but why is Zig in this discussion?
Simply because it's simple, and is perceived to target the same niche as Rust. I would have used Carp as a stand-in, but that's even more esoteric.
I'd argue that neither Rust nor Zig actually target that niche properly, Rust is too complex to get the implementation and models right, while Zig is too conservative and not having the right model as a goal.
Yet if I had to bet on either, it'd be on Zig simply because:
There are two methods in software design. One is to make the program so simple, there are obviously no errors. The other is to make it so complicated, there are no obvious errors.
-Tony Hoare
> Simply because it's simple
But by what metric? For instance, why do you consider comptime being simpler than macros? Macros are a dedicated tool built for one specific job (compile-time pre-processing). Comptime means re-using the same syntax for both compile-time and run-time execution. Yes it is pretty elegant, and I like it but it is more complex that the Rust model because you now have a tool that needs to work with both cases.
Same async, Rust has async functions and regular function, Zig has functions that are automagically coerced to the right type by the compiler. Is it ergonomic, probably (but not always), but it's much more complex to implement under the hood.
Last but not least: memory safety, Zig's goal is to achieve a reasonable level of memory safety by using a collection of built-in sanitization tools. Those tools are much more complex than the borrow-checker (which rely on 2 simple rules: move semantic and shared XOR mutable)and for the most part they don't even exist!
Zig's goal is developer ergonomics, and it's really cool to see radical experimentations in that space (and as I said, what zig has done with regard to cross-compilation is inspirational for every language author on this planet) but it's not simplicity. Otherwise, the Zig compiler probably wouldn't be 120% bigger than the rust one right?! (Rustc has 1.6 million loc[1], Zig has 3.5[2], more than twice as much.
I reported recently[1] that the zig compiler is 140,331 lines.
[1]: https://mobile.twitter.com/andy_kelley/status/14539106409432...
(I'm a big fan of your work btw, I'm just really sad that a few people on this forum seams to like Zig for the sole reason it's not Rust)
All what code exactly? Here's a file-by-file breakdown of the line count:
Let's wait until Zig gets a proper specification before trying to argue that we're doing a good job. Rust is much more mature and I'm sure people have found, fixed and standardized things that we probably have never even experienced once, ever.
Right now the compiler doesn't even implement the full language.
(This is handwavy and imprecise, but it's the easiest way I can explain it.)
However if you start manipulating atomics directly, while the atomics themselves are safe the behaviour can get quite complicated below seqcst. The latest Crust of Rust is rather illuminating to the subject for people like me who had not had to really dive into the logic and implications.
Given that Rust's "shared xor mutable" is guarded with mutexes, atomics or other synchronization primitives, memory reordering can't violate Rust's memory safety guarantees (assuming the sync primitives are not buggy, of course).
But this does not guarantee that your program logic is correct if you are sloppy with memory ordering. If you do two atomic stores, they are not guaranteed to occur in that order if they use "relaxed" ordering.
This gets more complex if you're writing device drivers and communicating with peripheral devices using memory mapped i/o.
I think rust gets away because atomic ptr is a raw ptr so it can only be dereferenced in unsafe code. So in practice relaxed atomics require unsafe where it matter (which makes sense to me).
The memory model is something different: it governs how memory operations are ordered in a concurrent system. You can have a program with no memory safety issues, but nonetheless is not correct (that is, it doesn't do what you expect it to do) because you have specified the wrong memory ordering of some operations.