Hard Mode Rust (2022)
matklad.github.io
matklad.github.io
* The fact that there is an `Oom` error that can be thrown means that there is no increase in reliability: you still don't know how much memory you're going to need, but now you have the added problem of asking the user for how much memory you're going to be using — which they are going to be guessing blindly on!
* This is because the memory usage is not much more predictable than it would be in easy mode Rust. (Also that "mem.lem()/2" scratch space is kind of crappy; if you're going to do this, do it well. Perhaps in allocating the correct amount of scratch space, you end up with a dynamic allocator at the end of your memory. Does that sound like stack space at the start of memory and heap space at the end of memory? Yes, your programming language does that for you already, but built-in instead of bolted-on.)
* Furthermore, the "easy mode" code uses lots of Box, but if you want you can get the benefits of RAII without all the boxes by allocating owned vectors scrupulously. Then you get the benefit of an ownership tracking system in the language's typesystem without having to `unsafe` your way to a half reimplementation of the same. You can get your performance without most of the mess.
* Spaghetti can be avoided (if you so desire) in the same way as the previous point.
What you do achieve is that at least you can test that Oom condition. Perhaps what you actually want is an allocator that allows for simulating a particular max heap size.
When I read "So we do need to write our own allocator", I winced. Almost every time I've had to hunt down a hard bug in a Rust crate, one that requires a debugger, it turns out to be someone who wrote their own unsafe allocator and botched the job.
The problem with that approach is that you end up with a “function coloring problem” akin to async/await (or `Result` for that matter). Function that allocate becomes “red” functions that can only be called by a function that itself allocates.
Like Result and async/await it has the benefit of making things more explicit in the code, but on the flip side it has this contaminating effect that forces you to refactor your code more than otherwise, and also cause combinatorial explosion in helper functions (iterator combinators for instance) if the number of such effect is too high, so there's a balance between explicitness and the burden of adding too many effects like these (or you need to go full “algebraic effect” in your language design but then the complexity budget of your language takes a big step, it's unsuitable for either for Rust or Zig which already have their own share of alien-ness (borrowck and comptime, respectively)).
A global allocator is IMHO worse than either alternative though.
Also, from my experience with Zig, it's not such a big problem. It's actually a good thing to know which code allocates and which doesn't, actually improves trust when working with other people's code
With Zig aiming at the kind of code you can write in C it doesn't surprise me that it works pretty well. (Also I'm a bit sceptical about the actual future of the language, which IMHO came a good decade too late, if not two: I feel that Zig could have succeeded where D couldn't, but I don't see it achieving anything nowadays, as the value proposition is too low IMHO. Except as a toolchain for C cross compilation actually, but not as a language).
Just fyi, console games settled on global allocator. Because everyone allocates and game needs to run at say 7 out 8Gb used consistently. It makes folks passing around pointers to allocators completely wasting their time and space. There are small parts of code with explicit pooling and allocator pointers but they are like 5% of total code.
It is funny when C++17 standard got PMR allocators that make a dream of explicit passing allocators around come true then folks noticed that 8 bytes in every string object are not that cheap. There are very small islands of usage of PMR allocators in the library ecosystem.
It does not make global allocators universal truth. It just shows that tradeoffs are different.
In general, the idea that each individual object is uniquely allocated doesn't make sense since objects of one type almost never come alone - especially in games, and you definitely don't want to carry allocator pointers in each invidiual object around, at most pass them into explicitly called creation and destruction functions.
Games typically only have few lifetime buckets (a frame, an active map region/zone, an entire map/game session, or static lifetime for the entire duration of the game), each of those can be handled by an arena allocator that can be flushed at once without calling individual destructors (because 'objects' should really just be dumb data items).
Most of this doesn't fit into the memory-management-via-RAII idea of automatically self-destructing objects when they are no longer referenced of course (because then the object - or at least the smart-pointer - indeed needs to know how it was allocated).
I do know though that a lot of ancient game- and game-engine codebases still do this OOP-inspired 'object spiderweb' (e.g. each object living in its own heap-allocation and referencing other objects via smart pointers - but that is really not how it should be done since the late 90s when memory latency became an issue).
Passing allocators hit a brick wall because most of the code is threaded tasks and to make something that will go into another async API like GPU has you need to allocate in non-blocking way. Which means that each specialized allocation and each non specialized allocation (like you need to pass something to tasks further in task graph) need to be non blocking on unknown statically thread. Having global multiple allocators does not make it easier to test and reason about. It just means that passing things as arguments is not useful if you are not 'calling code' synchronously for the most part. TLDR task graphs and async APIs make code look alien to people outside of gamedev. That's a fact of life.
Object graphs are independent from that. I can not say gamedev has resources to polish object graphs as much as in old smaller console times or like embedded folks would like. I have to confess that lots of our objects are not even in the C++ code anymore. They are in runtime of visual language that designers used... We live fast and ship mostly broken things... /end of rant
It's not. It makes a world of difference. Having a hard guarantee that you are not dragging any transient dependency is often a big deal. Not to mention the maintenance guarantees.
> This criticism is aimed at RAII — the language-defining feature of C++, which was wholesale imported to Rust as well.
> because allocating resources becomes easy, RAII encourages a sloppy attitude to resources
Then lists 4x bullet points about why it is bad.I never once heard this criticism for RAII in C++. Am I missing something? Coming from C to C++, RAII was a godsend for me regarding resource clean-up.
Towards the end of the 90s, game development transitioned from C to C++, and was also infected by the almighty OOP brain virus. The result was often code which allocated each tiny object on the heap and managed object lifetime through smart pointers (often shared pointers for everything). Then deep into development after a million lines of code there's 'suddenly' thousands of tiny alloc/frees per frame and all over the codebase but hidden from view because the free happens implicitly via RAII - and I have to admit in shame that I was actively contributing to this code style in my youth before the quite obvious realization (supported by hard profiling data) that automatic memory management isn't free ;)
Also the infamous '25k allocs per keystroke' in Chrome because of sloppy std::string usage:
https://groups.google.com/a/chromium.org/g/chromium-dev/c/EU...
Apart from performance, the other problem is debuggability. If you have a lot of 'memory allocation noise' to sift though in the memory debugger, it's hard to find the one allocation that causes problems.
Isn't it pretty common in gamedev to use bump allocators in a pool purely because of this. Actually isn't that pretty common in a lot of performance critical code because it is significantly more efficient?
I feel like RAII doesn't cause this, RAII solves resource leaks but if you turn off you brain you are still going to have the same issue in C. I mean how common is it in C to have a END: or FREE: or CLEANUP: label with a goto in a code path. That also is an allocation and a free in a scope just like you would have in C++...
Yes, but such allocators mostly only make sense when you can simply reset the entire allocator without having to call destruction or housekeeping code for each item, e.g RAII cleanup wouldn't help much for individual allocated items, at most to discard the entire allocator. But once you only have a few allocators instead of thousands of individual items to track, RAII isn't all that useful either since keeping track of a handful things is also trivial with manual memory management.
I think it's mostly about RAII being so convenient that you stop thinking about memory management cost, garbage collectors have that exact same problem, they give you the illusion of a perfect memory system which you don't need to worry about. And then the cost slowly gets bigger and bigger until it can't be ignored anymore (e.g. I bet nobody on the Chrome team explicitly wanted a keystroke to make 25000 memory allocations, it just slowly grew under the hood unnoticed until somebody cared to look).
Many codebases might never get to the point were automatic memory management becomes a problem, but when it becomes a problem then it's often too late to fix because that problem is smeared over the entire codebade.
Generally the expensive part of memory allocation isn't calling malloc, it's the underlying operations, when the underlying operation is +sizeof(T) and free is a noop you can happily keep using RAII and not care.
CLARIFICATION: I'm saying having a T* member is an antipattern, make it std::unique_ptr with a custom deleter or it should be in a container, struct of arrays...
Destruction should be explicitly performed at a specific time and place, and not happen decentralized in some random location of the code when a reference goes out of scope (I realize that this dismisses the whole idea of RAII and destructors and garbage collection in general, but ¯\_(ツ)_/¯).
> It's a really naive allocator that can be very fast due to the tiny amount of housekeeping involved, but you have to live with a pretty heavy constraint: there's no "free" operation for individual request - you just destroy the whole thing.
To me, this sounds like an arena allocator. I remember seeing them in Apache SVN source code a squillion years ago. What's the big deal? For RAII, we can just do placement new in the ctor, and "arena free/delete" in dtor -- which might do nothing in practice.Also, was the original article focused on gamedev? I don't see any mention of it. For us normie programmers working on CRUD apps, we don't need bump/arena allocators. Ctor->new, RAII/Dtor->delete is just fine for us.
Yes arena allocators is the word I couldn't remember, a bump allocator is a way to implement an arena allocator and probably the most popular.
Also I 100%, in general you shouldn't need an arena allocator, but if you are building a library you should be aware of allocators and make however you are allocating memory configurable so that if you do need to some day change it it doesn't involve a complete rewrite.
Whereas I am expecting profiler data showing memory issues caused by RAII patterns, from your side, or any of those influencers on whatever platform they complain about.
I guess you would like me to prove that the lag is caused by RAII but notice that I didn't actually wrote that - that's your reading of it. I just noted that the symptoms seem to point to it.
This proves what about their resource usage?
They should know a couple of things about how to deliver something into production.
Additionally influencers complaining about them, usually never even wrote a basic hello world, let alone understand how memory management works.
So, in the context of a thread about resource usage, what does the commercial success of the tools prove?
My credentials are having been a IGDA member for several years, already been in places like SCEE SOHO studio, regular GDCE attendee during its time in UK, and first GCDE in continental.
And to turn you question around, it proves that arguing about RAII impact on commercial success of the tools is completly irrelevant without solid proofs, they are printing money.
This one is big. It is a lot of accounting work to individually allocate and deallocate a lot of objects.
The Zig approach is to force you to decide on and write all of your allocation and deallocation code, which I found leads to more performant code almost by default – Rust is explicitly leaving that on the table. C obviously works the same way, but doesn't have an arena allocator in the standard library.
Re: C++ vs Rust, it might be more of a pain in Rust because of this: https://news.ycombinator.com/item?id=33637092
The big reason is that RAII encourage fine grained resource allocation and deallocaion, while explicit malloc() and free() will encourage batch allocation and deallocation, in-place modification, and reuse of previously allocated buffers, which are more efficient.
The part about "out of memory" situations is usually not of concern on a PC, memory is plentiful and the OS will manage the situation for you, which may include killing the process in a way you can't do much about. But on embedded systems, it matters.
Often, you won't hear these criticisms. C programmers are not the most vocal about programming languages, the opposite of Rust programmers I would say. I guess that's a cultural thing, with C being simple and stable. C++ programmers are more vocal, but most are happy about RAII, so they won't complain, I can see game developers as an exception though.
> In language performance benchmarks, you often see C++ being slower than C.
This is the first that I have ever seen such a statement. And, I am sure we could legitimately write exactly the opposite: <<you often see C being slower than C++>>. Not much being said here. As I understand, C++ template code can be much faster than equivalent code in C because the compiler has more metadata for optimisation. The classic case is C's qsort() vs C++'s sort(), where the C++ version is much faster due to compiler optimisation.There was an HN thread a few months ago[0] debating whether RAII is reason enough to disqualify Rust as a systems language, with some strong opinions in both directions.
I have got a feeling that the division comes from two different system programming folks. Hard realtime folks see unbounded time and uh-oh its fatal. On other side soft realtime folks ask about probabilities and profiling data and its relatively fine for them if probability is low enough. Where both of the camps would agree if number of effects is bound somehow.
So it's interesting to see something very similar crop up here in a different language and domain: throw a true barrier down on when and how your code can request resources, as an explicit acknowledgement that the usage of them is a significant thing.
Rust sucks at arenas. I love Rust. But you have to embrace RAII spaghetti or you’re gonna have an awful, miserable time.
The only point I agree with is the deallocation in batches being more efficient.
> Lack of predictability. It usually is impossible to predict up-front how much resources will the program consume. Instead, resource-consumption is observed empirically.
I really don´t understand this point.
C++ and Rust share the same criticism
- poor readability
- poor code navigation
- poor compilation speed
I find it interesting how they call simplicity "hard mode", quite concerning
Yet, most implementations do not consider SIMD parallelism, or they do it in a non-robust fashion, trusting the compiler to auto-vectorize.
I did exactly this in my QR Code generator library, C port and second Rust port: https://github.com/nayuki/QR-Code-generator/blob/master/c/qr... , https://github.com/nayuki/QR-Code-generator/blob/master/rust...
Writing the Rust code felt way safer because the language took care of enforcing rules about borrowing parts of the buffer.
Sure, you still have garbage collection but it’s generally for intermediate, short lived values if you follow this pattern of allocating resources in main and divvying them out to your pure code as needed.
You can end up with some patterns that seem weird from an imperative point of view in order to keep ‘IO’ scoped to main, but it’s worth it in my experience.
Update: missing word
The threadpool is usually the easy part of concurrent algorithm's, with the concurrent queue being one of the hardest to implement in a performant manner. The is a "draw the rest of the owl" moment...
Maybe I should practice it a bit more.
I just can't. If you're ignoring the UB checker, what are you even doing in Rust? I understand that "But it should be fine, right?" is sarcastic, but I don't understand why anyone would deliberately ignore core Rust features (applies both to RAII and opsem).
This is a skill issue.
C, C++, Rust, it doesn't matter - if you have a systems problem where you are concerned about resource constraints and you need to care about worst-case resource allocation, you must consider the worst case when you are designing the program. RAII does not solve this problem, it never did, and anyone who thinks it does needs to practice programming more.
It is not "hard" to solve though. It's just another thing your compiler can't track and you need to be reasonably intelligent about. There are sections of a program that cannot tolerate resource exhaustion, for which you set some limit, which is then programmable in some way, which then uses some shared state to acquire and release things of which RAII may help. But you still must handle or otherwise report failures.
This is a bread and butter architecture concern for systems programming and if you're not doing it, get good at it.
(I write mostly C code myself, but I'm also not immune to f*cking up now and then even though I write code for decades now).
To be clear, of course I'm also not immune to fucking up :-).
But beyond obvious hard differences in what languages in question allow about a given topic (of which indeed there are none between C/C++/Rust/Zig on the topic of memory utilization & co), everything is still just a skill issue.
More seriously, afaict the article in no way even mentions RAII solving anything about resource constraints (not even as a hypothetical thing to disprove); only about RAII simplifying doing basic management logic (at the cost of worsening reasonability about resources!).
Thinking about resource constraints is unquestionably harder than not thinking about them, which is what the article is about.