Are there many large programs (outside of embedded) where it ever really makes sense?
I ask because I am a fan of C, and I am sure there will always be old legacy C code, but is manual memory management essentially antiquated?
Are there many large programs (outside of embedded) where it ever really makes sense?
I ask because I am a fan of C, and I am sure there will always be old legacy C code, but is manual memory management essentially antiquated?
That said, to my mind "large program" and "manual memory management" do not go together. It's difficult to comprehensively reason about large programs in toto in any language. In some respects C is better than many other languages in this regard as there are only a few ways components can reasonably fit together--not many features and complex semantics to cut through. However, when it comes to manual memory management in C it's effectively impossible to reason about the program as a whole, even if all other traps, like aliasing, are taken out of the picture.
So even if I have a large application written entirely in C, it's never built that way, and that's not how I will conceive of it. It's built of discrete components with very minimal--preferably zero in most cases--cross dependencies. In fact, each component is typically maintained as either a single source file or an entirely separate library so that from soup-to-nuts the boundaries (functional and technical) are consistent. Components themselves, even single-file components, will typically be structured similarly internally. I try to avoid utility libraries with ad hoc functionality or even narrow, cross-cutting functionality. And I emphasize strictly functional interface boundaries. So, for example, if I use a bump allocator in a functional component, this fact is completely opaque and irrelevant outside the component. All of this sometimes results in more code duplication than people might normally be comfortable with. But it makes it far easier to reason about each component in isolation, and then in turn reason about their interactions.
A language like Rust is literally built around this type of code flow and data discipline. So in many respects I'm making a strong argument for using Rust instead of C. If you stopped reading here, no problem. But I would still point out that this type of architectural discipline has many other benefits, like making it easier to mix-and-match languages (e.g. C + Lua + Swift) or to do major refactors. Some of the rich semantics of Rust and its ecosystem can easily lead one stray. To my way of thinking, something like Cargo is almost all liability for projects that expect any longevity. And while I think generational memory arenas are awesome, if and when the semantics of the arena leak beyond a functional interface boundary (whether "internal" or "external"--rarely a substantive distinction to me), then that presents a major problem.
IOW, there shouldn't be "large programs" in any language. There should just be a bunch of small programs revolving around a central, cohesive problem model. (A high-level functional problem, not a technical problem like memory management.) And that should hold even if your final artifact is an enormous, static binary. One of the major benefits is that you're not stuck choosing just one memory management technique or even one language. One component can employ a bump allocator in a streaming parser, and another written in a language with mark & sweep GC. And if you enforce strict functional and semantic boundaries so these choices can't leak, you can often mix-and-match them with relative ease. One implication is that sometimes it could be more prudent to write a particular component in C rather than Rust, such as if the Rust component would end up a giant ball of unsafe{} blocks.
This perspective is rather peculiar and unpopular these days, however. When baseline assumptions are an autocompleting IDE and seamless importation of all other software in an ecosystem, the calculus is entirely foreign. So I wouldn't expect it to resonate with many people.
On Google Earth, garbage collection pauses weren't really acceptable, but we also wanted to add new functionality without being slowed down. Rust is great for a lot of reasons, but as many know, it can be slower on the feature velocity front.
In some AAA games, they need much more flexibility than the borrow checker thinks is appropriate. At some point, one has so many unsafe blocks that undermine the guarantees of surrounding code, that it makes more sense to use a different paradigm.
I know it's taboo to say this, but in many real world use cases, memory safety bugs aren't that bad and have the severity of any other logic error. Having a few more bugs per release might be worth it, if it means not dealing with the borrow checker's restrictions around abstraction, encapsulation, and polymorphism. You just fix them and push a new release.
Also, the problem is nicely mitigated by tools like Address Sanitizer, and will get even better with upcoming advances such as memory tagging and generational references, all which detect memory problems with surprising efficacy and accuracy.
In short, manual memory management is the right choice when Rust's drawbacks are a bit too much, and benefits aren't enough, for a particular situation.
Just my two cents!
I don't think it's true of many real world use cases and if anything is only true of perhaps trivial use cases, where an application is run on a computer that has no other sensitive information and is not connected to the Internet.
I think what's true is that many programs don't ever reach a point where they're popular enough that someone can reliably exploit it. The threat of memory related errors has not much to do with use case and more to do with whether someone cares enough to identify that error, can reliably reproduce it, and whether the application that exhibits that error is on enough computers that exploiting it is profitable. If those three conditions are satisfied then it matters not whether the program is a video game, a word processor, an image viewer or whatever.
(And the funny thing about "trusted" communications: they're trusted until an attacker figures out how to store and forward untrusted data, at which point a logic error in the server itself is now nicely composable with code that "didn't need" to be particularly memory safe.)
If you're running a single threaded cloud lambda, you could allocate into an arena and never free, so you don't get UAF or DF.
It would be nice if languages had Address Sanitizer built into the language. Maybe some sort of system of checking who owns and borrows memory at any point in time.
This may not be a big deal for a toy game like Tetris. In a multiplayer AAA game it could lead to a bunch of bitter consequences though, from malware pwning customers' computers to ruined multiplayer experience because the bug would affect the server.
I'd say that a memory safety bug is not like a blemish on the surface of a building, unsightly but tolerable. It's more like a crack in the foundation, which may look tiny on the surface, but who knows how deep it goes.
My comment was about the tradeoffs we make for our memory safety guarantees, and whether they're always worth it.
It’s impossible to verify, but some of the workarounds in very large c apps seem like they have about as much overhead as A modern GC.
On Google Earth, even sub-millisecond pauses would have been too much, we had to fight pretty hard for every sliver of performance we could get.
The effectiveness of some key optimizations were sensitive to how tight we could make tail latencies for important operations, so a GC always had an unacceptable cost to the extent it occasionally introduced a stall at an inopportune moment. If GC stalls were more like 10 microseconds, consistently, that would be much more viable. I have no idea if that is realistically achievable.
GC's don't exist in abstract, and there are plenty of GC enabled languages that also have value types and manual memory management capabilities.
There is no need to throw everything away when the right GC enabled language is chosen in first place.
Though for completeness’s sake I will add that in case of the JVM, variable tail latency still can be experienced due to the JIT compiler.
I have some doubt that Rust feature velocity is meaningfully worse than C++ or C. I would bet good money the other way to be honest.
It also depends on one's familiarity with the languages, I've used both pretty extensively and can't comment on the velocity when just learning.
Just my experience, YMMV.
When you're just learning it Rust can definitely be a lot less productive and a lot more frustrating. But once you get to an expert level I think there's just no contest, and you can churn out code high quality code significantly faster than you can in C++. (At least that's how it is for me, and I've been programming C++ for over 10 years and I'm don't intend to ever go back.)
Indeed. IME (around 10y C++, mostly c++14 and c++17, and 6y Rust), I measure x2-x3 velocity in Rust, a lot of which comes from the tooling, a part of which from ub-freedom and much easier and safer parallelism, and the rest from the expressive type system and functional polish.
This is for initial development. Factoring in the maintenance over the years I suspect it tends to an order of magnitude of difference.
But contrary to expectations, I didn’t have many unsafe usages — they were all hidden under some abstraction (like allocate_object function of an ObjectPool), there is perhaps 7-8 in the whole (semi-partial) implementation? (Though it doesn’t yet have a GC, so there is that).
So all in all, in my anecdotal experience one can reasonably well hide the unsafe usages and then use the good expressivity of the language to combine these APIs in strictly safe ways. I also ran the program regularly through MIRI (a rust AST interpreter, which is similar to Valgrind in goals for my use-case) and whenever I found a memory bug I could pin-point it to a few lines of code at most, greatly speeding up debugging.
For many applications, allocations fall into two categories: One, temporary allocations that work well with a stack allocator (which is often going to be the main stack where the allocations and deallocations are handled automatically, but sometimes it's useful to have secondary stacks with different lifetimes). And two, larger allocation categories where you can pre-plan access. Many many things can just be allocated statically, others can be allocated from simple pools, and maybe a few things will need complex pools, or even compacting allocators where you end up using smart pointers or smart handles to track them. Once you're used to programming this way, figuring out allocators usually isn't where you spend most of your time. (And as others have said in this thread, if performance is a concern, you spend the same amount of time thinking about this stuff in managed languages, as well.)
* The Big Asterisk is that it's hard to program this way 100% correctly. While it's not hard to get code that works 99.999% of the time for normal usage, it's very hard to make sure nothing in your manually memory managed code is exploitable. Rust, for instance, adds a bunch of language features to fix this issue, and that's one interesting way to go. I do wonder, however, if for many applications it won't turn out that the sandboxing model of something like WASM is more productive. You can still code everything in good ol' C++, but the worst an exploiter can do is hit the edge of the sandbox and crash your application, instead of pwning the user's system.
That is still exploitable. Let’s say you have a web app and the js code interacts with the wasm output - if the latter is exploitable, js code may be as well and it can be catastrophic if it is some deeply personal thing, or your bank account. And that can all be controlled by data alone, e.g. if it is a PDF converter or something.
HOWEVER, partial manually managed memory is unlikely to go anywhere. It's just too powerful of an optimization tool, especially as memory is one of the things that's just not getting much faster. If you're trying to optimize something like a game engine, being able to traverse your objects as fast as possible means they need to be laid out in memory sequentially. This is being somewhat branded "data oriented design" in the game world ( https://en.wikipedia.org/wiki/Data-oriented_design ), but it's broadly applicable.
Similarly you'll see things like arena allocators in all sorts of places (like protobufs https://developers.google.com/protocol-buffers/docs/referenc... ). This kinda blurs the line between "manual memory management" and "garbage collection" since it's sort of both, depending on which side of the usage you're on (individual allocations end up essentially GC'd, but the entire heap is manually managed). Again, it's just too fast of a tool to lose.
The ability to jump between worlds is a very strong strength of C++, and it's something Rust can mostly pull off as well with the unsafe escape hatch.
Ignoring the awkward issue of circular references, yes. But that's not always trivial to ignore.
In Rust I try to avoid `Arc` if I really don't need it due to its behavior that is hard to track in certain situations. I use it to share items over thread boundaries by value, that's it.
Codebases that use shared_ptr for everything because they think it ‘solves’ having to manage memory are smelly, it shows that no one has thought about the ownership model. Not to say there isn’t a solid use case for it, but some developers use it by default because they don’t fully understand the trade offs/why it’s important to think about ownership of managed memory.
As an example in our medium to large code base we have two places we use a shared_ptr, completely off the hot path and for long lived objects that aren’t copied more than a few times, but need to share ownership of the pointee - in a similar use case of your use of Arc.
Almost all of the time unique_ptr, with a single owner, is exactly what you want to replace the memory management side of a dynamically allocated object. From there you will pass around raw pointers or references to the object managed by the unique_ptr (for functions that need to ‘borrow’ the object).
Something interesting about shared_ptr is how it relates to weak_ptr, although I’ve never seen weak_ptr in anything out side documentation!
We've been going around this issue by storing flat collections instead of structuring the data as a tree. We for example have parents and children in adjacent vectors, and parent only needs to know the ids in the children vector to be able to iterate over them. In this case you can do a data structure storing the id values that are `Copy`, meaning this whole data structure can be `Copy` and is super easy to reason about.
The other use-case is a database connection pool. When you take a connection out from the pool, you get a reference to the pool itself. The pool is quite often internally inside an `Arc`, and the connection gets wrapped into a guard object holding a `Weak` to the pool. With RAII when the connection drops out from the scope, the `Weak` is checked is it still valid (is the pool alive) and if so, the connection is returned back to the pool.
I feel like it is not possible from low level language, but contrary, it is possible from managed languages. It is easier to give a small escape hatch in a GCd language to allocate n bytes and manipulate it as if you were in C, then to add some hacky semi-GC to C/C++/Rust. Don’t get me wrong, unqiue pointers are really cool and useful, but shared pointers for example can’t be that painfully implemented without full control over execution (thinking of circular dependencies here).
None of them let you do either placement allocations nor arena allocations.
.NET at least lets you do a linear array of structs, but that's nearly it.
Python also has similar techniques available. .NET indeed leads the bunch though. But arbitrary pointer arithmetics is not necessary for arena allocation and safe usage.
Hell, you can't even do linear object arrays in the JVM (still waiting on value types...), the most basic of memory optimization techniques.
Also for a JVM or Python, you can't avoid the extra overhead in an object (such as the ref count for python), which increases cache pressure and limits how much cache locality you can exploit.
Sure, value types will be a huge win, but it is not like you can’t store your RGB class’s data quite trivially into a huge linear structure as SOA or AOS as needed, and deserialize as needed.
Python got it mostly right, though a language with a smoother transition between 'manual memory managed kernels' and 'general purpose glue' is still missing.
Rust has a great type system and handles manual memory management remarkably well. Sadly, it is missing the garbage-collected yin to complement its manually managed yang.
If you are building large-scale, high-performance infrastructure, manual memory management is to some extent unavoidable. In many database engines, for example, memory lifetimes and object lifetimes are distinct and independent concepts. There are classes of algorithm optimization that require strict and explicit understanding of the address space your process claims to own. The silicon does not respect the object model of your programming language, but you still need the silicon to manipulate memory directly (as it understands it) for performance reasons. This seems complicated but understanding it well enough to leverage it has significant performance implications.
That said, I don't see a lot of heap allocation in things like database engines. Almost all the memory used is allocated and deployed at bootstrap, and most of the remainder is on the stack. Not a lot of attack surface for memory management problems to show up. Other applications may be different.
A single bad pointer arithmetic, out of bounds access can wreak absolute havoc.
As far as I'm concerned, it's automatic memory management (with lots of control). There's also the other end of the automatic memory management spectrum, garbage collection.
CVE database is updated almost every week on those "best" practices.
Why though?
I use C++ and I use a variation of malloc and free all the time. The arguments I hear against malloc is:
1. You'll leak memory.
This is super easy to solve in literally 300 lines of code or less. Just have a vector of your memory allocations in debug mode. Then call a custom allocator and free function that adds an allocation and removes an allocation from this list. At the end of the program, print out any allocations that are in the list that never got freed. You can use macros to print the exact file and line number that allocated the memory and then fix it very quickly.
2. It's not typesafe.
So just use new if you're really concerned with type safety and then overload the operator to get the benefits I described above.
3. It's unsafe.
How is this any more unsafe than RAII? Aren't there still the same risks involved?
4. Nullptr exceptions.
Once again, same problem with smart pointers. Technically if you're not transferring ownership you should be passing a const reference or a raw pointer rather than a smart pointer anyways. So the end result is the same.
Maybe I'm missing some huge thing, but I have not heard any convincing argument for using shared pointers over malloc. (And yes I know you should really use new if you're not using POD because of constructors and all that stuff, but I'm equating new and malloc for brevity).
Lastly, I've found that when I use shared pointers I get lulled into a false sense of security and get sloppy. It's a good thing to really think about why you're making a specific memory allocation. Doing it manually forces you to consider whether there's a better option like static local storage or something. In addition, as long as you have a memory tracker implemented like my answer to 1, you never have to spend time worrying about memory. Just today, I forgot to free some memory and I was immediately notified when I ran my program and it printed a warning after I closed it with the leak. I cleaned it up in 5 minutes, end of story.
Edit: I forgot to mention, what's the one huge pro to using new/delete or malloc/free over smart pointers? Explicit control over when memory gets allocated and deallocated. If you use smart pointers everywhere, won't you eventually have the same issue as GC? They go out of scope at random points in your programs lifetime, and you get "random" lag spikes from memory cleanup. The whole point of using C++ is to have control over when this stuff happens! If you don't care when it happens, why not use the GC?
>
> This is super easy to solve in literally 300 lines of code or less. Just have a vector of your memory allocations in debug mode. Then call a custom allocator and free function that adds an allocation and removes an allocation from this list. At the end of the program, print out any allocations that are in the list that never got freed. You can use macros to print the exact file and line number that allocated the memory and then fix it very quickly.
I've got similar macros for C, but it is still less work to simply use valgrind.
Valgrind is easier, but for real-time applications I've found that it takes a huge performance hit for some reason, at least the last time I used it.
It does, because it runs the app in an emulator.
What I did in the past (around 2005) was to write replacement malloc/calloc/realloc/free calls, put them into a libmymalloc.so, and then set LD_PRELOAD to point to libmymalloc.so.
This way I could switch between my malloc routines and the system malloc routines without recompiling the program.
std::vector<int> allocation_counts;
auto get_mem(int n) {
auto* memory = your_malloc(n);
allocation_counts.push_back(n);
return memory;
}
This function has a trivial leak of "memory" that an RAII version wouldn't have, can you see it?...assuming that there's a matching free_mem(ptr, n) which undoes the actions of get_mem() of course. Also I'm ignoring the cofusing use of auto vs auto*, I haven't caught up with modern C++ enough to know whether this even makes sense ;)
...also, what would the better RAII version of this function look like?
std::vector::push_back can throw an exception, if for instance you just allocated 2 gigabytes of memory before and it does not have space to reallocate its internal storage:
std::vector<int> allocation_counts;
auto get_mem(int n) {
auto* memory = your_malloc(n);
// if this throws: the function exits there and "memory" is lost forever
allocation_counts.push_back(n);
return memory;
}
A simple way to "RAII-fy" it would be, assuming a `your_free`: // Typedef this somewhere
using your_ptr = std::unique_ptr<void, decltype([] (auto p) { your_free(p); })>;
auto get_mem(int n) {
// encapsulate in your_ptr
auto memory = your_ptr{your_malloc(n)};
// if this throws: memory's destructor is called which frees the memory
allocation_counts.push_back(n);
return memory;
}
Or if you want your custom RAII type for some reason: // memory.h or something
struct mem {
mem() = default;
// copy is not meaningful here
mem(const mem&) = delete;
mem& operator=(const mem&) = delete;
// move
mem(mem&& other): ptr{std::exchange(other.ptr, nullptr)} { }
mem& operator=(mem&& other) { ptr = std::exchange(other.ptr, nullptr); return *this; }
// allocation & deallocation
explicit mem(int n): ptr{your_malloc(n)} { }
~mem() { your_free(ptr); }
void* ptr{};
};
// your code
std::vector<int> allocation_counts;
auto get_mem(int n) {
auto memory = mem{n};
// if this throws: mem's destructor is called which frees the memory
allocation_counts.push_back(n);
return memory;
}But this is far fetched because people doing manual memory management don't usually rely on exceptions, and just abort the program when this happens, so the leak is pointless.
I think you really, really underestimate the amount of codebases where a large part of the code is "normal" C++ with exceptions and someone does one specific module with manual malloc / free out of cargo-culting some RAII fears.
Hell, even OP above talks about using vector along with their custom memory management scheme (which throws no matter if you use -fno-exceptions or not, as the actual "throw" call is in the precompiled libc++.so ; the code in <vector> just calls std::__throw_bad_array_new_length() indiscriminately : https://gcc.godbolt.org/z/fno8fYP7E )
I think the actual problem is that std::vector is doing its own memory management under the hood, which makes it hard to control what's going on from the outside. E.g. a design flaw of the C++ stdlib, not related to RAII.
plenty of cases where this isn't the case.
-fno-exceptions just means that you want to not be able to throw for your own code.
-fno-unwind-tables / -fno-asynchronous-unwind-tables makes your own code un-unwindable which is only safe if you are in entire control of the whole codebase, which is very rare in anything that is not simple games - no plug-in system, or any dynamic code loading or JIT'ing is safe as you don't know which user-provided function you call is going to throw.
Likewise, you could be in the opposite case: imagine you're making a C++ plug-in for, say, 3DS Max, where error reporting is done through exceptions ; even if your own plug-in code is built with -fno-exceptions and does not itself throw at any point, you still want exceptions that may be coming from a 3DSMax function you called to propagate to your code back to the exception handler provided by 3DS Max - and for this RAII of your own types is needed if you don't want your own resources to be leaked.
When I got introduced to C++ via Turbo C++ 1.0 for MS-DOS in 1993(!), I saw no longer a reason to do that other than being forced to deal with legacy C code.
Turbo C++ already came with collection classes, and RAII has been a thing since C++ no longer was named C with Classes.
Unfortunately re-education takes generations when a language happens to be copy paste compatible with C.
The problem I see personally with unique_ptr is that it's still doing small-scale memory management. It only removes some of the typing. (And for non memory management related tasks, it add significant amount of typing as well, compared to naked pointers).
Explicit memory management is very important when performance matters and must be predictable (for instance in "soft realtime" applications like games), you want to get the expensive memory management calls out of the hot paths as much as possible and the language must provide the tools to make this possible. When seen under this lense (that memory management should be explicit and not hidden from the programmer), there isn't all that much difference between C, Zig, C++ or Rust, all those languages provide the features for explicit memory management, even if some also provide optional features for automatic memory management.
https://en.wikipedia.org/wiki/Resource_acquisition_is_initia...
(Example: you'll understand this when you have experience with any graphics APIs, RAII wrappers don't really help you that much and will actually make your life way harder...)
My solution for general resource management, is to manage your resources manually on a central store (as you should be), get references to these objects using either pointers or indices, and use techniques like generational references to catch use-after-free scenarios at runtime. Makes things way simpler and performant.
Even something as simple as a constraint solver can easily find that memory management overhead dominates if they don't do something interesting. Work per deallocation can be extremely low in that domain, and anything like checkpointing can make the new memory access patterns sufficiently non-trivial that the GC isn't likely to be able to recognize it.
A lot of software probably doesn't care though I don't think.
Even C++ RAII is already an improvement over spreading the code with malloc/free, since 40 years by now.