TigerBeetle uses Deterministic Simulation Testing to test and keep testing these paths. Fuzzing and static allocation are force multipliers when applied together, because you can now flush out leaks and deadlocks in testing, rather than letting these spillover into production.
Without static allocation, it's a little harder to find leaks in testing, because the limits that would define a leak are not explicit.
Here's an overview with references to the simulator source, and links to resources from FoundationDB and Dropbox: https://github.com/tigerbeetledb/tigerbeetle/blob/main/docs/...
We also have a $20k bounty that you can take part in, where you can run the simulator yourself.
The techniques mentioned above will (perhaps surprisingly) not eliminate errors related to OOM, due to the nature of virtual memory. Your program can OOM at runtime even if you malloc all your resources up front, because allocating address space is not the same thing as allocating memory. In fact, memory can be deallocated (swapped out), and then your application may OOM when it tries to access memory it has previously used successfully.
Without looking, I can confidently say that tigerbeetle does in fact dynamically allocate memory -- even if it does not call malloc at runtime.
I'd be curious if you have any resources/references on what is considered good practice in that now, then.
It's been a long time since I did much ops stuff outside a few personal servers, so it may well be my background is out of date... but I've certainly heard the opposite in the past. The argument tended to run along the lines that most software doesn't even attempt to recover and continue from a failed malloc, may not even be able to shut down cleanly at all if it did anyway, and the kernel may have more insight into how to best recover...so just let it do its thing.
It would be ideal to not tickle the OOM killer, but it does happen. A great example would be Redis, using bgsave. There's a lot to criticize about redis and bgsave and I don't mean to defend its architecture. Its behavior is fairly extreme so it provides an example useful for illustration. Because Redis forks its entire in-memory state and writes itself to disk, it will sometimes appear to the system as if it has doubled in memory use. It's a huge, sudden memory pressure event, exacerbated by any writes forcing CoW allocations while the bgsave runs.
Many other app servers or database systems can have similar sudden memory pressure. You basically don't ever want a primary app on a system to be killed, so it's usually best to just disable OOM killer entirely on these processes.
There's often no failed malloc() in these situations. Often you will see OOM in situations where no malloc() has ever failed, because malloc() is just allocating address space without any attached memory, which is nearly free, and which almost always succeeds. The failure will occur later when the allocated page is first written to -- causing a page fault and an actual allocation to happen. There's no associated function call or system call. Page faults are triggered simply by accessing memory. This is why the OOM killer exists, as there's no function available to return a failure to when memory can't be produced. Such is the idiosyncratic behavior of lazily-allocated memory in modern virtual memory systems.
tldr: Malloc never fails, because malloc allocates address space not memory. Memory writes trigger failures, because writes create page faults and trigger the actual allocations. But the failures often express behavior elsewhere, via the action-at-a-distance magic of the OOM.
We also target Linux with the same codebase, but the environment is a bit different. On Linux we're using a system slice and we're not doing minidumps.
1. Bound every whatever that requires memory to some value N
2. For all whatevers that can have a lifetime for your entire program, do that
3. For any whatevers that can't have a lifetime for your entire program (e.g. for a database server like TFA, a connection object would qualify), make a table of size N
4. When you need a whatever (again using a connection object as an example, when listen() returns) grap an unused entry from the table. If there is no unused entry do the appropriate thing (for a connection, it could be sending an error back to the client, just resetting the TCP connection, blocking, or logging the error and restarting the whole server[1]).
5. When you are done with a whatever (e.g. the connection is closed) return it to the table.
Those are the basics; there are some details that matter for doing this well e.g.:
- You probably don't want to do a linear-probe for a free item in the table if N is larger than about 12
- Don't forget that the call-stack is a form of dynamic allocation, so ensure that all recursion is bounded or you still can OOM in the form of a stack overflow.
I should also note, that for certain values of N, doing things this way is normal:
- When N=1 then it's just a global (or C "static" local variable).
- When N=(number of threads) then it's just a thread-local variable
1: The last one sounds extreme, but for a well understood system, too many connections could mean something has gone terribly wrong; e.g. if you have X client instances running for the database, using connection pools of size Y and you can have 2XY live connections in your database, exceeding N=2XY is a Bad Thing that recovering from might not be possible
The approach I've seen for larger fixed-sized tables is a free list: when the table is allocated, also initialize an array of the slot numbers, with an associated length index. To allocate, pop the last slot # off the array and decrement the length index. To free, push the slot # onto the end of the array and increment the length index. Checking for a free slot is then just testing for length index > 0.
Writing reliable systems means spending a lot of time thinking about what and how failures can occur, minimizing them and dealing with the cases you can't get rid of. Simple is nearly always better.
It also helped to be an EE by education and not a CS major. CS majors can get blinded by their abstractions and forget that they're working on a real piece of HW with complex limitations which require design tradeoffs.