Stupid Smart Pointers in C
blog.kevinalbs.com
blog.kevinalbs.com
C programmers are better off with either of these two techniques:
* Use __attribute__((cleanup)). It's available in GCC and Clang, and we hope will be added to the C spec one day. This is widely used by open source software, eg. in systemd.
* Use a pool allocator like Samba's talloc (https://talloc.samba.org/talloc/doc/html/libtalloc__tutorial...) or Apache's APR.
(I didn't include using reference counting, since although that is also widely used, I've seen it cause so many bugs, plus it interacts badly with how modern CPUs work.)
Interesting. I'm not very proficient in C, this looks like some sort of finalizers for local variables?
This being C, it's not without its problems. You cannot use it for values that you want to return from the function (as you don't want those to be freed), so any such variables cannot be automatically cleaned up on error paths either. Also there's no automated checking (it's not Rust!)
Note it's {...} scoped, not function scoped, which makes it more useful than Golang's defer.
[1] https://gitlab.com/nbdkit/nbdkit/-/blob/8b36e5a2ea331eed2a73...
if complex_nested_condition {
defer cleanup()
} defer if complex_nested_condition { cleanup() } else { noop() }
In Go, you could do: defer func(run bool) {
if !run { return }
}(condition)
Which admittedly wastes stack space with a noop function in the false case, but whatever.I feel like the number of times I've needed conditional defers is almost zero, while the number of times I've had to make a new function to ensure scoping is correct is huge.
Of especial note, 'mu.Lock(), defer mu.Unlock()' not being scope-based is the largest source of deadlocks in code. People don't use 'defer' because the scoping rules are wrong, code panics before the manual unlock call, and then the program is deadlocked forever.
I would say this is only half true. With some macro magic you can actually also return the values :)
https://github.com/systemd/systemd/blob/0201114bb7f347015ed4...
To be fair though, you probably meant without any such shenanigans.
It’s odd that the suggestion for a feature lacking in C is to use a non standard but well used supported path. c’s main selling point (IMO) is that it _is_ a standard, and relying on compiler vendor extensions kind of defeats the purpose of that.
Well, yeah...
How do you think Annex K got in?
Let's be honest, how many compilers are available, and how many of those would you actually use?
The answer isn't more than 4 and the 2 compilers you are most likely to use among those already support this and probably won't stop supporting without a good alternative.
I like standardisation, but you have to be realistic when it helps you without a large real cost other than fighting your ideals for getting this into the standard first.
But then again you are probably not doing a whole lot of heap management in embedded code.
It is quite common for this code to have all variables be global and just not have any heap allocations at all. Sometimes you don't even have variables in the stack either (besides the globals).
1) You add embedded linux
2) You have multiple CPUs that need to communicate with each other.
3) Both of the above (then the complexity skyrockets)
(this happens in modern cars)
So basically if it has a full LCD display (instead of 7-segment with a few extra toggle lights) I don't buy it.
It is actually quite refreshing code to read and write, yes the quality is often bad, but it often feels like school projects where they are small enough you can hold the complete system in your head. And it usually doesn't integrate with anything else besides the eletronics in the board
I would consider the union initializer change (require adding -fzero-init-padding-bits=unions for old behavior) much more hidden and dangerous, which is not directly related to ISO C23 standard.
I would count it as doing maintenance work for the upstream, kudos for doing this!
Huh. Wonder what the benefits of that are. Bools being ints never struck me as a serious problem. I wonder if this catches people who accidentally assign rather than compare tho....
I'd target the latest C standard and won't even care to know how many old, niche compilers I'm leaving out. These are vastly different uses for C and obviously your gaols drastically change your standard or compiler targeted.
The libraries you listed are all full of platform-specific code, and also have plenty of compiler-specific code behind ifdefs (for instance the stb headers have MSVC specific declspec declarations in them).
E.g. there is hardly any real-world C code out there that is 'pure standard C', if the code compiles on different compilers and for different target platforms then that's because the code specifically supports those compilers and target platforms.
Call it a posix extension, fair enough. But if your reason for writing C is that it’s portable, don’t go relying on non portable vendor specific extensions.
It's the same thing with the web and browser vendors, there's a constant mismatch, browsers propose and implement things and they may get standardized, and the standard dictates new requirements which might get implemented by all vendors.
The point of standardisation is defining behaviour for the things that are implemented as exploratory improvements and should be implemented on the more conservative compilers.
It's your choice whether to target the standard or a few selected compilers, there's a cost for both options between being late to improvements vs the possibility of needing to revisit your code around each of the "extensions" you decided to depend on.
If in certain projects portability is somehow of upmost importance, then any discussion around looking through the standard's black box to reach out for new stuff is kind of useless.
The *actual* power and flexibility of C lies in the non-standard, vendor-specific language extensions.
[0] https://thephd.dev/c2y-the-defer-technical-specification-its...
[1] https://thephd.dev/_vendor/future_cxx/technical%20specificat...
defer seems to be making significant progress (having a passionate and motivated advocate in Meneide, and a full TS)
(EDIT: I'm wrong, see reply)
Might have been the previous attempt from years ago, because being block scoped (unlike go) literally has its own section in https://thephd.dev/c2y-the-defer-technical-specification-its...
That is not, as far as I know, how __attribute__((cleanup)) works. It just invokes the callback when the value goes out of scope. So you can't have malloc return an implicitly cleanup'd pointer unless malloc is a macro, in which case you can do the same with a defer block.
I mentioned gigabytes because of how mine specifically worked. It allocated chunks in powers of 2, so there was some % of memory that wasn't being used. For instance, If you only need 20 bytes for a string, you got back a pointer for a chunk of 32 bytes. Being just a game, and side project, I never gave it much thought, so I'm curious to hear your input.
But how did you determine what you could re-use? That's the hard problem, one that's equivalent to calling free() at the right time.
It's also a way, I suppose, to ensure your program has the memory it needs, depending on how you design it. Mine basically allocates all the expected memory I need at the start, while still being able to grab more if it needs. This was more of an issue back when you shared a server with lots of others.
(In fact, most libcs let you just redefine the malloc(), free(), and realloc() functions to point to your own allocator instead of the default one, so you don't have to rewrite all the functions calling them. E.g., the mimalloc allocator [0] can be configured as a drop-in replacement for the libc allocator.)
Re-reading your comment, about the "not really freeing anything" -- I beg to differ, as when you do a real free(), it frees up memory for any program to use. As I already mentioned one of the benefits is once you have control of who can use that memory and aren't risking it not being available - disregarding some program that constantly grows in memory.
This is an extremely naive description of not only what the libc free() function does, but also of how memory management works in general.
The free subroutine deallocates a block of memory previously allocated by the malloc subsystem. Undefined results occur if the Pointer parameter is not an address that has previously been allocated by the malloc subsystem, or if the Pointer parameter has already been deallocated. If the Pointer parameter is NULL, no action occurs.
All that is guaranteed, is a deallocation. Not how and when that memory will become available again.
Your "more detailed" understanding will break, and will cause headaches, on platforms ypu're not used to.
overcommit would like to have a word
libc already does that. What is it that yours is adding?
I'd say 25 years ago you could write your own naive allocator, make just a couple of assumptions for your use case, and beat libc. But no more.
One of the selling points of Java in the 90s was the compacting part. Because in the 90s fragmentation was a much bigger problem than it is today. Today the libc allocators have advanced by maybe tens of thousands of PhDs worth of theory and practice. Oh, and we have 64bit virtual address space, which helps with some (but not all) of the problems with memory fragmentation.
See this post from Ian Lance Taylor about why Go didn't even bother with a compacting GC: https://groups.google.com/g/golang-nuts/c/KJiyv2mV2pU?pli=1
Have they, though? Looking at the blame on glibc's malloc.c, the most substantial change in the last 10 years has been the addition of memory tagging. Apart from that, I mainly just see a bunch of small tweaks to the tcache, fastbins, etc., and the basic logic is broadly the same as it was 20 years ago. Similarly, the MallocInternals wiki page [0] appears mostly as it did in 2016, except for some more explanations of the tcache added in 2018.
I can easily believe that lots of work has been done on malloc/free-style allocators, but from what I've seen of most libc allocators, they hardly stand at the forefront of this work. (Except for the ones that just vendor some standalone allocator library and keep it up to date. But that describes neither Windows nor Linux.)
And of course, if you're writing your own allocator, you can relax some of the constraints, e.g., you can drop most of the multithreading support and not have to fool around with locks and shared arenas and whatnot.
For fun I wrote up a very basic memory manager and benchmarked it against free/malloc. On my Mac (M3) the manager was 3x faster. On a random kubernetes pod (alpine) it was 33x faster. Performance increases as memory size goes up.
The "tl;dr" summary is that `free_on_exit()` replaces the caller's return address with a trampoline that calls a `do_exit()` function that frees all the memory marked to be freed "on exit", and it uses its own stack of cleanup handler closures so-to-speak. When returning all the allocations are freed and then the original return address is returned to.
Really, don’t do this, it’s a portability nightmare. If you’re going to write C stick to code that easy to run under MSVC.
mingw is radically more problematic than MSVC. Don’t use mingw.
Hey, I wrote this GCC plugin to automate reference counting in C. Any bug reports welcome.
Can you expand on this?
In reality, your milage will heavily vary. If you don't have contention (you don't commonly share objects across multiple threads concurrently), it's likely that reference counting will perform very well. Whether this is the common case really depends on the kinds of software you write.
I've seen this happen in cases arrays of objects are allocated, ie one per thread, and then handed to a thread pool to work on.
Even if heap allocated, if the object is just a reference counter and a few pointers, the memory allocator can fit several of them next to each other causing them to share cache lines, which causes the performance issues with atomic operations.
Depends on implementation of things of course, but can be a pitfall.
Objects often come in batches of the same type and similar maximum lifetime, so let's make use of that.
Instead of tracking the individual lifetimes of thousands of objects it is often possible to group thousands of objects into just a handful of lifetime buckets.
Then use one arena allocator per lifetime bucket, and at the end of the 'bucket lifetime' discard the entire arena with all items in it (which of course assumes that there are no destructors to be called).
And suddenly you reduced a tricky problem (manually keeping track of thousands of lifetimes) to a trivial problem (manually keeping track of only a handful lifetimes).
And for the doubters: Zig demonstrates quite nicely that this approach works well also for big code bases, at least when the stdlib is built around that idea.
A lot of problems break down to:
* we need this effectively forever (i.e. until config reload)
* we need this very briefly when processing a task or request
Sometimes you need a cache that has intermediate lifetimes, but that is a much smaller problem to deal with, and you can often cope with manual memory management for that
Hook any file handles and other resource cleanup functions into the same pools and you have a pretty easy life.
[0] https://thephd.dev/c2y-the-defer-technical-specification-its...
(not super experienced Go developer)
Adding `func() {}()` scopes won't break existing code usually, though if you use 'break' or 'continue' you might have to make some changes to make it compile, like so:
https://go.dev/play/p/_Gq4QYtyMmp
see, no other issues, works exactly like you'd expect
[1] https://thephd.dev/c2y-the-defer-technical-specification-its...
edit: TIL: Looks like the [drop trait][1] is an opt-in destructor that the Rust compiler understands.
* cannot return errors/throw exceptions * cannot take additional parameters (and thus do not play well with "access token" concepts like pyo3's `Python` token that proves the GIL was acquired -- requiring the drop implementation to re-acquire the lock just in case)
I think `defer` would be a better language construct than destructors, if it's combined with some kind of linear types that produce compiler errors if there is some code path that does not move/destroy the object.
It is first argument to all functions.
If no errors, function proceeds - if error, function instead simply immediately returns.
When allocating resource, resource is recorded in a btree in the state.
When in a function an error occurs, error is recorded in state; after this point no code executes, because all code runs only if no errors.
At end of function is boilerplate error, which is added to error state if an error has occurred. So for example if we try to open file and out of disk, we first get error "fopen failed no disk", then second error "opening file failed", and then all parent functions in the current call stack will submit their errors, and you get a call stack.
Program then proceeds to exit(), and immediately before exit frees all resources (and in correct order) as recorded in btree, and prints error stack.
It would also need to be a function that will truly be implemented as one following the ABI, which usually happens when the function is exported. Often times, internal functions won't follow the platform ABI exactly.
Just changing the compiler version is probably enough to break anything like this.
Save the return address highjacking stuff for assembly code.
---
Meanwhile, I personally have written C code that does mess with the stack pointer. It's GBA homebrew, so the program won't quit or finish execution, and resetting the stack pointer has the effect of giving you a little more stack memory.
I see undefined behaviours.
they walk
they talk
they [0?W0OF??0?r??reeBSD
they don't know they're undefined.No, it does not. Smart pointers are useful to help model lifetimes and ownership, but the real killer feature is RAII. Add that to C (standardized) and you can make smart pointers, and any other memory management primitive you need. Smart pointers are not a solution, they are one of many tools enabled by RAII.
2) I recently discovered the implementation of free_on_exit won't work if called directly from main if gcc aligns the stack. In this case, main adds padding between the saved eip and the saved ebp, (example). I think this can be fixed some tweaking, and will update this article when it is fixed.
I do not believe the article was updated, suggesting that the "tweaking" was far more complex than the author expected...
...which doesn't surprise me, because the overall tone is one of a clever but far-less-experienced-than-they-think programmer having what they think is a flash of insight and realizing thereby they can solve simply a problem that has plagued the industry and community for decades.
One person's hack-that-should-be-avoided-at-all-costs can be someone else's secret sauce.
the way i do this in C looks like
initialize all resource pointers to NULL;
attempt all allocations;
if all pointers are non-NULL, do the thing (typically calling another routine)
free all non-NULL pointers
realloc(ptr, 0) nicely handles allocations and possible-NULL deallocationsfor example, the object could be managing an open file, or an open socket
void my_thing_free(MyThing *thing) {
fclose(thing->file);
free(thing);
}
assuming an associated "my_thing_new" that only returns a valid pointer when both the allocation and the fopen succeeded.A lot of the motivation behind inventing golang seems to be to remove C++ toys, because it's too hard to get people to not use them if they're there.
C is not 100% compatible with C++.
There's a whole heap of incompatibilities that you can hit, that will prevent a lot of non-trivial C programs from compiling under C++. Things like character literals being a char in C++ and an int in C. Or C allowing designated initialisers for arrays, but C++ not.
I haven't used it personally yet, but it addresses the same issue with a different approach, also related to stack-like lifetimes.
I've used simple reference counting before, also somewhat relevant in this context, and which skeeto also has a nice post about: https://nullprogram.com/blog/2015/02/17/
Use C++ or __attribute__((cleanup)) instead.
Or maybe the asinine "thing", is some folks lack of text comprehension skills that can't distinguish an experiment from a best practice recommendation despite a title and content that clearly does not invite that.