A Defer Mechanism for C
gustedt.wordpress.com
gustedt.wordpress.com
One of the key interest of C compared to more recent languages is that everything is explicit. With a finger you can follow the code as it runs and know exactly what is going on and when exactly.
At the opposite, there is c++ that does a lot of things automagically. And it is often hard to understand why you suddenly get a segfault out of blue, just because a destructor was randomly called at the wrong time.
There's nothing implicit about defer.
That's quite implicit. You can no longer reason about a local piece of code; you now have to know its lexical nesting up to top level to see if it's inside a guard block that might trigger hidden behavior.
Plus, to handle the free or the leak if you had forgotten to free a resource at the exit point would also require to "know its lexical nesting up to top level".
Between goto and longjump and co, C has much worse non-local behavior than defer.
By your definition only [1] "come from" would be non-local.
You can't goto out of a function, and you know there's exactly one such label inside it. If goto isn't local, then neither are function calls, since the function could be defined anywhere in the codebase.
You can with a goto expression and a label address available - though the behavior is undefined in C, so bets are off.
And you can with longjump/setjump more explicitly.
Anyway, what could possibly happen at the end of a while loop, besides going back to the begining for the test?
Besides, if you have a bug in your code, you will have to look at the whole block anyway.
You have to check every line of a while loop to know what’s going on. Heck, another thread could hold a pointer aliasing the loop variable and then mutate it, causing the loop to terminate for no obvious reason.
Which C doesn't have.
They are no less explicit than the actions performed at the end of iterations of for or while loops. In general, for C, understanding the behavior of code requires understanding “is it in a block, and if so what kind of block”; guard blocks would be not generally different.
However, for some things you can only free/return resources if you successfully created that resource. At which point you would need to use something like a stack.
It's not really random, but it's quite hard to follow.
(std::shared_ptr should be used only when absolutely required and should be kept under close watch the whole time.)
Should I be alarmed?
There’s a good chance that most of the shared_ptr’s can be replaced by unique_ptr, which has less runtime overhead. More importantly, it documents the intent of the programmer regarding ownership semantics, and the timespan during which the object should be remain valid.
If you have a situation where a value needs to be shared between multiple other values, but also the value is trivial enough that it doesn't matter when it's destroyed or what it does when it's destroyed, but also the value is not so trivial that you can copy it instead of using a refcounting pointer, then std::shared_ptr is fine.
Unless you know you must share objects with multiple potential owners, though, std::unique_ptr is a lot wiser.
>but also the value is trivial enough that it doesn't matter when it's destroyed or what it does when it's destroyed"
that precludes the value from having other refcounting pointers that could then create cycles and leak.
But more that there is so much magic and abstractions that it is very hard for a dev to have a clear view of what is going on and what to expect. He has to 'guess' instead of just read the code. For that, I guess that it is similar to the current question of accountability of decisions made with deep learning algorithms.
As an example of my point, I would refer to the 'garbage collection' issues of a language like 'java'. GC will happen at a logically defined point like 'dirty mem > 100m' but from the developer point of view, his logic could suddenly lag unexpectedly in a middle of a simple operation because the GC was triggered by internal magic. It is very hard for a dev to be able to determine the memory usage at different points of the code and so have a certitude of when this operation could happen.
Two languages could hardly be less alike than Java and C++ - the latter is not garbage collected and where destructors are called is completely predictable.
It is more deterministic in C++, but you can still call exit(). (which is even more straightforward than the refcount cycle)
in the specific case of shared_ptr, which are a small part of codebases, if they are even used - for instance an immense amount of C++ GUI programs use Qt which doesn't use shared_ptr-like ownership semantics but instead a tree-of-objects model which does not have this issue. In contrast in Java / C# any object that has a reference to another is at risk.
when I put some object on automatic storage in C++ I do so because I explicitely want it to go away when its scope is left (either by reaching } or through an exception)
What is "I am thinking of using C?"
With macros this is not exactly true. I gesture towards the GObject system for an extremely complex situation in big important production software.
Oh yeah?
A possible explicit solution I can think of is to introduce a new function attribute that gives the name of their ‘cleanup’ function (so that the compiler would know fopen needs a fclose, for example) and a compiler that uses these attributes to issue a warning if a function has a path that calls a function and doesn’t either call its cleanup function or returns its result.
I don’t know whether that would cover all bases, though.
That said, it's still debatable if it's useful, given that you can achieve the same thing with the
struct some_resource resource;
do {
resource = allocate(...);
if (!resource) break;
} while(0);
if (resource) dispose(resource);
I can see it as a good thing because you have the dispose statement next to the allocate statement, which makes the logic easier to follow, but the implementation may have caveats which make it actually harder to reason about, e.g. see the other thread about capture value vs capture reference - C will most probably need to capture by reference, which means that modifying "resource" later on changes the meaning of the deferred statement.However, this is not without precedence in C. For example, just look at the for loop:
for (clause1; expression2; expression3) statement
expression3 is executed after statement.
I think the best defense of this syntax is that it makes writing basic iteration a bit nicer without having to add boilerplate (the iconic `for (i = 0; i < n; i++)`) but then I would argue that the real problem is that C is severely lacking in the iteration department and this is a rather obvious hack (that languages like Javascript felt the need to copy wholesale, for some insane reason).
Just like every C feature: if you know how it's implemented (and optimized), you will know what is going to happen. All big C compilers already support stack frames and unwinding, so it's not even entirely novel functionality.
While you may object it's not entirely obvious how unwinding is going to be implemented, OTOH in functions with complex control flow it can be easier to understand which `defer`s are going to be run, as opposed to following nested `else` statements, or reasoning about the program state from all `goto cleanup` locations.
I would suggest functions like atexit and pthread_cleanup_push as well established counterexamples. Granted this is not a perfect comparison because these are implementable without extending the c language; however, I think they have the same general idea of "defering" cleanup. I think the proposed defer functionality is actually more readable because the defer command will be written much closer to where it will be exited. Compare this to atexit() which may be put anywhere.
This has not been true in many years. C appears to be such a language, but optimizing compilers have learned how to find and optimize undefined behavior in code that most C programmers don't realize is unsafe. As a result there can be a considerable gap between the code as clearly intended, and the code that will be generated.
See https://blog.regehr.org/archives/213 for more on this. Including real world examples of things like validation checks being elided by the compiler, leading to real-world vulnerabilities in programs which clearly have checks to avoid exactly those vulnerabilities.
Probably. It's one of Go's lesser ideas.
C already has a "defer" mechanism in "exit", to close out files and such. Of course, the final I/O status gets lost.
I don't agree! There's a lot of behavior that's implicit and you essentially need to internalize the C standard/compiler behavior to follow. Weak typing, for instance; defaults for memory access/fencing, handling faults, runtime semantics with respect to initialization of the process, cleanup via atexit, and signal handling.
Memorable, sure, but hardly explicit.
Could u point to an example of programmers doing this ?
Although I guess Cheney on the MTA counts.( https://dl.acm.org/doi/10.1145/214448.214454 )
If anything RAII is unfit for C (which doesn't have classes), and in itself, a kludge (it's an idiom, not a language feature).
So it's immediately better compared to the current C situation.
As for compared to RAII? Well, thats one failure mode for defer (forgetting it), whereas there are dozens of ways to mess RAII...
In comparison to RAII, “defer” as a language feature is a kludge. Mainly because RAII completely solves the problem of correct resource cleanup for API users while “defer” does not.
So I don't see the comparison...
Destructors probably cannot be usefully imported into C while preserving the simplicity of the language: then you'll want at least unique_ptr to manage your memory with destructors, then probably shared_ptr, and some ownership semantics as you pass those around, et cetera. In that sense, defer strikes a balance between simplicity and usefulness. However, there is still a comparison.
RAII is the best for ensuring that things get cleaned up even if it does lead to more boilerplate to add the pattern to things. But even .NET's IDisposable feels better than defer. It's established practice to implement it if your class contains something that must be disposed. And it's easy for tools to check if you haven't disposed of something that implements IDisposable.
All that said, I find defer to be better than nothing. In C, the best I can do for cleanup is having all the cleanup functions called at the end with labels at the various points I need to jump to.
In terms of usability, destructors allow for resource management that is far easier at the call site than any other language. "with" blocks such as in python require the call site to be modified to include the explicit time of destruction, whereas C++ gets that for free from its existing scoping rules. "try/finally" such as in Java also requires modifying the call site, but at least has reasonable default behavior if accidentally omitted. "defer" feels like it has the worst of both worlds. It requires the call site to be modified when a resource is owned, and also requires the caller to know what the corresponding cleanup function is for any allocation.
For example, I cannot write:
lock (mutex);
Expecting to declare an anonymous lock with the mutex passed in, as this gets parsed as a declaration of a lock called mutex (hopefully lock doesn’t have a default constrictor and I at least get an error). Instead I have to bake my lock: lock l(mutex);
Often I have objects that only exist for RAII and it’s a little bit annoying to have to give them dummy names.You can do
lock(mutex), some-lock-protected-expression;
but it is a bit too cute and limited.On the other hand, being able to name the RAII object is often very useful for example if you need to dismiss them early, which happens often in transactional code (unlocking a mutex before the end of scope is a relatively common occurrence).
Also specifically for mutexes, a nice pattern is to bind object and mutex together to guarantee that the object is always used with the lock:
sync<my_object> foo;
{
auto l = foo.lock(); // starts the critical section
l->do_something()
l->do_something_else();
} // it ends here
// or if you only need to hold the mutex for a single call:
foo->do_something(); // critical section lives only for the full expressionI didn't say it wasn't. Being able is perfectly fine, being forced is not. As I mentioned in another comment, the problem isn't the extra typing, the problem is that it requires us, the programmer, to remember to name it even when we don't feel we need and to remember that "foo (bar)" isn't calling the constructor of foo and passing in bar, its calling the default constructor and declaring bar. That's too many gotchas and rather error prone!
I believe there was a cppcon talk where the speaker said that it was a very common bug in Facebook, despite that they have linter rules to catch this case. Its also a rather insidious bug, because the code will run seemingly normally, just... it never actually locks anything. A rather hard problem to debug too.
Actually, I think it's worse than that; IIUC, it will block until the mutex is unlocked, then run. Under light load, it will get away with this, but if the load is heavy enough, it will either guarantee that every process waiting for the lock runs at once, trampling each other's work, or almost-serialize them, making the race condition even more intermittent than if they didn't lock at all. Which failure mode you hit depends on how your scheduler works.
#define LOCK(mutex) lock lock##__LINE__(mutex)
Or even C compatible scoped locks: #define SCOPED_LOCK(mutex) \
for (int i##__LINE__ = lock_mutex(mutex), 1; \
i##__LINE__ --; \
unlock_mutex(mutex))
Use like SCOPED_LOCK(foo->mutex) {
do_stuff();
}The problem I'm describing is less about the effort of adding a variable name -- that's really not a big deal -- its that its super error prone to remember and these bugs slip through all the time. Requiring a variable name (either explicitly or through a macro like yours) is error prone.
With lock() being a function taking a lambda.
lock([]{
// code locked with anonymous mutex
});RAII doesn't help with a generic pair of init and deinit functions, such as malloc() and free() for example, unless you wrap your mallocs into "memory objects" or something. You can't do anything with RAII unless you wrap your stuff into objects which just forces the OO crap on everything regardless of whether it's a good fit for the OO paradigm.
But even in RAII, classes are merely just a way to bind init and deinit together. As a result, you can't "forget" to free the resource. (Unless you use the new operator...) But this is merely an interface issue: defer could be of the form
void *block = defer malloc(SZ) with free(block);
or something similar that requires the init and deinit calls to be part of the defer expression itself.Nevertheless, a defer is clearly a useful language feature that C actually lacks. Currently C doesn't offer the code any attaching points to the lexical scope of the program. You can fake it with a special for(;;) statement but it's a kludge and even by wrapping it into a macro it's very hard to make it a generic solution.
people have been writing RAII scope_exit classes for at least 20 years to do exactly this.
> You can't do anything with RAII unless you wrap your stuff into objects which just forces the OO crap on everything regardless of whether it's a good fit for the OO paradigm.
wrapping something in an object does not make it OO. For example there is nothing object oriented about std::unique_ptr.
c'mon, here's the article from 2000 that introduces ScopeGuard : https://www.drdobbs.com/cpp/generic-change-the-way-you-write...
That's Microsoft's fault; it's 2020 already, and unless things changed since I last looked, their compiler still doesn't have full support for the C99 standard.
Even then, some of the C99 syntactic features have been somewhat widely adopted (at least when not compiling for Windows); off the top of my head, we have "//" comments, declarations in the middle of a block, and designated initializers.
Everything that was moved into optional in C11, like VLAs, is not planned to ever be supported.
I see people tend to write C89+designated initializers but this is a code smell IMO. Initial data should always be 0, especially data that lives in BSS/data section.
Looking at the article, this specific implementation doesn’t quite go all the way, requiring a special guard {} block instead of working anywhere. A better implementation, working in any context, would be easy but has to be baked into the compiler.
Edit: just saw the proposal for inclusion into the C standard. It is unfortunate that they are considering 1) requiring a guard block and 2) deferring clean up to the end of the function instead of end of the scope. #1 is needless syntax and #2 would make the feature useless within loops by causing explosions and contractions in memory usage.
Without one, you couldn't 'defer' to the end of a containing block from within an if/while/do/for/etc. block.
When/where the deferred code is executed has to be specified some way so an explicit marker for the defer 'scope' is surely required without severely limiting the utility of the feature.
guard {
void *ptr = malloc(12);
if (ptr) {
defer free(ptr);
// Use ptr
}
// Use ptr some more
} // free ptr here
Without guard keyword: {
void *ptr = malloc(12);
defer { if (ptr) free(ptr); }
if (ptr) {
// Use ptr
}
// Use ptr some more
} // free ptr hereIn a more realistic case, say some object was handed off so it shouldn't be cleaned up anymore, you can just add a should_cleanup flag and check it in the defer block. There's really no need for guard.
Their other example for guard, running defer in a loop, requires closures. This means it requires dynamic memory allocation and so the complexity goes through the roof.
Dlang already has scope guards (Appendix C in the OP) https://dlang.org/spec/statement.html#scope-guard-statement, and does so at block scope. For some code it's a more elegant solution (keeps acquire and release code close together), without introducing the usual C++ RAII issues.
This proposal contains a specification for complete stack unwinding in C. It doesn't just specify defer, but also panic and recover, which jumps between functions and cleans up guard blocks across stack frames. It's essentially exceptions for C. This is frankly horrifying and I can't believe the C standards committee is entertaining this.
They claim that this is a separable feature from defer, so maybe they intend for defer to be mandatory and panic/recover to be optional. But much of the design of defer is to support their panic/recover mechanism. This makes it much more complicated than attribute((cleanup)). If they want defer to be taken seriously, they should move all of the panic/recover stuff to a separate proposal. I suspect if they did this, a lot of the design recommendations they've made for defer wouldn't make sense on their own.
Also, the defer syntax they're proposing seems nicer in terms of syntactic sugar (no need for a stub function), but at the same time the semantics seem ... weird. Having it tied to a variable makes more sense, and sidesteps a whole bunch of weird situations (like loops with defer statements).
Hey, standards committee’s gotta eat.
'defer'? I occasionaly use it in Swift to clean up resources; s’okay there, I guess, though I’m not convinced it’s better than Python’s 'with' block. But in C?
One of C’s few distinguishing strengths is that the language is relatively† small and stable, and well understood. For that kind of cleanup there is already 'goto', which again is small, stable, and well understood. I just used it for that the other day: it works, it’s fine; I’m a grown-up.
Yeah, sure, 'defer' is “safer” and “more elegant”… did we mention this is C? That ship sailed fifty years ago. Don’t try to make C into something it’s not: that’s C++’s job so go fill your boots there instead.
I just posted Tony Hoare’s excoriation of ALGOL68 the other day, but clearly it’s needed again:
http://zoo.cs.yale.edu/classes/cs422/2011/bib/hoare81emperor...
Simplest solution: track down the ruddy C standards committee and beat them in the head with a leather-bound copy of Zawinski's Law (wrapped around a large gold brick), till either they’re dead or they leave C be. It does what it was designed to do, and that along is reason enough not to dick with it just because they’re bored and struggling to justify their continued existence.
The only thing C needs to do is keep on working. That will only get harder the more crap they pile on top. A good artist knows when to stop.
Which brings us to…
“panic/recover”
K&R give us strength! Tell these frustrated wannabe language designers to go make their own damn language, instead of screwing up someone else’s!
Okay, now I’m done. And get off my lawn!
Speaking from experience, GCC/Clang extensions are pretty useful when you can afford to use them and it would be nice to have some of that stuff standardized because a lot of code I see in the wild is already using them. The time for warning about C becoming a mess of incompatible extensions has already passed unfortunately, and the only group that can solve this is a standards committee.
This is exactly why everyone hates using modern C++. Their standards committee has run amok with every fancy feature that every new language comes out with in the past two decades. And what has that wrought? A huge language that no one fully understands and is hard to read unless you happen to be the guy that wrote it and time since writing < two weeks.
C is simplistically beautiful. I can write it and understand exactly what the assembly will look like on the backend of the compiler (with the exception of nasty macros). Don't mess with that. If it doesn't feel like it belongs, that's probably because it doesn't! Know when to stop.
speak for youreself ?
> Their standards committee has run amok with every fancy feature that every new language comes out with in the past two decades. And what has that wrought?
given the amount of languages that reimplement C++ features (and sometimes have to be dragged by the feet to get them, e.g. generics and interface methods in java), I'd say that people who don't understand c++ are doomed to reinvent it :)
I can because C already includes setjmp/longjmp as an exception mechanism that gets used in plenty of real code. The problem here is that those don't unwind the stack so they would break and cause memory leaks when using defer statements.
The C compiler vendor for some embedded CPU with homegrown C compiler for their in-house OS also has a seat at ISO table.
While people here might not care, ISO does.
Don’t go complaining that C++ is “too complicated” and then be hauling its complexity into C, because all you’ll end up with is a bloated schizophrenic mess that is neither a good C nor a good C++ [alternative].
guard {
for (int i = 0; i < n; i++) {
defer foo(i);
}
}
Now the compiler has to:1. implement some side of capture/closure mechanism to keep all the 'i's to the end of the guard block
2. do dynamic allocation to store the closures so they can be executed at the end did the scope
1 seems like too much work for such a feature, and 2 is a massive no. Implicit dynamic allocation, in C?
And all of this for nothing. The guard syntax doesn't give any reasonable benefits. They should have just kept it simple; defer happens at the end of the scope, and it just takes 'i' by "reference". It's a shame because it's a feature I would really like to have.
Still, I agree that this is not "nice and clean". Ultimately it doesn't feel like something that belongs in C.
guard {
int i;
for (i=0; i<n; i++) defer foo(i);
}
I would expect foo(n) to be called n times. That is, variables won't be captured, it would work literally as if you wrote "foo(i)" n times at the end of the guard block. It's a footgun, but everything in C is a footgun and this behaviour is not surprising to me as a C developer.So I expect the implementation would be
1. If a defer block references variables which do not live to the end of the guard block, the program is malformed (compiler error).
2. Compile each defer block as if it was written at the end of the guard block. What order the defer blocks are stored in is implementation-defined
3. Put a pointer on the stack each time a defer statement executes pointing to the compiled defer statement.
With this each defer statement is just a few bytes of overhead added to the stack and you can do it in a loop, though it may not do what you expect if you come from Go. For the example I expect the compiler to emit code similar to:
struct defer_node {
void **label;
struct defer_node *next;
};
struct defer_node *defer_start=0;
int i;
for (i = 0; i < n; i++) {
struct defer_start *defer_new = alloca(sizeof(struct defer_node));
defer_new->label = &defer_stmt0;
defer_new->next = defer_start;
defer_start = defer_new;
}
while (defer_start) {
goto *defer_start->label;
defer_stmt_exit:
defer_start = defer_start->next;
}
return;
/* or break; or continue; whatever is appropriate for the enclosing block */
defer_stmt0:
foo(i);
goto defer_stmt_exit;
This makes defer more or less just a mechanical transformation of code that can be expressed with existing C primitives and without requiring dynamic memory. I think it would be a good addition to C's structured programming elements.This still makes it really easy to overflow your stack, especially if for example your loop count above can come from user input. This is dangerous for all the same reasons that variable-length arrays are dangerous.
But the order they're executed in would have to be defined, lest we build ourselves another major bug generator.
for i := 0; i < 10; i++ {
func(i int) {
defer foo(i)
// other code
}(i)
}This case of the defer in the loop is frequently cited, probably because it is a problematic case. However, I looked at a lot of real code and the only case I found of resources being allocated in a loop they were allocated at the beginning of the loop and deallocated at the end. Another option we are considering is to use the scope for the guarded block. In this case, deferred statements would be executed at the end of each iteration of the for loop which would be ideal for this sort of code. For example, you could rewrite this function using defer:
https://github.com/openssl/openssl/blob/a829b735b645516041b5...
like this:
for (;;) {
raw = 0;
ptype = 0;
i = PEM_read_bio(bp, &name, &header, &data, &len);
defer {
OPENSSL_free(name);
name = NULL;
OPENSSL_free(header);
header = NULL;
OPENSSL_free(data);
data = NULL;
}
...
} else {
/* unknown */
}
} // end for loop, run deferred statementsI would strongly agree that the dedicated guard{} block is a bad idea and this should just tie into the innermost scope. I see what this is trying to do ("if (...) { foo.x = malloc(); defer free(foo.x); }") but you don't solve a UI problem (as in, user interface for the programmer to their code) by adding more weird UI.
("Worst" example for shooting yourself in the foot: defer inside macros. Programmer then forgets that the macro contains a defer, and the defer defaults to the function-implicit outer guard block. But it's really in a nested loop. That's gonna be a fun week of debugging... much less of a risk when you have the guarantee in terms of innermost scope.)
void reverse_bitstream(void * read_ctx, void * write_ctx) {
guard {
while(has_more_bits(read_ctx)) {
int b = read_bit(read_ctx);
if (b == 0) defer write_bit(write_ctx, 0);
else if(b == 1) defer write_bit(write_ctx, 1);
}
}
}To have a good example where defer inside a loop actually does something, you'd have to capture the value of i. We also have a macro for that in the reference implementation, but that has a much more specialized scope of use.
Good to see it standardized of course...
#define DEFER_START(scope_name) void * scope_name = &&scope_name##_end;
#define DEFER(scope_name, iter_name, func) void * iter_name = scope_name; { \
scope_name = &&iter_name ##_exe; \
if (0) { iter_name##_exe: {func} goto *iter_name; } }
#define DEFER_END(scope_name) scope_name##_start: goto *scope_name; scope_name##_end: {}
#define DEFER_EXE(scope_name) goto scope_name##_start;
...
{
DEFER_START(scope);
void * p = malloc(4);
if (!p) DEFER_EXE(scope);
DEFER(scope, free_p, { printf("freep\n"); free(p); });
void * q = malloc(40);
if (!q) DEFER_EXE(scope);
DEFER(scope, free_q, { printf("freeq\n"); free(q); });
DEFER_END(scope);
}
you can even macro the free_p/ free_q names with line numbers to shorten the macros, IE DEFER(scope, {}); #define guard_name3(name, line) name ## line
#define guard_name2(name, line) guard_name3(name, line)
#define guard_line_name(name) guard_name2(name, __LINE__)
#define guard { void * defer_scope = &&guard_line_name(defer_scope ##_exe); \
if (0) { guard_line_name(defer_scope ##_exe) : defer_scope = 0; }\
if (defer_scope != 0)
#define guard_end goto * defer_scope; }
#define guard_break goto *defer_scope;
#define defer(func) void * guard_line_name(defer_scope_item) = defer_scope; \
defer_scope = &&guard_line_name(defer_scope_item##_label); \
if (0) { guard_line_name(defer_scope_item##_label): { func } goto * guard_line_name(defer_scope_item); }
guard {
void * p = malloc(10);
if (!p) guard_break;
defer({ printf("freep\n"); free(p);});
void * q = malloc(10);
if (!q) guard_break;
defer({printf("freeq\n"); free(q);});
void * fail = 0;
if (!fail) guard_break;
defer({printf("freefail\n"); free(fail);});
guard_end;
}The results from 387 responses (to a Twitter poll) show a 2:1 preference for the value being read at the time the deferred statements are executed (66.9%) rather than when the defer statement encountered(33.1%).
Since both options - by value and by reference - may be viewed as reasonable or desirable, neither should be a default. Instead, they both should be using a special syntax. So if you write something likes this -
guard {
void * p = malloc(...);
defer free(p);
}
it simply won't compile. Instead you'd need to say something like this - guard {
void * p = malloc(...);
defer free(^p); // evaluate now
}
guard {
void * p = malloc(...);
defer free(p^); // evaluate later
}
This may also be reused later for specifying lambda captures... should lambdas ever make their way into C.IMO it would've been better (read, cleaner) to merge lambdas and function pointers into a single language construct. Throw in the partial application too and we'd be have a natively supported concept of a "callable" instance -
void foo(int tick);
void bar(int tick, int tock);
void do_something( void (* progress)(int tick) );
do_something( foo );
do_something( bar(,1) );
do_something( void (int tick) { /* lambda */ } );
The syntax is approximate, but for the code that is using callback-based flows this would've been very handy.If they add this, it would be better to restrict it to the current scope, to avoid forcing the language to create complexity to handle defers created via loop in the background.
Of course, that leads to the question of what to do when you loop over a defer with a goto, which would make sense if you were constantly appending function pointers and arguments to the frame similar to alloca.
That could lead to fun things like a function getting inlined into a for loop and then having its defers that would have been in function instead piled together to blow the stack.
They'll have to make sure they account for that in the spec and implementations.
Mostly, it would be easier to remember that C should not be using such a pattern, and that if you have need of defers and various similar magic to be baked into the language and compiler, what you're doing is probably inappropriate for implementation in the C language in the first place.
Why not have `defer free(p)` capture at deferral execution time, and `defer free(capture(p))` capture at statement execution time?
defer[p] free(p);
Instead of specifying it separately for each variable use.I would suggest doing it "the hard way", one thought would be treating it as an array (comptime known length) of u8
'void* x = defer_free_malloc (...);'
Not at all sure of the proper syntax, but at least tying it together with something like
void * const x = DEFER(malloc(...), free);
would make sense. In a more high-level language with interfaces, there should be a way to tie together the creation and destruction into a combined type, and just do obj = acquire SomeType(...);
And then 'acquire' should take care that the type to the right implements the interface, and automatically add a deferred call to the proper freeing method. Hm. I guess I just invented some kind of garbage collection, bummer.Or you know... free the things you need to free and move on with your life
And leave memory leaks, buffer overflows, and bugs in the process, and we've done for the past 40+ years...
Yes, which is a much better formulation.
So it adds a new way to misinterpret the code: is cleanup deferred or not?
That's the case only the first time you write code:
{
foo *f = new_foo(); // step 1
/* lots of code */ // step 3
free(f); // step 2
}
vs {
foo *f = new_foo(); // step 1
defer free(f); // step 2
/* lots of code */ // step 3
}
Sure, in both cases you can forget step 2. But what about review? With defer, the init and cleanup code are besides each other, and a missing defer would be immediately suspicious. Without defer, you'd have to check the end of the block to make sure the cleanup code is there. The absence of the cleanup code wouldn't jump to your eyes the same way the absence of defer would. In the long run, this makes defer significantly harder to forget.---
Another significant advantage of defer is that it can handle several exit points. Imagine this code:
Foo f = new_foo();
defer free(f);
if (!f) {
return FAIL_FOO;
}
Bar b = new_bar();
defer free(b);
if (!b) {
return FAIL_BAR;
}
Baz z = new_baz();
defer free(z);
if (!z) {
return FAIL_BAZ;
}
/* business logic */
/* business logic */
/* business logic */
return SUCCESS;
Now the same, without defer: Foo f = new_foo();
if (!f) {
free(f);
return FAIL_FOO;
}
Bar b = new_bar();
if (!b) {
free(f);
free(b);
return FAIL_BAR;
}
Baz z = new_baz();
if (!z) {
free(f);
free(b);
free(z);
return FAIL_BAZ;
}
/* business logic */
/* business logic */
/* business logic */
free(f);
free(b);
free(z);
return SUCCESS;
You really don't want to repeat yourself like that, you'd be liable to forget something. Now we could use `goto` and a return value: ReturnValue retval = SUCCESS;
Foo f = new_foo();
if (!f) {
retval = FAIL_FOO;
goto cleanup;
}
Bar b = new_bar();
if (!b) {
retval FAIL_BAR;
goto cleanup;
}
Baz z = new_baz();
if (!z) {
retval FAIL_BAZ;
goto cleanup;
}
/* business logic */
/* business logic */
/* business logic */
cleanup:
free(f);
free(b);
free(z);
return retval;
Better, except maybe the fact that Q/A hates you. All is not lost, you can still please them with a single exit point (pattern seen in the real world): ReturnValue retval = SUCCESS;
Foo f = new_foo();
if (f) {
Bar b = new_bar();
if (b) {
Baz z = new_baz();
if (z) {
/* business logic */
/* business logic */
/* business logic */
} else {
retval FAIL_BAZ;
}
free(z);
} else {
retval FAIL_BAR;
}
free(b);
} else {
retval = FAIL_FOO;
}
free(f);
return retval;
To be honest this may be the worst of them all.---
The only real contenders for this use case are defer and goto, and even then I think I prefer defer.
Language features should be orthogonal. A new language feature should add something that is not possible or extremely painful to do with the existing language features. I just don't see how these minor syntax adjustments warrant a new feature, especially one with as much complexity and corner cases as this defer proposal.
(The real answer to "why defer?", of course, is that the authors need it to implement panic/recover. This proposal should stop masquerading as a defer mechanism for C and instead call itself what it really is: exceptions for C.)
My, I didn't think it was possible to miss the point like that. Are you even arguing in good faith? Let's examine for a moment the 3 other alternatives.
First, we get the "repeat ourselves" problem: when I have several exit points, I must clean up at each exit. And if I edit the code in any way, (for instance by adding yet another check), I must review everything that has been initialised until this point and clean it up there again. This might be okay if I have only 1 or 2 exit points, but if I have more this is clearly unacceptable.
Second, we have goto. We replace our exit points by a goto cleanup. That one at least can scale. I don't like it however for three reasons. First, the cleanup code is at the end, far from the init code, so checking that the two pairs together correctly is inconvenient. Second, I need to manage an additional variable for the return value. Third, goto is banned in a lot of places, no matter how convoluted the alternatives may be.
Third, we have this monstrous pyramid if else that wastes horizontal space, requires you to re-indent everything at the slightest edit, separates cleanup code from init code, and is just plain ugly. The only thing going for it is the single exit point, and frankly it isn't much.
---
Those "alternatives" are anything but. They're what we have to do when faced with a limited language that doesn't express what we want to say. Workarounds, not solutions.
> minor syntax adjustments
Your perspective must be seriously warped if you're calling the function-wide reorganisation I spoke of "minor syntax adjustments". Or you're not arguing in good faith.
> The real answer to "why defer?", of course, is that the authors need it to implement panic/recover.
That is a separate point, which I think I agree with. Me, I just want a way to trigger an instruction when we exit the current scope. It's the necessary complement to `break` and `return`, which provide ways to exit scope before the end of the block. We could get rid of them, and apply a straightjacket structured programming discipline of course, but personally, I don't think I'm ready to give up on `break` and `return`.
Clearly we disagree on whether a syntax change is minor. But first let me repeat the point I made that you ignored in between your accusations: it really is just syntax. Of the four examples in the second part of your post, if we assume defer is implemented like attribute cleanup and fix up the compile errors, your first and fourth example compile to the identical assembly code:
I would argue that the best solution is one you didn't present: move the "business logic" into a separate function, one that takes the necessary resources as arguments. This way you're no longer mixing up resource acquisition error handling with business logic, and the function that acquires the resources can use the nested if statement style (or any other style) with no downsides. No surprises here, it again compiles to the identical assembly code:
In my opinion the nested if style is better than using defer because it's completely linear with no backward jumps. But even if you disagree you can hardly complain about cleanup code being far from init code because the whole resource handling function is less than 20 lines of code regardless of what cleanup style you chose. It doesn't matter, which is why I argue that it's a minor syntax change not worthy of addition to C.
It's really not. When the impact of "syntax" are non-local like that, it's more than syntax. A compiler would handle this beyond the parsing stage. At the very least, it would seriously massage the AST to remove `defer` from it.
> if we assume defer is implemented like attribute cleanup and fix up the compile errors, your first and fourth example compile to the identical assembly code:
This is to be expected: they ultimately do the same thing, and optimisers are known to do significant, non-local transformations to the code.
> move the "business logic" into a separate function, one that takes the necessary resources as arguments.
So now I have a function with (likely) too many arguments, that's used only once, and my eyes have to jump around to get to it (or I have to reach for the F2 key). The pyramid may be more visible, but that's a meagre advantage.
> In my opinion the nested if style is better than using defer because it's completely linear with no backward jumps.
Not even a criterion in my book. I suspect you're having an overly operational mindset. A mindset I suspect has held the whole field back a couple decades. Don't think of it like a backward jump. It's meant to be viewed as deferred execution, triggered by scope exit.
> you can hardly complain about cleanup code being far from init code because the whole resource handling function is less than 20 lines of code
That was an example, dummy. In real code, I'd have more than 3 things to initialise, and their initialisation might not be as trivial (or as repetitive) as what I've shown here. That's when I really want to read the code from top to bottom, with concerns packed together. Defer/cleanup lets me do that. The other solutions, less so.
Look at it from the opposite direction: if x->y didn't exist today, and the billions of lines of existing C code all used (*x).y, would you support a proposal to add a new x->y operator to the language? I doubt it.
Do you not like the array subscript operator, either, since a[b] can be *(a+b)? How about a && b, you can replace that with (!!a) & (!!b), with an extra 'if' if you need the short circuiting.
Well, that's the whole point of syntactic sugar.
Not that it gives you something you can't already do, but that it gives you a succint and better way to do it.
>Language features should be orthogonal. A new language feature should add something that is not possible or extremely painful to do with the existing language features.
I beg to differ, based on your definition of "extremely painful". Many kinds of syntactic sugar are welcome, even when the previous native solution wasn't "extremely painful" but e.g. just tedious or error prone.
Notice that there are five different goto targets, each for a specific case of what has-and-has-not been allocated. The resources are memory, locks, and even TLB flushing. This code would probably be cleaner with defer.
If your context has a list of things to free, it can also have a list of things to close. Or whatever you want.
Many libraries do stuff like this and call the context creation/deletion lib_init/ lib_shutdown.
I just did a search and it seems https://crates.io/crates/cc comes close. If only an IDE could do the job automatically when it finds .c files in the project. Even better, it should find and install a local compiler when needed (C or any other).
https://github.com/vlisivka/crust/blob/master/crust-mem.h#L3...
(Warning: GPL3)
I switched to Rust and lost interest in C.
I some ways that is encouraging, because research in practical systems things is valuable and used to be one of the cornerstones of computer science.
RAII is one of the best things about C++, and I'm excited for similar functionality in C. GCC's __cleanup__ is a poor substitute for a fully-baked addition to the language.
From skimming through the paper, it looks like there's an open discussion about the 'guard' keyword and scoping. I know that the scoping rules are tricky.
Would it make sense for defer statements to be attached to a variable's scope instead of a scope block? It would look something like GCC's __cleanup__, except that it could run an arbitrary statement/block instead of a callback. Scope-level defer() could be specified by attaching to a depth number.
If anybody involved in the paper is reading this, what would you think about this syntax?
//---------- Attaching a defer() to a variable's scope -----------//
int main(void) {
int *dummy = malloc(sizeof(int));
defer (dummy) {
printf("This statement prints second.\n");
free(dummy);
}
printf("This statement prints first.\n");
}
//------- Attaching a defer() to the current block's scope -------//
int main(void) {
int *dummy;
do {
dummy = malloc(sizeof(int));
defer (0) {
printf("This statement prints second.\n");
free(dummy);
}
printf("This statement prints first.\n");
} while (0);
printf("This statement prints third.\n");
}
//-------- Attaching a defer() to a parent block's scope ---------//
int main(void) {
int *dummy;
do {
do {
dummy = malloc(sizeof(int));
defer (1) {
printf("This statement prints third.\n");
free(dummy);
}
printf("This statement prints first.\n");
} while (0);
printf("This statement prints second.\n");
} while (0);
printf("This statement prints fourth.\n");
}
This syntax would eliminate the need for an explicit guard keyword, and would also make for a straightforward porting process for all code that currently uses __cleanup__. It also feels a little more C-like to me, in that it resembles the look-and-feel of other control-flow statements.In my example, defer() with a variable-name would attach itself to the variable's scope, and would execute when the variable leaves scope.
And defer() with an integer would attach itself to [current_scope_level - target_value]. So a defer(0) would trigger at the end of the current block scope, and defer(1) would trigger at the end of the parent block scope. Combined with generic {} blocks, you could get the same behavior provided by the paper's suggested guard keyword.
while(TRUE)
{
a = malloc(sizeof *a);
if(a == NULL)
break;
*a = malloc(1024);
if(*a == NULL)
break;
if(!do_some_other_tests(a))
break;
return a;
}/* cleanup /
if(a == NULL)
return;
if(a != NULL) free(*a);
free(a);
alt:void * * a = NULL;
switch(TRUE)
{
defaultt :
a = malloc(sizeof *a);
if(a == NULL)
break;
*a = malloc(1024);
if(*a == NULL)
break;
if(!do_some_other_tests(a))
break;
return a;
}/* cleanup /
if(a == NULL)
return;
if(a != NULL) free(*a);
free(a);