Why not just do simple C++ RAII in C?
thephd.dev
thephd.dev
Of course, but if you bothered at all to understand the constraints, you would have seen it is not actually that simple in our case.
And my project was several orders of magnitude simpler than the C standard.
Engineer (thinking): (No, you idiot, there isn't, because it's broken! I told you, all options have been tried, and this was the least painful way of doing it. Yes, it's not the ideal solution, but there's no other way, unless the upstream vendor decides to fix the issue on their end!)
Engineer: thanks, I'll look into it :)
We ended up supporting the brain-dead solution, but that team has now experienced 100% turn over since then.
It reads like you had experts giving you advise on how to improve things, and instead not only did you ignored their advise but you went to the extent of mindlessly disparaging their help.
They were doing it just to boost their egos and most of the teams in the company learned to ignore them. When the company ownership changed, the "A-Team" was the first under the chopping block because the new owners correctly saw that the high status they had was simply due to inertia of being first devs at the company and were not fullfiling any meaningful role in the present.
I’ve met dozens that don’t know their head from their ass. And always, always when you describe the problem constraints, they mumble and disappear.
(But at the time, I basically joined what was still essentially a startup just after they had been acquired by a larger company. I think the titles like 'architect' might have come from the larger company, but the competence came from them still being the same people as at the startup.)
We do exist, I promise. ;) But in my case at least, the Eye of Sauron can only keep so many things in sight at a time...
My sides. This is the most hilarious and accurate summary of every org where I’ve worked.
Thank you.
At some point we had to wear deodorant and a collared shirt, boom we became engineers.
(Yes, in a large enough corp, export control is a source of a surprisingly large amount of extra work...)
That gave me a good 5 minutes of chuckling and smiling. Thank you.
Good luck explaining that to the A-Team.
In all seriousness just accept their advice and see it for what it is. Someone trying to help you with limited view of the scope. As long as they don't impose their view I think your take is extremely bad.
One variant that I think might work even better than RAII or defer in a lot of languages is having a thread local "context" which you attach all cleanup actions to. It even works in C, you just define cleanup as a list of
typedef void(cleanup_function*)(void* context);
which is saved as into a thread local. Unlike RAII, you don't need to create a custom type for every cleanup action and unlike the call-with pattern from functional programming, lifetimes of these lists can be non-hierarchical.
However, I'm still glad to see defer being considered for C. It's a lot better than using goto for cleanup.Like the autorelease pool found in Objective-C of yore? I always liked that solution and sometimes implemented in plain C too.
test "arena allocator" {
var arena = std.heap.ArenaAllocator.init(std.heap.page_allocator);
defer arena.deinit();
const allocator = arena.allocator();
_ = try allocator.alloc(u8, 1);
_ = try allocator.alloc(u8, 10);
_ = try allocator.alloc(u8, 100);
}I've started to use simple memory arenas in C and it just feels so damn _nice_.
There's basically a shared lifetime for most of my transient allocations, which are nicely bounded in time by a "frame" of execution. Malloc/Free felt like a crazy amount of work, whereas an arena_reset(&ctx) just moves a pointer back to the first entry.
Another person pointed out that arenas are not destructors, and this is a great point to make. If you're dealing with external resources, moving an arena index back to the beginning does not help - at all.
#ifndef LIB_MALLOC
#define LIB_MALLOC malloc
#end
if it allows you to customize allocation strategy at all, which is not a given.But this only allows you at compile time to provide your own stateless global allocator. This is very different in Zig, which has a very strong culture of "if something needs to allocate memory, you pass it a stateful, dynamically dispatched allocator as an argument". You COULD do that in C, but virtually nobody does.
I've written reasonable amounts of both, and it's just different. For instance, in Zig, you can create a HashMap using a FixedBufferAllocator, which is a region of memory (which can be stack allocated) dressed up as an allocator. You can also pass it an arena and free all at once, or any other allocator in the standard library, or implemented by you, or anyone else. Show me a C library with a HashMap which can do all three of these things. Everything which allocates takes an allocator, third-party libraries respect this convention or will quickly get an issue or PR either requesting or implementing this convention.
Ultimate solution? No, but also, sort of. The ability to idiomatically build a fine-grained memory policy is a large portion of what makes Zig so pleasant to use.
Is it actually complicated? There’s only the rule of 0 - either your class isn’t managing resources directly & has none of the 5 default methods defined explicitly (destructor, copy constructor/assignment, move constructor/assingment), or it manages 1 and exactly 1 resource and defines all 5. Following that simple rule gives you exception safety & perfect RAII behavior. Of all the things in C++, it seemed like the most straightforward rule to follow mechanically.
BTW, the rule of 3 is from pre-C++11 - the addition of move construct/move assignment makes it the rule of 5 which basically says if you define any of those default ones you must define all of them. But the rule of 0 is far stronger in that it gives you prescriptive mechanical rules to follow for resource management.
It’s much easier to do RAII correctly in Rust because of the ecosystem of the language + certain language features that make it more ergonomic (e.g. Borrow/AsRef/Deref) + some ownership guarantees around moves unless you make the type trivially copyable which won’t be the case when you own a resource.
It is. There is no point in arguing otherwise.
To understand the problem, you need to understand why it is also a solution to much bigger problems.
C++ started as C with classes, and by design aimed at being perfectly compatible with C. But you want to improve developer experience, and bring to the table major architectural traits such as RAII. This in turn meant you add support for custom constructors, and customize how your instances are copied and destroyed. But you also want to be able to have everything just work out of the box without forcing developers to write boilerplate code. So you come up with the concept of special member functions which are automatically added by the compiler if they are trivial. However, forcing that upon every single situation can cause problems, so you have to come up with a strategy that suits all use cases and prevents serious bugs.
Consequently, you add a bunch of rules which boil down to a) if the class/struct is trivial them compilers simply add trivial definitions of all special member functions s that you don't have to, but once you define any of those special member functions yourself them the compiler steps back and let's you do all the work.
Then C++ introduced move semantics. This refreshes the same problem as before. You need to retain compatibility with C, and you need to avoid boilerplate code, and on top of that you need to support all cases that originated the need for C++'s special member functions. But now you need to support move constructors and move assignment operators. Again, it's fine if the compiler adds those automatically if it's a trivial class/struct, but if the class has custom constructors and destructors then surely you also need to handle moves in a special way, so the compiler steps back and lets you do all the work. On top of that, you add the fact that if you need custom code to copy your objects around, surely you need custom code to move them too, and thus the compiler steps back to let you do all the work.
On top of this, there are also some specific combinations of custom constructors/destructors/copy constructors/copy assignment operators which let the compiler define move constructors/move assignment operators.
It all makes absolutely sense if you are mindful of the design requirements. But if you just start to onboard onto C++ and barely know what a copy constructors is, all these aspects are arcane and sadistic. If you declare nothing then your class instances are copied and moved automatically, but once you add a constructor everything suddenly blows up and your code doesn't even compile anymore. You spotted a bug where an instance of a child class isn't being destroyed properly, and once you add a virtual destructor you suddenly have an unrelated function call throw compiler errors. You add a snazzy copy constructor that's very performant and your performance tests suddenly start to blow up because of the performance hit if suddenly having to copy all instances instead of the compiler simply moving them. How do you sort out this nonsense?
The rule of 5 is a nice rule of thumb to allow developers to have a simple mental model over what they need to do to avoid a long list of issues, but you still have no control over what you're doing. Things work, but work by sheer coincidence.
I wouldn't be so dramatic. House of cards don't stay put by coincidence !
Well, I don’t know how to respond to this. I clarified what the rules actually are (< 1 paragraph) and following them blindly leads to correct results. You’ve brought in a whole bunch of nonsense about why C++ has become complex as a language - it’s not wrong but I’m failing to connect the dots as to how the rule of 0 itself is hard to follow or complex. I’m kind of taking as a given that whoever is writing the code is mildly familiar enough with C++ to understand RAII & is trying to apply it correctly.
> The rule of 5 is a nice rule of thumb to allow developers to have a simple mental model over what they need to do to avoid a long list of issues, but you still have no control over what you’re doing. Things work, but work by sheer coincidence.
First, as I’ve said multiple times, it’s the rule of 0. That’s the rule to follow to get correct composition of resource ownership & it’s super simple. As for not having control, I really fail to see how that is - C++ famously gives you too much control and that’s the problem. As for things working by sheer coincidence, that’s like your opinion. To me “coincidence” wouldn’t explain how many lines of C++ code are running in production.
Look, I think C++ has a lot of warts which is why I prefer Rust these days. But the rule of 0 is not where I’d say C++’s complexity lies - if you think that is the case, I’d recommend you use another language because if you can’t grok the rule of 0, the other footguns that lie in wait will blow you away to smithereens.
So it's not nonsense?
I think GP clearly laid out the base principles that lead to emergent complexity . GP calls this "coincidence" to convey the feeling of lots of complexity just narrowly avoiding catastrophe in a process that is hard to grok for someone getting into C++. GP also gave some scenarios in which the rule of 0 no longer applies and you now simply have to follow some other rule. "just follow the rule" is not very intuitive advice. The rule may be simple to follow but the foundations on which it rests are pretty complicated, which makes the entire rule complicated in my worldview and also that of GP. In your view, the rule is easy to follow therefore simple. Let's agree to disagree on that. Again, being told "you need to just follow this arbitrary rule to fix all these sudden compiler errors" doesn't inspire confidence in ones code, hence (I think) the usage of "coincidence". If I were using such a language, I'd certainly feel a bit nervous and unsure.
I think that's what they said themselves:
>> It all makes absolutely sense if you are mindful of the design requirements. But if you just start to onboard onto C++ and barely know what a copy constructors is, all these aspects are arcane and sadistic
IMO not knowing why something works (in any language) is an unpleasant feeling. Then if you have the chance you can look under the hood, read things - it's exactly why I'm reading this thread - and little by little get a better understanding. That's called gaining experience.
> Again, being told "you need to just follow this arbitrary rule to fix all these sudden compiler errors" doesn't inspire confidence in ones code, hence (I think) the usage of "coincidence"
That's exactly what other languages like Haskell or Rust are praised for. Why does C++ receive a different treatment when it tries to do the same thing instead of crashing on you at runtime, for once?
You making a trivial change, and suddenly there are entire new classes of bugs all over your code is an aspect that does really not receive any praise. People using those two languages work hard on avoiding that situation, and it clearly feels like a failure when it happens.
The part about pointing problems at compile time so the developer will know it sooner is great. And I imagine is the part you are talking about. But the GP was talking about the other part of the issue.
There is a neater design in rust with its own tradeoffs: destructors are the only special function, move is always possible and has a fixed approach, copying is instead .clone(), assignment is always just a move, and constructors are just a convention with static methods, optionally with a Default trait. But that does constrain you: especially move being fixed to a specific definition means there's a lot you can't model well (self-referential structures), and that's a core part of why rust can have a neater model. And it still has the distinction you are complaining about with Copy, where 'trivial' structures can be copied implicitly but lose that as soon as they contain anything with a destructor or non-trivial .clone().
And in C++ it's pretty easy to avoid this mess in most cases: I rarely ever fully define all 5. If I have a custom constructor and destructor I just delete the other cases and use a wrapper class which handles those semantics for me.
I'm sorry, that is not true at all.
Nothing forces you to add implementations, at least not for all cases. That's only a simplistic rule of thumb that helps developers not well versed on the rules of special member functions (i.e., most) to get stuff to work by coincidence. You only need to add a, say, custom move constructor when you need it and when the C++ rules state the compiler should not generate one for you. There's even a popular table from a presentation from ACCU2014 stating exactly in which condition you need to fill in your custom definition.
https://i.sstatic.net/b2VBV.png
You are also wrong when you assert this has nothing to do with C++'s heritage. It's the root cause of each and every single little detail. Special member functions were added with traits and tradeoffs for compatibility and ease of use, and with move semantics the committee had to revisit everything over again but with an additional layer of requirements. The rules involving default move constructors and move assignment operators are famously nuanced and even arbitrary. There is no way around it.
> There is a neater design in rust (...)
What Rust does and does not do is irrelevant. Rust was a greenfield project that had no requirement to respect any sort of backward compatibility and stability. If there is any remotely relevant comparison that would be Objective-C, which also took a minimalist approach based on custom factory methods and initializes that rely on conventions, and it is a big boilerplate mess.
I'm afraid you're complaining about entirely unrelated things.
It's one thing to claim that C++ structs have this or that trait. It's a entirely different thing to try to pin bugs and developer mistakes on how a language is designed.
It is not very complicated at all; just a discipline to follow (or not if you know what you are doing) once learnt - https://en.cppreference.com/w/cpp/language/rule_of_three
Incidentally i use it as 4/6/0 by including the default ctor in the set.
100% of the using structs like they are C structs vs using class as objects is cultural not a part of the language.
C++ is somewhat unique in that it started out as a few extra features on top of C before gradually splitting off and mutating into a totally separate programming language.
But Cfront was released circa 1983 and you basically just wrote C, but it added a bit of new syntax that generated extra C behind the scenes. Object-oriented programming was still fetal in 1983! It didn't get really hyped until the mid-90's. So C++ kind of mutated for decades as this gross appendage on C until it became this whole separate blob that ate half of programming. It was 15 years later when the C++98 "standard" started trying to reign in Dr. Stroustrup's monster.
Then in 2005 we threw away all our textbooks that were like "Look! `Apple` derives from `Fruit`! `Car` derives from `Engine`! This is going to change the world!" because adding object-orientedness to everything became uncool when our bosses became fans of Java. But by this point the C++ blob had taken on a life of its own...
So yeah. Very few programming languages have a story as long and insane as C++.
Objective-C++ likewise on top of CFront.
Until like with CFront, they became selfhosted compilers.
Groovy code is Java code, regardless of targeting the JVM, the same syntax is supported and extended with dynamic capabilities.
Object Pascal was created for Lisa project, exactly in 1983.
Tom Love and Brad Cox created Objective-C in 1984.
I think this take is completely wrong. There is nothing cultural about it. C++ was created as a strict superset of C, and thus from the inception it supported all features made available in C. This design goal remains true up to this day, and only started to diverge relatively recently when C was updated to include features that were not supported (yet) by C++.
When someone declares a plain old struct in C++, they are declaring a struct that is perfectly compatible and interoperable with C. This is by design. From the inception.
Just asking, for a friend.
Just a friendly reminder that two leading underscores wont protect your member functions in C++. Even if people insist that those are totally not supposed to be private in python.
Whenever I say "I'm no longer attached to all that private stuff", people always reply, "wait until you work on a large code base". I work on a million line+ code base. Whatever.
This argument aside, I'm not a total philistine. RAII is awesome but C++ is full to the boot with crusty stuff to keep the compatibility. I always feel there is a language better than anything trying to come out.
These days I'm for minimalism, most of my structs are aggregates of public members, but sometimes you really want to make sure to maintain the invariant that your array pointer and your size field are in sync.
Of course neither double nor single underscore will stop anyone who wants to touch your privates badly enough. Which is big part of the python philosophy: You're not stopped from doing inadvisable things. Instead there's a strong culture around writing "pythonic" code, which largely avoids these pitfalls.
In python, if any of this gives you an trouble you can just replace the stuff in the class dict with your own functions. You don't even need to cast.
This is not really the case. See https://en.wikipedia.org/wiki/Compatibility_of_C_and_C%2B%2B for a non-exhaustive list.
It is true that both sides agree that compatibility is an important goal, but it's only a goal, not something that's 100% the case.
I think this glances over what structs actually are in C++, and unwittingly portrays them as something different.
Structs in C++ are definitely exact like structs in C. Or they can be, if that's what you're aiming for. If you include a C header file that defines a struct in a C++ program, you build it, and you use instances of that struct to pass them to C programs, everthing just works.
The detail you need to be mindful of is that C structs support a subset of all the features supported by C++ classes, and once you start to use those features C++ also allows implementations to forego some constraints.
If you expect to use a struct in C++ but still define it in a way that you make it include features that are not supported in C then you can't pin that on the language.
https://learn.microsoft.com/en-us/cpp/cpp/trivial-standard-l...
Using C-like structs is a very common use case, to the point that the standard explicitly defines the concept of standard layout and builds upon that to specify the concept of a standard layout type. A struct/class that is a standard layout type, which means it's a POD type, corresponds exactly with C structs. They are explicitly defined in terms of retaining interoperability with other languages.
Still. There's always extern "c".
But yes, if you make extra sure (under threat of footgun) that your struct only has simple types in it and doesn't use virtual or define any ctors/dtors or use protected/private or use inheritance and all of its members follow those rules etc etc, maybe you can treat it like a C struct. But the C++ Standard is telling a different story.
Keep in mind, I'm not blaming you for ignoring all these complications if at the end of the day the compiler seems to give you the behavior you expect. But the fun of C++ is that it's kind of two programming languages in one: the language the Standard defines, and the language the typical programmer thinks it is.
[0] There was std::is_pod, but it was deprecated because it doesn't reflect how the Standard actually defines things. A bit of a cruel joke, dangling that in front of us and then yanking it away.
References:
1) Trivial, standard-layout, POD, and literal types - https://learn.microsoft.com/en-us/cpp/cpp/trivial-standard-l...
2) No more plain old data - https://mariusbancila.ro/blog/2020/08/10/no-more-plain-old-d...
Keep in mind, my original comment was pretty much just drawing a line through TFA, which also argues that you can't cleanly map C++ object concepts onto C structs. C++ has some backwards compatibility with C obviously but nowadays it's a totally separate language with an independent standards body (for better or worse). Specifying "do what C does" might have flown in 1998 but that changed a long time ago.
I am fully with Stroustrup in arguing that C++ should strive for as much compatibility with C as possible in the spirit of the original (see ref. at https://en.wikipedia.org/wiki/Compatibility_of_C_and_C%2B%2B...). But sadly the rest of standards committee don't seem to want this which i believe is a huge mistake. On the other side, the C standards committee should be very careful what inspiration they take from C++ in the evolution of the language since it was designed as a "minimal" language which was one of the main factors in its success. Whether people call it "primitive", "well behind other languages" etc. does not matter. You definitely don't want C turning into C++-lite. Hence IMO the conclusions stated in the last few paragraphs of the submitted article are quite right.
In a way, the whole C++ endeavor was doomed from the start. C was old and pragmatic and vague, a "portable assembly", and it was a shaky foundation to build C++ on top of. When the Standard tried to tighten things up, it just got more lopsided, full of hacks to fix hacks. But the alternate universe where C++ had a more pragmatic, laissez-faire design going forward probably isn't any better; maybe the "standard" would have become "do whatever GCC does"--or in the Darkest Timeline, "do whatever MSVC does".
I disagree that C++ "respecting its C roots" is viable. The C++11 and later Standards were trying to make the best of a bad situation, and that required leaving C behind because the C way of doing things doesn't fit with a higher-level language like contemporary C++. Especially when the language has multiple implementations that need to compile the same code the same way. The "C with classes" days are long over for most of us who have to use libraries expecting std::vector, smart pointers, and exception handling. We live in mortal fear of compiler writers smiting us for innocent things like punning through a union.
> You definitely don't want C turning into C++-lite
I agree. Trying to quickly hack classes or templates or whatever back on top of C would just start the whole C++ nightmare over again.
Hey! Them's fighting words! :-) "C++ as a better C" (which is what it started as) was/is/always will be needed and necessary. It gave you the best of both low-level and high-level worlds with full control and just enough complexity. Instead of implementing structs full of function pointers to design dynamic dispatch object models you just had the compiler do that for you while still retaining full control over other aspects. I still have some manuals that came with SCO Unix one of which was on the then newfangled C++ language. It had one chapter by Stroustrup himself (his original paper probably) on the C++ object model showing how vptrs/vtables are implemented and thinking it neat that the compiler did it for you. Also templates were just glorified macros then with none of the shenanigans that you see today. Hence moving from C to C++ was easy and its usage and popularity exploded. But with the infusion of lots of people into C++ land people who were not aware of the original vision/design/compatibility goal of the language started asking for the inclusion of more and more OO and modern language features. The result? The standards committee reinventing the language from C++11 onwards(and changing every freaking 3 years) and alienating the old C++ folks who made it popular in the first place. No doubt there are some benefits like increased design space and modern programming techniques but am not sure whether the increased complexity makes it all worth it. For me it is still C++98 with the addition of the STL and some simple generic programming techniques which is the sweet spot.
C++20 introduced `std::bitcast`, so I appreciate alias analysis getting all the help it can.
Not true. Using C structs themselves in C++ is very common - when you include the C header file, the relevant declarations are wrapped in "extern "C" {}" which gives structs C semantics. You can do this because C++ is backwards compatible with C.
Most of the time when you use a struct in C++ you're just ignoring most of the capabilities of objects (which is fine!). If you declare a struct in C++, you're getting an object. The only difference between the struct and class keywords in C++ is the default privacy of the members.
What I think you're trying to say is "a POD structure with no custom behavior is essentially identical in C and C++". That is mostly true, though if the struct contains a union, C++ has stricter UB rules (there might be other differences as well, but that's the one I can think of at the moment).
Again, "traditionally", one could (ab)use C++ as "C with extras". And it wasn't uncommon, especially in resource constraint usecases, to do just that. C++ without STL or templates, or even C++ without new/delete.
This "is not C++", agree. Would a subset be enough for "using it like C-with-RAII" ?
Given the details and pitfalls the original author lists, I suspect not. It's not just C programmers who "do strange things" and make odd choices. The language itself though "lends itself to that". I've (had to) write code that sometimes-alloca'ed sometimes-malloc'ed the same thing and then "tagged" it to indicate whether it needed free() or "just" the implied drop. Another rather common antipattern is "generic embedded payloads" - the struct definition ending "char data[1]" just to be padded out by whatever creates it to whatever size (nevermind type) of that data.
Can you write _new_ C code that "does RAII" ? Probably. Just rewrite it in rust, or zig :-) Can you somehow transmogrify language, compiler, standard lib so that you can recompile existing C code, it not to "just get RAII" then at least to give you meaningful compiler errors/warnings that tell you how to change it ? I won't put money on that.
https://www.youtube.com/watch?v=rX0ItVEVjHc
A classic which touches on such stuff.
You can do "manual" goto-based RAII in C, and it has been done for decades. The end of your function needs to have a cascading layer of labels, undoing what has been done before:
if (!(x = create_x())) {
goto cleanup;
}
if (!(y = create_y())) {
goto cleanup_x;
}
if (!(z = create_z())) {
goto cleanup_y;
}
do_something(x, y, z);
cleanup_z:
destroy_z(z);
cleanup_y:
destroy_y(y);
cleanup_x:
destroy_x(x);
cleanup:
return;
It just takes more discipline and is more error-prone maintenance-wise.It’s not clear if you’re talking about defer or RAII
Like C, with its many hidden behaviors?
I would argue that if it needs to be spelled out in a separate document from the code you're reading, then it's hidden.
But that's already what linters/static analyzers are doing? But then, why not integrate those tools directly in a C++ compiler instead?
With cpp2/cppfront, Herb Sutter is already building some sort of a "sane" subset of the C++ language, maybe because you cannot achieve good practices without having a new syntax.
C++ seems to have the same problem of javascript: it has annoying "don't-do-that" use cases, although it seems insanely more complicated to teach good C++ practices.
Of course, this requires buying into a set of tooling and learning a lot of specific idioms. I can't say I've used it, but from reading the docs it seems sound enough.
The issue is developers that think they are useless tools.
This sounds like a great idea to me! Rust disables implicit copying for structs with destructors, and together with move-by-default, it works really well. Unlike PoD structs, you don't need to heap allocate them to ensure their uniqueness. Unlike copy constructors, you don't need to worry about implicit copies. Unlike C++ move, there's no moved-from junk value left behind.
But you could gain reusability of headers to be also used in C++, not needing to reinvent the wheel with new issues (e.g. variable lifetime), and a whole lot of existing experience with RAII.
"Disabling" is maybe not the right way to think about it. Rust only has "implicit copying" for Copy types, so you have to at the very least #[derive(Copy,Clone)] to get this, it's true that you can't (and therefore neither can a derive macro) impl Copy on types which implement Drop and that's on purpose but you're making a concrete decision here - the answer Rust knows is never correct is something you'd have to ask for specifically, so when you ask it can say "No" and explain why.
Lots of similar behaviour in C++ is silent. Why isn't my Doodad behaving the way I expected? I didn't need to ask for it to have the behaviour I expected but the compiler concludes it can't have that behaviour, so, it doesn't, and there's nowhere for a diagnostic which says "Um, no a Doodad doesn't work like that, and here's why!"
Diagnostics are hard and C++ under-values the importance of good diagnostics. Rust recently landed work so libraries can provide improved diagnostics when you try to call them with inappropriate parameters. For example now if you try to collect() an iterator into a slice, the compiler notices that slice doesn't implement FromIterator and it asks FromIterator to explain why this can't work, whereupon FromIterator notices you were trying to use a slice and emits a diagnostic for this particular situation - if you'd tried to collect into an array it explains how you'd actually do that, since it's tricky - the slice is impossible since it's not an owning type, you need to collect into a container.
The initial example in the article is anti-idiomatic, because it imbues the larger class with a RAIIness which can be limited to just one element of it:
struct ObjectType {
int a;
double b;
void* c;
ObjectType() : a(1), b(2.2), c(malloc(30)) { }
~ObjectType() { free(c); }
};
It's only the c member that really requires any special attention. In this particular case. So, there should be something like a `class void_buffer` which is a RAII class, and then: struct ObjectType {
int a;
double b;
void_buffer c;
ObjectType() : a(1), b(2.2), c(30) { }
};
and actually, let's just not sully the set of constructors, but rather have: struct ObjectType {
int a;
double b;
void_buffer c;
static ObjectType make() {
return ObjectType{ 1, 2.2, 30 };
}
};
and now instead of a complicated bespoke class we have the simplest of structs; the only complexity is in void_buffer.It abstracts the void_buffer into its own type with proper correct functions for creating, (maybe copying), moving, and destructing the buffer. With that you get a simple type that you can use elsewhere without needing to remember that you need to free() the buffer manually before the end of the scope, or needing to remember how to correctly copy or move the buffer elsewhere.
* Less code overall
* More reuse of classes as versatile/simple components, as opposed to a zoo of bespoke classes
* Classes which are simpler to understand and with more predictable behavior
This is true in the example above: With the corrected code, it's enough that I tell you "ObjectType is a simple struct; and one of its members is a buffer of untyped data". I don't have to show you the class definition; you know enough to understand what's going on. And you can use your void_buffer elsewhere.
1. It would be no more difficult to write and use than the larger class. After all, you can use the larger class as a void_buffer with some dummy extra fields.
2. You can put the class in a detail_ sub-namespace, or make it an inner class of ObjectType, and then people will avoid using it in other, general contexts.
There are 2 ways to get C++-style RAII into C. The first way is to wholesale import the C++ object system into C (which means name mangling, all the different flavors of constructors, destructors, etc). Conceptually this would work, but it's never going to happen, because implementing that would be literally more work than an entire conforming C99 compiler.
The second way is to just use some special function attributes to signify that a function runs when an object is created on the stack / popped off the stack. This won't work either because the C++ object system also solves lots of other problems that this simpler system just ignores (such as, what happens when you copy an object that has a constructor function).
When I started reading it, the first thing that came to my mind was the issue with copying the structs. The article started looking at the issue, but didn't really follow further with the changes needed to make it work, which is that you start needing to introduce tracking which instance is responsible for the resources and providing a way to transfer that responsibility (a.k.a. ownership and move semantics).
amateur C++ coder
Another way to think about it: even if you had defined constructors and destructors for a struct, you have not solved when to call them. C++'s answer to that question is its sophisticated object model. C does not have one, and in order to answer that question, it must. It's worth noting that RAII was not a feature that was intentionally created in C++. Rather, astute early C++ developers realized it was a useful idiom made possible by C++'s object model.
I think the actual question should be "can C get automatic memory management like in C++ without having the equivalent of C++'s type system"?
Though I can't put my finger on it, my intuition says it can, if the interested people are willing to look deep enough.
This doesn’t make sense: you don’t need runtime introspection to do this?
Introspection (reflection) would go even further and provide at runtime all the information that you have at compile time about an object. But that's not required for assignment and destruction operations to work.
C doesn't have any of that, so a struct copy is just a shallow copy, a bit by bit copy of the entire struct contents. Which works pretty well, except for pointers/references.
So, no, runtime introspection is not needed, but runtime dispatch may be needed.
I haven’t seen this distinction laid out so clearly before:
Every other language worth being so much as spit on either employs deep garbage collection (Go, D, Java, Lua, C#, etc.) or automatic reference counting (Objective-C, Objective-C++, Swift, etc.), uses RAII (Rust with Drop, C++, etc.), or does absolutely nothing while saying to Go Fuck Yourself™ and kicking the developer in the shins for good measure (C, etc.).
GC, ARC, RAII or GTFO, those are the options. That’s right!
I always come away from these discussions with more respect for Objective-C -- such a powerful yet simple language. I suppose Swift is the successor but it feels very different.
Although, Obj-C only really came into its own once it finally gained automatic reference counting, after briefly flirting with GC. At that point it was already being displaced by younger and more fashionable languages.
C11 provided a few worthwhile improvements (i.e., a proper memory model, alignment specification, standardized anonymous structures/unions), but so many of the other additions, suggestions, and proposals I’ve seen will just ruin the minimal nature of C. In C++, a simple statement like `a = b++;` can mean multiple constructors being called, hidden allocations, unexpected exceptions, unclear object hierarchies, an overloaded `++`, an overloaded `=`, etc. Every time I wish I had some C++ feature in C, I just think about the cognitive overhead it’d bring with it, slap myself a couple times, and go back to loving simple ole C.
Please don’t ruin C.
Correctness can be established well enough - even if guaranteed automatically - in a language with UB.
No, this is wrong. It's a common misconception though. You would only want that in a hypothetical world where all computers are exactly the same.
Undefined and implementation defined behavior is what allows us to have performance at all. Here are some simple examples.
Suppose we want to make division by zero and null pointer dereference defined. Now every time you write a/b or *x, the compiler will be forced to emit an extra branching check before this operation.
Something much more common---addition. What about signed overflow? Do you want the compiler to emit an overflow check in advance? Similar reasoning for shift instructions.
UB in the language specification allows compilers to optimize based on the assumption that the programs you write won't have undefined behavior. If compilers are not able to do this, it becomes impossible to implement most optimizations we rely on. It's a very core feature of modern language specifications, not an oversight you can fix by thinking about it for 10 minutes.
Some of these checks could be removed by languages with better compilers and likely more restrictions. That is the better approach. As a user, I don't want to run code that is potentially unsafe and/or insecure.
You would not want to force these by default, nobody wants it. You can not statically determine them unnecessary in for the vast majority of code, even stuff as simple as `print(read(a) + read(b))`.
I know that languages like Java have a NullPointerException which they can throw and handle for situations like this, but they're also built on a highly specified virtual machine architecture that is consistent across hardware platforms. This also does not guarantee that your program is safe from crashing when this exception gets thrown, as you have to handle it somewhere. For something as general as this it will probably be in the Main function, so you might as well let it go unhandled as there's not that much you can do at that point.
For a language like C++ it is simpler, easier, and I would argue more correct, to just let the hardware handle the situation, which in this case would trigger a memory error of trying to access invalid memory. As the real issue is probably somewhere else in the code which isn't being handled correctly and the bad data is flowing through to the place where it accesses the null pointer and the program crashes.
To add to that in a lot of cases the program isn't crashing while trying to access address 0, it's crashing trying to access address 200, or 1000, or something like that, and putting in simplistic checks isn't going to catch those. You could argue that the check should guard against accessing the lowest 1k of memory, but then when do you stop, at 64k? Then you have an issue with programs that must fit within 1k of memory.
Leaving it unspecified is the better choice.
Given that has proven to be a completely false assumption, I don't think there's a justification for compilers continuing to make it. Whatever performance gains they are making are simply not worth the unreliability they are courting.
This part is correct. The problem is in how to deal with this. If you want the compiler to correctly deal with code having undefined behavior, often the only possibility is to assume that all code has undefined behavior. That means, almost every operation gets a runtime branch. That is completely incompatible with how modern hardware works.
The rest is wrong, but again, this is a common misconception. Language designers and compiler writers are not idiots, contrary to popular belief. UB as a concept exists for a reason. It's not for marginal performance boosts, it is to enable any compiler based transformation, and a notion of portability.
A good example is WebAssembly*—address 0x00000000 is a perfectly fine and well-defined address in linear memory. In practice though, most code you’ll come across targeting WebAssembly treats it as if dereferencing it is undefined behavior.
* Of course WebAssembly is a compiler target rather than a language, but it serves as a good example of the point you’re making.
If a pointer can be null, it must be an optional pointer, and you must in fact check before you dereference it. This is what you want. Is it ok to write a program which segfaults at random because you didn't check for a pointer which can be null? Of course not. If you don't null-check the return value of e.g. malloc, your program is invalid.
But the benefit is in the other direction. Careful C checks for null before using a pointer, and keeping track of whether null has been checked is a manual process. This results in redundant null checks if you can't statically prove (by staring at the code and thinking very hard) that it isn't null. So in practice you're likely to have a combination of not checking and getting burned, and checking a pointer which was already checked. To do otherwise you have to understand the complete call graph, this is infeasible.
Zig doesn't do any of this. If it's a pointer, you can safely dereference it. If it's an optional pointer, you must check, and then: it's a pointer. Safe to pass down the call stack and freely use. If you want C behavior you can always YOLO and just say `yoloptr.?.*`.
Overflow addition and divide by zero are safety checked undefined behavior, a critical concept in the specification. They will panic with a stack trace in debug and ReleaseSafe mode, and blow demons out of your nose in ReleaseFast and ReleaseSmall modes. There's also +% for guaranteed wraparound twos-complement overflow, and +| for saturating addition. Also `@addWithOverflow` if your jam is checking the overflow bit. Unwrapping an optional without checking it is also safety-checked UB: if you were wrong about the assumption that the payload carries a value, you'll get a panic and stack trace on the line where you did `yolo.?`.
Shift operations require that the right hand side of the shift be a type log2(Type.bitwidth) of the left hand side. Zig allows integers of any width, so for a: u64, calling a << b requires that b be a u6 or smaller. Which is fine: if you know values will be within 0..63, you declare them u6, and if you want to shift on a byte, you truncate it: you were going to mask it anyway, right? Zig simply refuses to let you forget this. Addition of two u6 is just as fast as addition of the underlying bytes because of, you got it, safety-checked undefined behavior. In release mode it will just do what the chip does.
There's a common theme here: some things require undefined behavior for performance. Zig does what it can to crash your program if that behavior is exhibited while you're developing it. Other things require that you take some well-defined actions or you'll get UB: Zig tracks those in the type system.
You'll note that undefined behavior is very much a part of the Zig specification, for the same reasons as in C. But that's not a great excuse to make staying within the boundaries of defined behavior as pointlessly difficult as it is in C.
The debug modes you mention are also available in various forms in C and C++ compilers. For example ASan and UBSan in clang will do exactly what you have described. The question is, then whether these belong in the language specification or left to individual tools.
Language specification is unavoidable when using said language.
no such thing ever existed.
For a bunch of languages outside the C-centric world, specifications don't exist.
The intuitive distinction is that the second one is for compiler/library developers, and the former is for users.
A specification can not leave any room for ambiguity or anything up to interpretation. If it does (and this happens), it is treated as a bug to be fixed.
If you can make it work in a way that has acceptable performance characteristics, every systems language will adopt your technique overnight.
Huh, doesn't that sound familiar?
This is not the case. It's two's compliment overflow.
Also, since we're being pedantic here: it's not actually about "debug mode" or "release mode", it is tied to a flag, and compilers must have that flag on in debug mode. This gives the ability to move release mode to also produce the flag in the future, if it's decided that the overhead is worth it. We'll see if it ever is.
> Huh, doesn't that sound familiar?
Nope, it is completely different from undefined behavior, which gives the compiler license to do anything it wants. These are well defined semantics, the polar opposite of UB.
So if we're going to be pedantic, it's safe Rust which has defined semantics for basically everything. A considerable accomplishment, to be sure.
Okay, here is an example showing that rust follows LLVM behavior when the optimizer is turned on. LLVM addition produces poison when signed wrap happens. I'm a little bit puzzled about the vehement responses in the comments wow. I have worked on several compilers (including a few patches to Rust), and this is all common knowledge.
define noundef i32 @square(i32 noundef %x, i32 noundef %y) unnamed_addr #0 !dbg !7 {
%_0 = add i32 %y, %x, !dbg !12
ret i32 %_0, !dbg !13
}
Let's compare like to like, here's one with equivalent C++ code: https://godbolt.org/z/Y4MnGeof4The C++ output:
define dso_local noundef i32 @square(int, int)(i32 noundef %0, i32 noundef %1) local_unnamed_addr #0 !dbg !99 {
tail call void @llvm.dbg.value(metadata i32 %0, metadata !104, metadata !DIExpression()), !dbg !106
tail call void @llvm.dbg.value(metadata i32 %1, metadata !105, metadata !DIExpression()), !dbg !106
%3 = add nsw i32 %1, %0, !dbg !107
ret i32 %3, !dbg !108
}
> LLVM addition produces poison when signed wrap happens.https://llvm.org/docs/LangRef.html#add-instruction
> nuw and nsw stand for “No Unsigned Wrap” and “No Signed Wrap”, respectively. If the nuw and/or nsw keywords are present, the result value of the add is a poison value if unsigned and/or signed overflow, respectively, occurs.
Note that Rust produces `add`. The C++ produces `add nsw`. No poison in Rust, poison in C++.
Here is an example of these differences producing different results, due to the differences in behavior: https://godbolt.org/z/Gaonnc985
Rust:
define noundef zeroext i1 @test() unnamed_addr #0 !dbg !14 {
ret i1 true, !dbg !15
}
C++: define dso_local noundef zeroext i1 @test()() local_unnamed_addr #0 !dbg !123 {
tail call void @llvm.dbg.value(metadata i32 undef, metadata !128, metadata !DIExpression()), !dbg !129
ret i1 false, !dbg !130
}
This is because in Rust, the wrapping behavior means that this will always be true, but in C++, because it is UB, the compiler assumes it will always be false.> I'm a little bit puzzled about the vehement responses in the comments wow.
You are claiming that Rust has semantics that it was very, very deliberately designed to not have.
The more interesting part is that the mode can be individually modified on a per-block basis with the @setRuntimeSafety builtin, so it's practical to identify the performance-critical parts of the program and turn off safety checks only for them. Or the opposite: identify tricky code which is doing something complex, and turn on runtime safety there, regardless of the build status.
That's why this sort of thing should be part of the specification. @setRuntimeSafety would be meaningless without the concept of safety-checked undefined behavior.
I would say that making optionals and fat pointers (slices) a part of the type system is possibly more important, but it all combines to give a fighting chance of getting user-controlled resource management correct.
Given the topic of the Fine Article, it's worth briefly noting that `defer` and `errdefer` are keywords in Zig. Both the test allocator, and the GeneralPurposeAllocator in safe mode, will panic if you leak memory by forgetting to use these, or rather, forget to free allocations generally. My impression is that the only major category of memory bugs these tools won't catch in development is double-free, and that's being worked on.
—— I'm still entirely unconvinced.
The thing is, wrap-around is not only well-defined, it's common, and EXPECTED.
Example:
static inline u32 __hash_32_generic(u32 val)
{
return val * GOLDEN_RATIO_32;
}
and dammit, I absolutely DO NOT THINK we should annotate this as some
kind of "special multiply".
—-Full thread: https://lore.kernel.org/lkml/CAHk-=wi5YPwWA8f5RAf_Hi8iL0NhGJ...
Signed integer overflow, on the other hand, is undefined. The compiler is allowd to assume it never happens and can re-arrange or eliminate code as it sees fit under that assumption.
How many lines will this code print?
for (int i = INT_MAX-1; i < 0; ++i) printf("I'm in danger!\n");No, it's really not. Do this experiment: for the next ten thousand lines of code you right, every time you do an integer arithmetic operation, ask yourself if the code would be correct if it wrapped around. I would be shocked if the answer was "yes" in as much as 1% of the time.
(The most recent arithmetic expression I wrote was summing up statistics counters. Wraparound is most definitely not correct in that scenario! Actually, I suspect saturation behavior would be more often correct than wraparound behavior.)
This is a case where I think Linus is 100% wrong. Integer overflow is frequently a problem, and demanding the compiler only check for it in cases where it's wrong amounts to demanding the compiler read the programmer's mind (which goes about as well as you'd expect). Taint tracking is also not a viable solution, as anyone who has implemented taint tracking for overflow checks is well aware.
For the kernel, which deals with a lot of device drivers, ring buffers, and hashes, wraparound is often what you want. The same is likely to be true for things like microcontroller firmware and such.
In data analysis or monte carlo simulations, it's very rarely what you want, indeed.
For example, I opened up https://elixir.bootlin.com/linux/latest/source/drivers/firew... as a random source file in the Linux kernel, and I didn't see a single line where wraparound would be correct behavior.
There are definitely cases where wraparound behavior is correct. There are also cases hard errors on overflow isn't desirable (say, statistics counters), but it's still hard to call wraparound the correct behavior (e.g., saturation would probably work better for statistics than wraparound). There are also cases where you could probably prove that overflow can't happen. But if you made the default behavior a squawk that wraparound occurred, and instead made developers annotate all the cases where that was desirable to silence the squawk, even in the entire Linux kernel, I'd suspect you'd end up with fewer than 1000 places.
This is sort of the point of the exercise--wraparound behavior is often what you want when you think about overflow, but you actually spend so much of your time not thinking about it that you miss how frequently wraparound behavior isn't what you wanted.
If wraparound is ok for that particular multiplication, tell the compiler that. As a sibling comment says, this is seldom the case, but it does happen, in particular, expecting byte addition or multiplication to wrap around can be useful.
The actual expectation of the vast majority of arithmetic in a computer program is that the result will be correct in the ordinary schoolyard sense. While developing that program, it should absolutely panic if that isn't the case. "Well defined" doesn't mean correct.
I don't understand your objection to spelling that `val *% GOLDEN_RATIO_32` is. When someone sees that (especially you, later, coming back to your own code) it clearly indicates that wrapping is expected, or at least allowed. That's good.
For example: Rust will silently wrap signed integers in release mode even when it’s considered a bug and crashes in debug mode.
You have missed my point.
Every other example you mention is done by rust in release mode and the performance impact is minimal, so I would say it's a good counterexample to your claims that defining these things would hamstring performance (signed integer overflow especially is an obvious no-brainer for defining. Note that doesn't necessarily mean overflow checks! Even just defining the result precisely would remove a lot of footguns).
This is wrong, because you would define them to have the behavior that the architecture in question does, so no changes would be needed. For integer division this would mean entering an implementation-defined exceptional state that does not by default continue execution (on Linux, SIGFPE with the optional ability to handle that signal). For dereferencing a pointer, it should have the same semantics as a load/store to any other address--if something is there it works normally, if the memory is unmapped e.g. for typical Linux x86 programs you get SIGSEGV (just as you would for accessing any other unmapped address).
Suppose now, there are two architectures with slightly differing behavior.
Can the compiler still optimize signed x + 1 > x to true?
And even then there are tools to help define much of that - if you want well defined wrapped signed integers, great. If you want to trap on overflow, there's an option for that. Lots of compiler warnings and other static analysis tools (that would just be default-rejected by the compiler today if it didn't have historical baggage, but they exist and can be enabled to do that rejection).
Yes, there's many issues with the ecosystem (and tooling - those options above should be default IMHO), but massively overstating them won't actually help anyone make better software.
And other languages often have similar amounts of "undefined behavior" - but just don't document it as such, relying on a single implementation being "Defined Correct", and hope they're not actually being relied on if anything changes. Just like C, only undocumentated.
This is what I mean by it becoming "meme" - things like "Undefined Behavior" or "Memory Safety" have become a discussion-ending "Objective Badness", hiding the real intent - being "Languages I Do No Like" (or, most often, are a poor fit for the actual job I'm trying to do. Which is fine, but not rejecting that those jobs actually exist).
But they mean real things that we can improve in terms of software quality, and safety - but that's rarely the intended result when those terms are now brought up. And many things we can do right now with existing systems to improve things, to not throw away huge amounts of already well-tested code. To do a staged improvement, and not let "perfect" be the enemy of better.
> You don't want C, you want a language that actually gives defined semantics to all combinations of language constructs.
So, Zig?If a function has an error type (indicated by a ! in the return type), you have a few options. You can use `result = try foo();`, which will propagate the error out of the function (which now must have ! in its signature). Or you can use `result = foo() catch default;` or `result = foo() catch unreachable;`. The former substitutes a default value, the latter is undefined behavior if there's an error (panic, in debug and ReleaseSafe modes).
Or, just `result = foo();` gives `result` an error-union type, of the intended result or the error. To do anything useful with that you have to unwrap it with an if statement.
It's a different, simpler mechanism, with much less impact on performance, and (my opinion) more likely to end up with correct code. If you want to propagate errors the way exceptions do, every function call needs a `try` and every return value needs a ! in the return type. Sometimes that's what you need, but normally error propagation is shallow, and ends at the first call which can plausibly do anything about the error.
It also has tagged unions as a general mechanism for returning one of several enumerated values, while requiring the caller to exhaustively switch on all the possibilities to use the value. And it has comptime generics ^_^. But it doesn't use them to implement optionals or errors.
However, if I were to request a feature to the core language it would be: NAMESPACES. This would clean up the code significantly without introducing confusing code paradigms.
I guess I should have reworded. I don’t expect that feature in C, but if I were to reinvent C today I would keep it the same but add namespace and mangling.
Adding an explicit prefix to every function call is a lot boilerplate when it’s all added up.
In fact on some targets the assembler name of identifiers doesn't always match the C name already.
Although as someone almost always explicitly qualifies names, typing foo_bar is not very different from foo::bar; the only minor advantages are that you do not have to use foo:: inside the implementation of foo itself and the ability to use aliases.
surely not. How do you differentiate these two functions?
void fooN(void);
namespace N { void foo(void); }You would mangle it as something like foo$N depending on the platform.
Yes, nothing like that is possible in C
I assume you haven't looked at the expansion of errno lately?
edit: also
whereas you can see the user-defined macro definition of "b" at the top of the file. you can't blame the c language for someone choosing to write something like that. sure it's possible, but its your choice and responsibility if you do stupid things like this example.
- macros are also standard C++ features too, so this point doesn't differentiate between those languages
- i'm failing to adequately communicate my point. there's a fundamental difference practically and philosophically between macro stupidity and C++ doing things under-the-hood. of course a user (you, a co-developer, a library author you trusted) can do all sorts of stupid things. but it's visible and it's written in the target language - not hard-coded in the compiler.
yes - sure, good luck finding the land-mine "b" macro if it was well buried. but you can find it and when you do find it, you can see what it was doing. you can #undef it. you can write your own version that isn't screwed up, etc.
you can do none of those things for operations in c++ that occur automatically - you can't even see them except in assembly.
I specifically reject this. Constructors, exceptions, and so on are as similarly visible at the source level as macro definitions.
And thanks to macros, signal handling, setjmp, instrumentation, hardening, dynamic .so resolution, compilers replacing what look like primitive accesses with library functions, any naïve read of C code, is, well, naïve.
I'm not claiming C++ superiority here [1], I'm trying to dispel the notion that C is qualitatively different from C++ form a WYSIWYG point of view, both theoretically and in practice.
[1]although as I mentioned else, other C++ features means that macros see less use.
but i will also emphatically reject your position: "Constructors, exceptions, and so on are as similarly visible at the source level as macro definitions"
no they are not. you can certainly see what the macro is doing - you see it's definition, not just it's existence. whereas in c++ you have to trust that language/compiler to:
- build a vtable (what exactly does this look like?)
- make copy ctors
- do exception handling.
- etc.
none of these are explicit. all of them are closed and opaque. you can't change their definition, nor add on to it.
at issue at hand is both "magic" and openness. c gives relatively few building blocks. they are simple (at least in concept). user libraries construct (or attempt to construct) more complex idioms using these building blocks. conversely c++ bakes complex features right into the language.
as you note, there are definitely forces that work against the naïve original nature of c. macros, setjmp, signal handling, instrumentation, hardening, .so resolution, compilers replacing primitive accesses, etc. but all of those apply equally to c and c++. they are also more an affect of the ABI and the platform/OS than either language. in short, those are complaints and complexities due to UNIX, POSIX, and other similar derived systems, not c or c++ the language itself.
c has relatively few abstractions: macros, functions, structured control flow, expressions, type definitions. all of these could be transformed into machine code by hand, for example in a toy implementation. sure a "good" compiler and optimizer will then mangle that into something potentially unrecognizable, but it will still nearly always work the way that the naïve understanding would. that's why when compilers do "weird" things with UB, it gets people riled up. it's NOT what we expect from c.
c++ on the other hand has, in the language itself, many more abstractions and they are all more complex. you aren't anywhere near the machine anymore and you must trust the language definition to understand what the end effect will be. how it accomplishes that? not your problem. this makes it squarely a high-level language, no different than java or python in that facet.
i explicitly reject your position that "that C is qualitatively [not] different from C++ from a WYSIWYG point of view, [either] theoretically [or] in practice."
to me, it absolutely is. it represents at lower level interface with the system and machine. c is somewhere between a high-level assembler and a mid-level language. c++ is truly high-level language. yes, compilers and os's come around and make things a little more interesting than the naïve view of c in rare cases . but c++? everything is complex - there is not even workable illusion of simplicity. to me this is unfortunate because, c++ is still burdened by visible verbosity, complexities, land-mines, and limitations due to the fact that it is probably not quite high-level enough.
this is all very long winded. you and many other readers might think i'm wrong. the reason i'm responding is not to be argumentative, but because it is that it's by no means a "settled" question and there are certainly also plenty of people that see it a very different way. which i think is fine.
1) labels as values in standard 2) control over memory position offsets, without linker script
other than that a few more compiler implementations offering things like checked array bounds, and a focus on correctness rather than accepting the occasional compiler bug
the rough edges like switch fallthrough are rough, but easy to work around. They don't need fixing (-pedantic fixes it already, etc)
maybe more control over assembly generation, such as exposing compilation at runtime; but that is into the wishful end of wishlists
Exactly this. C++ folks should not approach C like a "C++ lite". I appreciate the authors candid take on the subject.
As for defer, there is some existing precedent like GCC and Clang's __attribute__((cleanup)), but - at least for me - a simple "goto cleanup;" is usually sufficient. If I understand N3199 [1] correctly, which is the authors proposal for introducing defer in C, then "defer" would be entirely a compile-time construct. Essentially just a code transformation to inject the necessary cleanup at the right spots. If you're going to introduce defer to C then that does seem like the "best" approach IMO.
[1] https://www.open-std.org/jtc1/sc22/wg14/www/docs/n3199.htm
And yes, I also agree that C++ has WTF insanity, like 17 or so initialisation quirks, exceptions in general (primarily to address failures in constructers, surely there must be a better way, also OOM / bad_alloc is a relic from the past), and unspecified sizes for default built in types (thats C heritage).
Unless you're writing inline assembly or intrinsics or something like that, the semantics of your target architecture are quite irrelevant. If you're reasoning about the target architecture semantics that's a pretty good indication that what you're writing is undefined behavior. Reasoning about performance characteristics of your target architecture is definitely ok though.
its just mean if you need that logic, in C you would write lots of verbose less safe code.
If you miss a destructor event, without configuring the addon "yes I really meant that", the addon halts the compilatoin at best, or returns nonzero for ci at worst.
Edit: I just reread this comment and realized the beginning of it could come across as a bit condescending even though that wasn’t at all my intention. I’d edit it out, but I don’t like doing that, so my apologies if it did come across that way!
It's a bit confusing to have a 'thing' mention one mechanism in its name, but actually being valuable by ensuring some other mechanism
The reason to do this is precisely so that the resource can be cleaned up at destruction of the object. So even if you had an acronym like RASALTOI, it would still probably be misleading
Indeed! When I was first learning C++, I found the term "RAII" quite confusing too. However, after years of experience with this term, associating "RAII" with its intended meaning has become second nature.
Having said that, there is at least one way to make better sense of "RAII" and that is considering the fact that in RAII, holding a resource is a class invariant. The resource is acquired during construction (initialisation) and released during destruction (which happens automatically when the object of the class goes out of scope). Throughout the object's lifetime, from construction to destruction, maintaining possession of the acquired resource is an invariant condition.
Although sounds simple in principle, this can get complicated pretty quickly, especially in the implementation of the copy assignment operator where we may need to carefully delete an existing resource before copying the new resource received by the operator. Problems like this led to formulating more techniques for carefully managing the resources while satisfying the class invariant. One such technique is the copy-and-swap idiom.
None of this is meant to justify the somewhat arbitrary term though. In fact, there are at least two better alternative names for RAII: Scope-Based Resource Management (SBRM) and Constructor Acquires, Destructor Releases (CADR).
You are proposing to change the C language. The risk is great even the smallest change will break the existing code. If you can't convince all of the stakeholders, it's better not to change it. Keep the status-quo.
Oh man, I hear ya. And in a lot more domains than computer language design. Is it inexperience? Impatience? The tendency for search results to be filled with low-quality and high-recency content? The prioritization of hot-take blog posts and Reddit comments over books?
there is dedicated mechanism to achieve RAII-likeness in .NET: try-finally construct
There is no such thing as IL finalizers. There are object finalizers which are highly discouraged to be used on their own.
Their most frequent application is a safety measure for objects implementing IDisposable where not calling Dispose could lead to memory leak or other form of resource starvation that must be prevented.
For example, a file handle is IDisposable, so it is naturally disposed through using statement but should a user make a mistake in a scenario where that handle has non-trivial lifecycle, once the object is no longer referenced, its finalizer will be called upon one of the Gen2 GCs by a finalizer thread, preventing the file handle leakage even if its freeing is now non-deterministic:
// Disposed upon exiting the scope
using var okay = File.OpenHandle("file1");
// IDEs will complain if you do this, but if you insist,
// the implementation will prevent you from shooting your
// foot off even if it'll hurt until next Gen2 GC
var leaked = File.OpenHandle("file2");... actually writing code that gets the job done ... in C++.
You can add `defer` instead, but regardless, this has nothing to do with C++. You can implement safety features without having to copy the arguably worst language in the world, C++. I like C++, I wrote many larger projects in it, but it sucks to the very core. Just add RAII to C.
This is a response to people contacting / criticising them asking for destructors instead of defer.
The author even acknowledges halfway through that it’s basically a strawman:
> It’s not a bad argument; after all, the entire above argument hinges on the idea of stealing from C++ entirely and copying their semantics bit-for-bit.
To me, only after that does it engage with the underlying concept in a way which is engaging and convincing. But you’ve had to trawl through 2500 words to get to that point.
{
void *buffer = malloc(SIZE_MAX);
if (buffer) {
if (!do_stuff(buffer)) {
free(buffer);
return;
}
more_stuff(buffer);
free(buffer);
}
return;
}
If you wanted something like that in C it doesn't need to emulate C++ style RAII with classes and strongly typed constructors. It could look like something like, for example, where you just define pairs of allocator and free functions: allocdef void *autobuffer(malloc, free);
...
{
autobuffer buffer(SIZE_MAX);
if (buffer) {
if (do_stuff(buffer)) {
return;
}
more_stuff(buffer);
}
return;
}
The implementation would effectively be a Lisp style macro expansion encoded in the C compiler (or preprocessor) that would just basically write out the equivalent of the first listing above.I find this an interesting thought experiment, basically types that you'd opt in to RAII. Just have a feeling that you'll need to define some notion of ownership to make it work.