The C Programming Language: Myths and Reality
lelanthran.com
lelanthran.com
The advantage of hiding your state behind an opaque struct with builders and accessors is that you can change the size and layout of said struct without it being a breaking API change. The code remains binary compatible even, no need for a recompile if you're shipping a shared lib. This is something just using private members doesn't achieve since with private members the compiler still knows and uses the layout of the struct, it just forbids access to it.
That's why you can even find C++ libraries use this idiom even though C++ obviously has `private`. It's about having a stable, opaque API.
On the other hand because of this added indirection, there's usually a greater performance hit to accessing these opaque structs since code can't be inlined. With private since the compiler can still see inside the struct, it's able to more aggressively optimize the code. You can also store the objects directly on the stack without requiring malloc.
IMO the right way to have private members in C structs is... to document that members shouldn't be touched directly, perhaps using a special naming convention or embedding the publicly-accessible members in a dedicated sub-struct to prevent confusion.
There is somewhat common PIMPL idiom to work around the binary compat issue. Iirc there were some macros floating around to make it easier to manage.
In my experience, the impl class is usually defined in the main class's cpp file.
Reminds me of that one time when glibc broke the whole of Debian for s390 architecture by changing the fields in the jmp_buf struct (which is public): [0].
There is an anti-pattern where header files are used for all the declarations needed internally by the source file. Including (pasting verbatim with the preprocessor) that file from another module would bring in all the unnecessary declarations.
If I code in Rust or C++ I can use namespacing and public/private to give every single object in the codebase a clean interface, but in C doing that is just frustrating, not to mention potentially inefficient.
// in <opaque_foo.h>
typedef struct opaque_foo opaque_foo_t;
size_t opaque_foo_sz(void);
void opaque_foo_init(opaque_foo_t* foo)
// in your code, which you could write a helper macro for if you were so inclined
char opaque_foo_mem[opaque_foo_sz()];
opaque_foo_t * my_foo = (opaque_foo_t*) opaque_foo_mem;
opaque_foo_init(my_foo);alternatively, gcc supports VLAs in unions, but I don't think clang does, but that makes it extra annoying to do.
edit: apparently you can probably apply the may_alias attribute to the type? Or you could try using transparent_union. No idea if clang supports either...
[1] https://nullprogram.com/blog/2019/10/28/, discussed at the time at https://news.ycombinator.com/item?id=21374863
// in your code, which you could write a helper macro for if you were so inclined
opaque_foo_t *my_foo = alloca(opaque_foo_sz());
opaque_foo_init(my_foo);There is no way for a compiler to infer the alignment requirement of a struct because it does not see its definition. You would always have to align the char buffer by hand but you cannot do it for the same reason compiler cannot.
What you can do is to always greedily align the char buffer to the strictest (largest) fundamental requirement for that platform - in other words alignas(max_align_t).
And this is actually what alloca() does for you by default. From https://gcc.gnu.org/onlinedocs/gcc/Other-Builtins.html
> The object is aligned on the default stack alignment boundary for the target determined by the BIGGEST_ALIGNMENT_macro.
* I may be running in an environment where allocating memory is an asynchronous operation. A callback based API forces me to block, which can cause unpleasant side effects like halting an event loop.
* In my experience, some libraries with custom allocation hooks forget to define one or both of the two basics in the callback signature: a "context" or "user data" parameter, and a way to return an error.
The proper solution is to decouple memory allocation from object initialization entirely. There are two different approaches to this:
1. Expose get_foo_size() and get_foo_align() functions that return the size and alignment that the opaque foo struct needs (at runtime, of course). Then I as the user can allocate that memory, and initialize my opaque foo objects in-place:
size_t foo_size = get_foo_size();
void* buf = alloc_aligned_memory(foo_size * 1000, get_foo_align());
for(size_t i = 0; i < 1000; ++i) {
int err = foo_init(buf + i * foo_size, /* params here */);
}
2. Define foo_init(void* buf, size_t len, ...) which attempts to initialize an opaque foo object in the buffer defined by [buf, buf+len). If the buffer does not have enough space, return an error. Otherwise, return the number of bytes actually used by the object.But yes, opaque structs do enforce that it will be treated as a plain pointer, and the compiler (usually) cannot treat it as an aggregate of variables.
The code that you develop can still access the member fields directly, and those accesses can be inlined and optimized aggressively by the compiler.
That seems like a straw man to me. He never said it was exactly equivalent - he said it provided encapsulation and isolation. 'private' is another language's mechanism which does that but they're obviously not identical.
Is this still true? Link-time optimization, which includes inlining, seems to be all the rage these days.
I'm also not sure that dynamic linkers are typically smart enough to do this type of optimization, but I admit that I'm not familiar with the state of the art in dynamic linking technology.
In C++ this is a popular idiom and a standard technique covered by multiple references under the name pointer to implementation (pimpl).
This is not a C thing beyond the point that C has pointers.
* It does not work on a field, but on the whole struct. Either all fields are public or all are private. The latter case forces you to write getters/setters for properties you want users to access, and that in C can be even more cumbersome as you need to write the definition in the .h and the implementation in the .c.
* It breaks, without an explicit and specific and error message, several actions. As it's mentioned in the article, malloc isn't possible, but neither are copies by value, sizeof() breaks... And those will break with an "incomplete type" error message, not a "this type is intentionally made private", which can add confusion in some situations.
* Completely incompatible with inlining code. Considering a lot of people still use C precisely for its performance, I think this can be a drawback in a lot of usecases.
I honestly think that hiding struct declarations should be done sparingly, and preferably limiting it to cases where it's actually necessary (for example, a library that doesn't expose internal struct fields so the same executable works with different versions of the dynamic library; or proprietary libraries that want to expose as little as possible). In the end it's still easy to bypass, and the distinction between header and code files already provide an indication of which functions you should use and which ones you shouldn't.
struct PrivateFieldsOfMyShinyClass;
struct MyShinyClass {
int somePublicData;
double morePublicData;
struct PrivateFieldsOfMyShinyClass *p;
};
> malloc isn't possible, but neither are copies by value, sizeof() breaksThose are implementation details which are deliberately being hidden.
> Completely incompatible with inlining code.
They are as incompatible with code inlining as public/private modifiers in C++ are. That is, LTO is your best friend here. Also, have you ever tried to maintain binary compatibility with several versions of a third-party C++ library that keeps adding/removing private fields to/from its classes?
However, having only a boolean private/public access state isn't generally satisfying either. It often leads to violation of the principle of separation of concerns when all the functions (methods) acting on certain "private" fields need to live in the same class.
In simple classes, like std::vector, it's possible to get away with private. But in many cases that are more complex than that, it seems to me that the best approach is still to expose the data and to be just very clear about the exact purpose of each member.
> Those are implementation details which are deliberately being hidden.
I know they're being hidden deliberately, but in C++ "private" doesn't break malloc (new), copies by value or sizeof. Or stack allocation, to add to the list.
> They are as incompatible with code inlining as public/private modifiers in C++ are. That is, LTO is your best friend here.
public/private aren't incompatible with inlining in C++. That is, you can call class functions that access private members and the compiler can inline those functions. Also, LTO is not always enabled by default, and doesn't always inline the things you want it to inline.
> Also, have you ever tried to maintain binary compatibility with several versions of a third-party C++ library that keeps adding/removing private fields to/from its classes?
I mentioned binary compatibility as one of the reasons one might want to do this. However, if you have a third party that doesn't care about API compatibility I doubt struct fields are the only thing they're going to change constantly.
At this point though, I think I honestly would prefer setters/getters.
struct MyClassPublic {
int x;
int y;
...
}
/* using it */
MyClass *myclass = myclass_create();
((MyClassPublic *)myclass)->x = 5; struct public_stuff {
...
}
struct private_stuff {
struct public_stuff public;
...
}
struct public_stuff *make_public(/* ... */) {
struct private *prv = malloc(sizeof *prv);
/* ... */
return &prv->public;
}
Then when passed in to functions, cast the passed in struct pointer to the private one.The public struct doesn't even have to be at the start if one make appropriate use of offsetof and ensuring valid alignment.
Nothing new under the sun...
That would make more sense, since you can use the header to craft an interface in which some components are public and some are private.
Clearly you can abstract away the parts of data that the API should not see.
Yes that tracks with the work I've done. It could even be so simple that only one value needs to be exposed through a getter (along with a couple "methods").
An ADC in an embedded system could operate like that.
You can, in a defined way (i.e., no invoking of UB). I just didn't put that in.
They called this "modular programming" where "classes" were represented by opaque pointers and actions could only be performed on those pointers using the functions defined in the module.
What's the benefit of "opaque" pointers?
For background, my C code is almost entirely on microcontrollers. So I'm looking at it from that point of view. If you're talking about event-based applications running inside a full operating system, I've always stepped up to something like C#, so I don't have much experience with function pointers for that kind of work.
But consider the position of developer who implements a shared library that is distributed in binary form only. In this case, the benefit of opaque pointers for the development of a library is that the implementation remains private at the source level. One could, of course, reverse-engineer the binary but few people would do it.
If you define your structures in a public header -- and this includes C++ classes and templates with private members -- one can easily see the implementation and, with a few casts, start munging the guts of your objects and baking in a hard requirement on a specific layout and / or version of your library.
But for dynamic linking, this is how you avoid breaking ABI while maintaining forward flexibility.
Many of the newer languages have modules or packages, but interface well with C, thus give the benefits of what the article was referring to. To include that various newer ones don't use class-based OOP, just as C doesn't.
I didn't mean to imply it is, I lead with:
>> All too often someone, somewhere, on some forum … will lament the lack of encapsulation and isolation in the C programming language. This happens with such regularity that I now feel compelled to address the myth once and forever.
It's only about the myth that C doesn't have any level of encapsulation or isolation.
To be clear: you tried to say "C offers encapsulation/isolation". People read "This solution is equivalent to 'private'", an almost completely unrelated statement, and then respond to that.
That could typically be classified as a "Straw man fallacy"[0], but I believe people who do this in many cases simply do not have the necessary reading skills to understand what proposition has been made, and therefore honestly believe themselves to be reasoning correctly (i.e. without fallacy).
Reading comprehension used to be a topic at school when I was a child. I suppose that's no longer the case??
>That could typically be classified as a "Straw man fallacy"[0], but I believe people who do this in many cases simply do not have the necessary reading skills
Fyi... the author's article that this thread is about has in bold heading: "Myth: C has no equivalent to “private”"
So, a reasonable interpretation of the text following that headline is how to use C Language constructs to dispel that myth.
Doesn't seem like "straw man" applies here.
With all that said, I don't think all that condescending talk about the lack of reading comprehension or skills, without actually going into the arguments themselves, is really necessary or positive.
You are correct; I should change the heading to "Myth: C provides no encapsulation". I don't want to do it now, as the discussion is still ongoing and making this change now while people are commenting on the 'private equivalence" aspect would be gaslighting those people.
> Even then, my points still apply: this isn't really equivalent to how encapsulation works in other languages due to the lack of granularity and the "extras" of all the usual language behavior that stops being supported.
So? The myth the article addresses is "C has no encapsulation", not "C has great encapsulation".
That C provides stronger encapsulation and stronger guarantees for upgrade migration is the specific myth that I am trying to address.
Yes, you can do something in C that somewhat resembles encapsulation. You could also do it in Python with a decorator that inspects the caller. You can also do inheritance and virtual functions in C. You could do anything in a Turing complete language. But usually when talking about "X language has Y" refers to whether the language has Y as part of the specification. In this case it isn't encapsulation as part of the specification (that's why it comes with all those downsides/extras) but as an artifact of other aspects of the language, that's why I say it's not really equivalent.
struct Thing { int i; };
In an implementation c file: struct _Thing
{
struct Thing;
int b;
};
struct Thing *new_thing()
{
struct _Thing *thing = malloc..
thing->i = 0;
thing->b = 1;
return thing; >* Completely incompatible with inlining code
You were right 10 years ago. Today we have LTOIn other words, the 2 different possible receptions to your post:
- YES, file-level modularity with opaque structs is _equivalent_ to class private members --> for those mindsets already sympathetic to C Language
- NO, using file-scoping rules and structs is not equivalent to class private members because it's a bunch of extra ceremonial syntax to implement a workaround. (The "Turing Tarpit"). It's using the opaque struct as a "design pattern" and as Peter Norvig famously said, "Design patterns are bug reports against your programming language."
C++ style (without "pimpl") requires recompilation of the whole dependent tree when adding a new private member function. It's encapsulation only in a formal sense
For the latter, the industry best practice is to avoid malloc(), except maybe at init time, and instead allocate memory statically. And in that use case, you break your code into modules, which can contain private data, public data, private functions, and public functions.
In other words, building an app out of C modules is a lot like building an app in a more modern language just using static classes, with no instantiation. And that design pattern — which is extremely common in the embedded world — we have a direct equivalent to the “private” qualifier, which is “static”, which restricts the rest of the app from accessing so-marked file-scope variables and functions.
Where this breaks down — as always with C — is when you need multiple instantiations of a module, which modern programming languages refer to as an object. The closest we can get in C is to pass the module’s public functions a struct with some sort of data structure containing the object’s n9n-static data. And the author explains, there are standard ways make that data structure opaque to calling code, but those are definitely workarounds to language shortcomings.
But the bottom line is that those language shortcomings — the lack of objects and a private qualifier for its members — are only shortcomings if you need those features, and in the embedded world, most applications don’t, they only require all the advantages offered by C. So as always, this is about picking the right language for the project, there’s no one size fits all.
Encapsulation - Hidden
Inheritance - Same fields in beginning
Polymorphism - Tagged struct
I mean, that's the point of everything, isn't it? giving names and defining good practices so that everyone can benefit from them? because even today you can find a LOT of codebases which are an imperative mess
But the reality is that the best practices of imperative programming where already much like object oriented programming and that OOP is a formalization of those practices.
That is why Oberon takes the spartan approach of only having extensible types, everything else is just like in Modula-2. Later descendants adopted a more mainstream approach.
Likewise how OOP is done in Ada or Modula-3 isn't quite like in mainstream approach.
Or when modules can be manipulated like variables, and given type signatures, we get Standard ML functors, with overlapping capabilities to OOP.
[1] https://dl.acm.org/doi/book/10.5555/1243380 (lousy interface, but PDF download available)
It's okay to admit C has shortcomings.
I did that, didn't I?
>> Just to be clear, C is an old language lacking many, many, many modern features. One of the features it does not lack is encapsulation and isolation.
These are my favorite types of debates and while I disagree with it for much better articulated reasons already mentioned, primarily the exploitative nature of header files used as security through obscurity, I think it stirs up a lot of debate and keeps the spaghetti meatball of knowledge around patterns/anti-patterns, "best-practices", etc. moving forward in a much more passionate way.
I think this is the truest sense of the term "hacker" which lives up to the title of the site pushing us into these debates. Putting stuff together that doesn't always work as intended or expected but arguing for and against it.
A long winded thank you to everyone, OP and all the threads responding.
I don’t know a lot about C but this interests me. Can anyone point me to where I might learn more about how this works?
You can give callers the public version and then cast it to the private version for internal usage.
I guess my point is that there’s a whole bunch of caveats in regards to safety that need to be considered in your solution.
Modern PL features are more or less wrappers around comparatively complex C code. Object inheritance is actually one that isn't too difficult or complex to implement: https://www.youtube.com/watch?v=443UNeGrFoM&t=4275s
True, and if you implement it as a preprocessor, that's exactly how C++ started in 1979 (Stroustrup's "C with classes" Cfront pre-processor: https://en.cppreference.com/w/cpp/language/history).
Of course, you could do much the same thing if you didn't use templates in C++, or if you were very disciplined and limited with your use of templates, but this seems to go against the grain of how I've seen C++ used.
I haven't used C++ much since modules were introduced. From doing some quick reading, it seems unclear whether they significantly improve compilations times in practice. I wasn't able to find anything which addressed how they relate to the ABI... I suspect this is a pretty complex topic. If you have any good references related to either of these points, I'd be very eager to read them.
Either way, "C++ with modules" appears to me to be unlikely to clear the bar set by "C with opaque types" (which, for all intents and purposes, can be done in C++) in terms of 1) ABI stability and 2) encapsulation. I consider point (1) to be related to point (2), since details which are leaked into the ABI are not encapsulated.
Importing the whole standard library as defined by C++23 import std; is quicker than doing a plain #include <iostream>, as shared per Microsoft employees in some talk they have done regarding upcoming C++23 support.
Unfortunely I can't remeber which one it was, but someone else can glady share the link.
Second by having template details marked as private on the module metadata, that isn't directly exposed to consumers.
As for the ABI specifically, that is compiler dependent anyway.
And you've just made it impossible for the users of your StringBuilder to pass it around by-value. Every instance has to be malloc'ed by your library, even though it's just a tiny, word-sized struct. Awesome! And each access needs to go through an additional pointer indirection. All this just to pretend that C supports proper encapsulation. Hooray!
I'm sorry that I'm targeting your blog post specifically, but it's just so stereotypical of C proponents, that can't (or won't?) realize that their favorite programming language is inherently limiting and limited along several very important dimensions. It makes me think that although some of them might be excellent programmers, they make for terrible software engineers.
Further, and quite separate, the term "software engineer" was coined on behalf of assembly programmers, whose language lacks even further features.
Finally, I am not sure what you mean when you say they make bad software engineers. Perhaps we have different definitions of software engineer.
Which also validates that any university teaching software engineering has a certain quality level, and portofolio of lectures, to create a general background across all subjects of engineering practices besides writing code.
Could you please give some examples of countries where "software engineer" is a protected title, together with the organizations responsible for conferring it? I think I looked it up in the past but wasn't able to find one. All I was finding was things like civil engineering certification (basically search engine garbage in the case of my search -- the quality of search engine results has really gone down the drain in the past 15 years or so, sadly).
No, C is limited because it's mutually exclusive: either encapsulation, or zero overhead by-value passing. Other languages, like C++ or Rust, allow both at the same time.
> You can do dummy defines of structs with the same size
Don't forget alignment. The general pattern is: https://godbolt.org/z/6je9Yb3rf.
If you have:
struct foo;
in foo.h and: struct foo {
int a;
double x;
...
};
in foo.c, you won't have sizeof(struct foo) available from bar.c, so your construction won't work in bar.c. You could define a function: size_t sizeof_foo(void);
which just returns sizeof(struct foo) from inside foo.c, but since this size is now only known at runtime, you'll need to resort to alloca or VLAs...In the .c file you define private_foo, same a public_foo except that the byte array is replaced with the actual members.
You static assert that size and alignment match and cast at function boundaries.
You hope not to have violated strict aliasing rules.
This is not completely unlike type erasure with small buffer optimization done by some c++ classes like std::function.
The big downside here is that you’re leaking the size of the details into your ABI which wouldn’t happen with a fully opaque type… I could see some uses for it but haven’t felt a strong enough need to reach for it before, although it has occurred to me.
Specially great in IoT instead of macros accessing directly IO ports.
In Modula-2, I can specify an API in its DEFINITION module, and after compiling it, client applications can use such an API without the implementation being ready yet, and still I can compile the client and check if it is free of syntax errors.
...and?
I'm guessing the implication here is that you'll trash performance by doing this. How can you assume that? The thing about optimizing code is, you don't know where your hot paths are until you profile your code. And, the one thing experience has taught me, my intuitions about what the hot spots will be rarely match reality. There's nothing wrong with that, complex systems are complex, and we have incredible profiling tools to eat through that complexity and highlight the hot spots for us.
Now, I know you'll probably go on about a death by a thousand cuts etc. The thing is, well-crafted modules typically don't encapsulate on a fine-grained level. You usually have larger systems that hide details. These systems are usually used a fraction of the time the rest of your program is. So the indirections end up usually being a very insignificant cost to the overall program.
And if you are coding in such a way that copying a string builder by value and/or the indirection imposed by encapsulating that information is a bottleneck, I highly doubt that "fixing" this by copying by value and/or removing the indirection will suddenly make your entire program performant.
> It makes me think that although some of them might be excellent programmers, they make for terrible software engineers.
You haven't actually highlighted any issues here and then go on to finish your argument with an ad hominem. Instead of attacking the competence of C programmers, you should illustrate the actual real world impact that this design philosophy results in. I know plenty of really slow Java libraries, and plenty of really fast C libraries that use this method of encapsulation. So if your argument is that using this method trashes performance, it's a poor argument that doesn't have many real world examples (unless you know of some off the top of your head).
Absolutely! So, you profile your program, and it turns out that 95% of the runtime is caused by malloc/free in tight loops, which you can't get rid of, because they're hidden behind an API which had to choose between encapsulation and efficiency.
> And if you are coding in such a way that copying a string builder by value and/or the indirection imposed by encapsulating that information is a bottleneck [...]
You don't seem to realize that the StringBuilder was just an example to illustrate this style of encapsulation? Oftentimes you want to encapsulate actual "value structs", where it is sensible to create millions of them in an array. In C, you're forced to choose between following good software engineering practices (=> encapsulation) and getting good performance.
I have literally never run into a library that was written so badly that using the library encouraged you to use the API to create millions of small objects. That's what I'm saying. Sure, this can happen, but in reality I've never seen it. Can you show me where this hypothetical scenario is occurring and trashing people's performance? We probably want to avoid using those libraries.
Instead, I usually see encapsulation used like it is in GLFW, or libcurl, or stbi. The encapsulation covers systems and not tiny objects, which encourages the user of the library to not make API calls millions of times or construct millions of tiny objects.
> You don't seem to realize that the StringBuilder was just an example to illustrate this style of encapsulation? Oftentimes you want to encapsulate actual "value structs", where it is sensible to create millions of them in an array.
I did realize this. Encapsulation is typically useful on larger systems. Once you get to the point of millions of objects, you usually have a larger system managing those millions of objects. And ideally, those millions of objects should be POD. If they're POD, encapsulating the data makes no sense at that point, because it makes more sense to encapsulate whatever is managing that data.
> In C, you're forced to choose between following good software engineering practices (=> encapsulation) and getting good performance.
This is a false dichotomy. There are plenty of large C projects that follow good software engineering practices (which is entirely subjective, what is "good"?). Look at any OS kernel, or the libraries I mentioned above.
So, once again, I'm curious if you know of any C libraries (ab)using encapsulation in the hypothetical scenario you've laid out. If there aren't any libraries that do this, then this is a non-issue and attacking the competence of C developers is entirely unwarranted since you've built up a strawman that doesn't exist in reality.
Most languages pass objects by reference (C# and Java chief among them).
Still, if you really want to pass by value -- even though you'll likely end up with pointer ownership problems -- you just add a few functions to the API to do so.
Creating an opaque type on the stack can be done, you just need a little more work.