Object-oriented techniques in C
dmitryfrank.com
dmitryfrank.com
If this project really did require re-inventing C++ in C, it must be justified by being fairly large. In which case, a low-end ARM (Cortex M based) microcontroller would have been entirely suitable. All ARMs have multiple C++ toolchains which support them.
If this is a microcontroller you end up lumbered with, and didn't choose, then fair enough, but this effort should be labeled for the hack that it is. This is not best or even good practice.
So the software engineers are the only ones on the project who get a say as to what chips get used?
> If this project really did require re-inventing C++ in C, it must be justified by being fairly large. In which case, a low-end ARM (Cortex M based) microcontroller would have been entirely suitable. All ARMs have multiple C++ toolchains which support them.
ARM is not entirely suitable for any application just because you have complex (for some definition of complex) software.
It's not just ARM - there's plenty of other CPU architectures with modern Clang, GCC, or other toolchains with C++ support. Given that this is a large code base (justifying this amount of work), this cannot be a "tiny" microcontroller, as in 8 bit, and needing to run on microamps. I cannot believe that there were not C++-capable microcontrollers available which would have done the job, without breaking the bank.
I am kinda-sorta happy to have C++ at all.
I think.
Cost of the software in embedded systems is quite often low compared to hardware costs. I ballpark lower limit for part cost at $10. I ballpark software development cost at $100k/y or $50/h. Let's say it is reasonable to get software done in a man-year. At 1M units of projected sales, software cost is at 1% of hardware cost (not project cost). At such ballpark estimate software engineers' opinion is only treated as a humble request.
Of course this varies from project to project. Increasing portion of embedded systems are mostly software. Portion of projects have hard-realtime requirements. Here software engineers are benevolent dictators. But that is by no means a general case.
The compiler I use does have an embedded c++ option but it's limited and has some issues working with the RTOS that's supplied with the compiler. Also with a total of 6K of ram available it's usually easier just to stick to writing plain C.
Why not just use the word "compiler"?
If you got rid of name mangling (and a few other C++ features that are besides the point), you could certainly call C++ from C. The thing is - your C calls would look exactly like what the article is suggesting. That's why it's important to learn this technique: it is what your C++ compiler is doing under the hood. Indeed, the very first C++ compilers were just preprocessors that transformed C++ syntax into the type of vtable + base class + first parameter indirection that you see here.
If current C++ to C transpilers don't handle naming well, that's a problem that should be worked on.
one advantage of using something like this is it forces you to really think about what makes something OO and gives you a deeper appreciation of it. i'm a java programmer now, but to this day i still prefer object oriented C to C++
I know you list the resources, but how does this approach add overhead? (how does it use those resources more than a "C" approach would?) Most of the constructs in C++ only cost anything if you use them, so you only pay if you need the feature; further, the runtime costs of most C++ features are pretty much exactly what's required for that feature… so if you were writing C, and you needed the same functionality, how would you avoid paying the same cost?
Also the first technique doesn't feel like OO in any meaningful way. Defining a state object and passing it to functions is exactly what you'd do in say, a functional programming language, or even in regular old imperative code.
Wrong. If some function is not virtual, there's no point to add it to vtable; so, it's just a regular function.
> Also the first technique doesn't feel like OO in any meaningful way.
Really? Of course I can be wrong, but for me, defining a state object and operate on it only through methods (i.e. functions that are given a pointer to state object) is exactly what is called an OO style.
I guess in simple cases like this the difference between OO and other styles are pointless. I'm just thinking that if I were to write it in a functional language, it'd probably a very similar API, though possibly/probably returning a new state vs modifying.
But in the simple case described in op (no inheritance), that is all a class is, with the exception of some syntactic sugar.
By the way, similar approach is used across the Linux Kernel: check, for example, the book "Linux Kernel Development" by Robert Love, especially the discussion on the virtual filesystem.
Most vtable like implementations in the kernel are nothing more than struct of function pointers, and this is almost exclusively.
When I was (or am) coding C, a lot of the OO principles helped me write better C code, as I understood better what the higher level constructs that I was working with was. (i.e., if there's a base class among essentially hiding among my structs that should perhaps be pulled out into a separate struct; do I need a vtable, or perhaps function pointers on the struct; how is a struct initialized, and how is it destructed, etc.)
While I've been a C++ programmer since my uni days, I think algorithms and data structures should come first, and I don't like the idea of "bending" programs to make them fit nicely within the OO paradigm.
I have experience of an early nineties project that was entirely structured like this and trust me, it ends up as a disaster.
C is not meant to be OO, consequently all the tooling in editors and the like do not understand the links between classes and methods.
What it ends up like is an impenetrable mess with all the mechanics of C++ exposed but being impossible to navigate.
again... please do not do this in any project... for the sake of any who will follow you in maintaining your code!
Ended up doing a similar approach when forced to use C instead of C++ for a data structures university project.
Goal achieved, but not an experience to repeat ever again.
Edit: Forgot to mention that so far I only used plain C instead of C++ when obliged to do so.
As ever, we have to apply common sense: as I mentioned in the article, 1-level inheritance is quite good, or maximum 2, but if we have more, it gets too verbose. Luckily, for most cases, 1-level is quite enough to create common API and reusable modules.
Maybe you overuse it? Of course if you try to build huge hierarchies like this, you most likely will end up using typecasts in client code, and the entire thing will quickly become a mess. The article proposes solutions that don't have typecasts in client code.
As I said before, we should apply common sense. Engineering is all about tradeoffs. If I were you, I wouldn't argue against any approach as aggressively as you just did.
The Linux kernel has some OO techniques in C. 2-part story:
https://lwn.net/Articles/444910/
https://lwn.net/Articles/446317/
I don't follow the Linux development mailing lists so I don't know if Linus Torvalds and other respected contributors regard those techniques as "disasters". Maybe they do.
Though, if your project is more of a toolkit, then a strictly procedural style is probably better.
When I hear about C programs that were a disaster because they used OO-like techniques, the operative word is usually "disaster" and not "C" or "OO". In other words, there were other things wrong with the project that had nothing to do with the selection of language or programming paradigm, like hiring inexperienced programmers or mandating a single programming paradigm for the whole codebase.
Good C code will have a lot fewer objects than say, Java, because most of the things you want to use C for, you can accomplish just fine with structs and functions. But you shouldn't shy away from a struct full of function pointers or an ADT with opaque pointer types and appropriately-namespaced member functions just because you're afraid a maintenance programmer won't get it. It's their responsibility to learn the language.
For similar reasons, "cute" macro-heavy C code tends not to survive long-term, because it breaks too much tooling.
I'm currently writing a binary parser library using C++ but without the STL. I'm able to use range-based for loops, zero-cost iterators, and other features without any dependencies on a C++ STL library. After optimization, the generated code is essentially what the C equivalent with all the boilerplate would compile down to. There are plenty of OS kernels written in C++, you just have to pick the appropriate C++ subset. ;-)
Regardless, I enjoy the stricter typing in C++ that generates essentially same code as C with OO tacked on but with less boilerplate.
I can't really think of a more straightforward way to implement CRC32, frankly, especially assuming that maybe you don't want to force the entire buffer to be present. If you did, you could `uint32_t crc32(unsigned char * buffer, size_t len);` — but you can implement that function in terms of the interface in the article, and if you pick up a bit of inlining from an optimizer, I think it would end up being just as efficient, too. (If you don't, you'll end up with some extra pointer dereferences, but again, the streaming interface supports streaming.)
And perhaps to stress what some other commenters are missing: it's not "you should always OO in C"; it's that if you want an interface to compute, while streaming, the CRC32 of a stream — you need to store that state somewhere, and you need to act on that state somehow. That state is an object: it has a setup, and potentially a teardown, and two relevant methods: feed it data, and get the computed hash.
If you want to build a system where you can change the hash function at runtime, you might need some kind of interface or inheritance. But if you end up needing a virtual call per byte (or per a small chunk of bytes), you're killing your performance.
And besides, there's so few hash functions you end up needing in practice that a simple switch-case would do the same with less boilerplate and probably give better performance because compiler optimizations can take place (virtual calls tend to prevent many optimization tricks).
I agree with GP... The choice of CRC was a poor example for an article about OO in C. Something related to filesystems or device drivers would be vastly more practical.
Given some of the other comments, I think this ought to be stressed though… OO doesn't strictly require inheritance, or vtables. A fair number of classes in the C++ STL don't use those features, for example.
It's not that the state is a single integer, either, it's that the concept of a state exists; having a separate type (with the functions surrounding that separate type) exist to embody that concept, and can change a function from an ambiguous "what goes in this uint32_t arg?" to a very obvious md5_state_t.
> or even constructors or destructors (apart from zero-initialization).
Just a nit: MD5 requires initialization that is more than just zero-initialization.
> probably give better performance because compiler optimizations can take place (virtual calls tend to prevent many optimization tricks).
This is about the most concrete argument against vtables (but hardly against OO in general) thus far, and I'd still love more explanation: what makes a branch more optimizable than following a pointer?
Embedded platforms tend to have size limits. Would this increase the size of the source code? Seems a bit verbose to me.
The more interesting question is: how much overhead? And the answer is: it depends on the MCU. For example, some 8-bit PICs don't have silicon support for indirect function call (i.e. by function pointer), so they work around this problem by saving function address to the stack and execution 'return' instruction. It works much slower and code size increases notably. I don't use these techniques on these chips.
But for 16- and 32-bit MCUs that I was working with, it works flawlessly. For most of our projects, the overhead is much less significant than the maintainability we get with this approach. As I said in another comment, engineering is all about tradeoffs.
While no doubt this is true, if you have an indirect function call in C++ OO code, how do you not have it in C? That you had it in C++ implies — I hope — that you needed it. The C code can't simply whisk that need away. (Or, if it can, so can the C++…)