Comparing C and C++ usage and performance with a real world project
nibblestew.blogspot.com
nibblestew.blogspot.com
As for the "memory leaks" --- I haven't looked at the source, but something whose runtime is very short-lived, like pkg-config, may be very well justified in allocating and never freeing, letting the process exit itself be the "ultimate free". I've seen and done this many times myself.
I've seen projects that turned from simple and straightforward to buggy (and harder to debug), slow, and bloated because someone decided they wanted to "use C++" and would try to make use of as many "modern C++" features as they could.
Converting an existing C program into C++ can yield programs that are as fast, have fewer dependencies and consume less memory. The downsides include a slightly bigger executable and slower compilation times.
My experience has been the complete opposite.
Highly recommend reading that short tale! This lore is slowly being forgotten, thanks for that link.
(Is there a canonical lore repository for this kind of thing, other than the Jargon file?)
Not to mention they added even more HW to work around the leaks. No wonder those projects always run over budget.
When exactly did pkg_config became missile software?
As they say in meme-land, "that escalated quickly".
Here's a novel idea: how about the appropriate level of effort/time/YAGNI-stuff based on the domain?
Or do you write one-time scripts with MISRA rules?
But I still think not freeing is pretty lame even in short-lived tools. If one is using a no-op deallocator at least the code is designed properly and could be repurposed.
The discussion is at https://news.ycombinator.com/item?id=14233542
This makes sense, and then somebody has a vision for a use case beyond the original imagination, goes to turn the code into a library and spends countless hours smartening up lazy resource management. Whether the original author should be "more responsible" is open for debate, but I've personally run into the above situation, and only mention for another perspective on The Life of Code.
Edit: s/coffee/code/
e: it may be worth noting that I can't for the life of me remember what the I and D stand for... clearly I am much more of a KISS kind of programmer.
That's THEIR problem then.
Why should the original author care for that?
It is their problem, clearly.
> Why should the original author care for that?
That's the question. Don't you think it's easy to think of reasons, though? What if the person porting the code to a library was a later version of the original author?
I said it's lazy resource management. I'm really not trying to condemn, here, just provoke questions.
Then they know what they need to add, and can add it now that they need it -- instead of having it slowing them down when they didn't need it
Keyword: "can".
To be fair in the other direction, it's a minority of interesting software projects that don't care about memory leaks. Also, there are ways to "skip" freeing memory in C++ safely using appropriate allocators. In actuality, you'd have your allocator, in its destructor, clean up itself and all its objects at the same time. It's nearly the same performance without sacrificing correctness or lowering standards with respect to leaking memory.
But sometimes the code grows and what was once a stand-alone executable is about to become a component in a larger executable. With C++ you get a correct component out of the box (if you use RAII consistently), but with C you have to audit all of the code and clean up the leaks.
If one must use this "trick", it's better to think things through, add the appropriate release calls and then somehow replace the release function with a no-op.
Anyway, you're welcome to take a C++ project and translate it to a fast C project with fewer dependencies and lower memory consumption. Then we can discuss facts instead of your personal opinions.
It is not unusual to have allocate-only heaps for exactly these reasons.
I've noticed that using a lower-level language, where abstractions have to be built explicitly, tends to cause one to rethink the problem and often come up with an even simpler and more efficient solution, by approaching it from another direction which using a higher-level language may not even allow.
An example of this I encountered several years ago was with several coworkers who were trying with utmost effort to optimise a piece of code which the profiler had indicated was taking a substantial amount of time --- and a lot of it consisted of memory allocation and copying. They tried lots of "classic" tricks like unrolling, inlining, even reorganising the layout of several classes in an attempt to be more cache-friendly. I looked at the algorithm and realised rather quickly that the code in question was not necessary at all; some trivial modifications to code elsewhere which was using it and deleting that code completely resulted in 30x faster performance and 1/10 memory usage. Due to their background, my coworkers were stuck in the mindset that it was necessary to perform all that convoluted processing, and neglected to see the bigger picture.
So if what's written here is true, I may be unwittingly baking bad practices into my C++ knowledge as a direct result of trying to accomplish the exact opposite…
Which leads me to the question: what is the "right way"? I've seen highly vocal critics of writing C++ as "C with extras", so I assume some middleground is where I need to target?
Generally there's a lot of emphasis on RAII, clear ownership semantics, leveraging more of the standard library as its grown, using lambdas, avoiding shared mutable state, judicious but not excessive use of inheritance and in particular avoiding implementation inheritane, encapsulation.
It's not so much a middle ground in the sense you are thinking. The people I'm talking about don't advocate developers, particularly non expert, going crazy with templates. People do that on their own.
Hope that helps.
The one and only thing that muddies this picture is other people. Once you are collaborating on a program(whether on the same team, through end-user code, or through a library or API call) all your tricks, preferences and conventions are subject to other people's inept groping and misunderstanding. And that is where you get into standardized best practices. They are basically guaranteed to not actually be the best practice, but they're the one you can compromise on.
Citation needed? Better developers are more aware, in any language. There are some cases where idiomatic C++ may introduce more indirection over C (though I can't think of any); there are plenty where idiomatic C introduces more indirection than C++. However, with a little more effort and awareness, the faster and more maintainable solution is always accessible.
> I've noticed that using a lower-level language, where abstractions have to be built explicitly, tends to cause one to rethink the problem and often come up with an even simpler and more efficient solution, by approaching it from another direction which using a higher-level language may not even allow.
Having lots of developers, re-implement many things, mostly just results in much more buggy code. Getting things exactly right is hard. Having a bigger standard library and safer abstractions is a huge edge.
I don't think your anecdote has anything to do with C vs C++. I think basically some negative experiences with so-so C++ devs has colored your thinking rather than technical reasons.
FWIW, this is in my bashrc along with a myriad of other programs that have lame defaults: alias valgrind='valgrind --leak-check=full --show-reachable=yes --track-origins=yes --track-fds=yes --error-limit=no'
The C++ standard library has more features though.
You can think of it like treating C as a garbage collected language, except the garbage collection cycle occurs only once at the end of the program :P
It really can be an effective trick. Deallocation isn't free, and under certain loads can be quite expensive.
The 1000+ leaks in the C version might actually be what's giving it the slight run-time advantage.
I've seen programs building ASTs with hundreds of millions of nodes, where all the nodes were allocated by a separate malloc call, and ref-counted... More than one-third of the startup time (which was counted in minutes) was calls to malloc and free. Some optimizations were made, but in the end we ended up reducing the size of the AST instead of fixing the allocations.
Last time I did any serious parser work I used a pool allocator so I could free all the nodes at once, so allocation was just a compare + increment operation. Although that was forced on me by the difficulties of error recovery in yacc.
Side note wrt deallocations being somewhat slow: could something like Boehm conservative GC speed that up, by grouping all the deallocations together, or by doing them on a separate thread?
https://lists.gnu.org/archive/html/coreutils/2014-08/msg0001...
That's about as "C" as C++ is. Why not Gneural or libpng or even GNU make ?
http://git.savannah.gnu.org/cgit/gneuralnetwork.git https://github.com/glennrp/libpng http://git.savannah.gnu.org/cgit/make.git/tree/
https://accu.org/content/conf2015/DanSaks-Embedded%20Program...
Language Design Implementation Relative Performance
either any inline 1 (fastest)
C++ polystate non-inline 1.56 x fastest
C++ bundled non-inline 1.65 x fastest
C polystate non-inline 1.70 x fastest
C bundled non-inline 1.79 x fastest
C++ unbundled non-inline 1.82 x fastest
C unbundled non-inline 1.95 x fastest
He furthermore argued that the biggest mistakes C++ developers did to kill the adoption of C++ for C programmers was to diverge from the previous line of "C++ is a better C" to "if you're using C++ as a better C you're doing it wrong"https://www.youtube.com/watch?v=D7Sd8A6_fYUI
(I have no skin in the game, I was just curious to see if it's worth looking at rust for embedded when I came across that talk)
https://dlang.org/blog/2017/08/23/d-as-a-better-c/
This is not in the sense of tossing away C coded programs wholesale and rewriting it in D, but incrementally using D here and there for parts of a C program. That way, you've always got a working, usable program.
if (existsCoffee)
writeln("Drink coffee");
http://ddili.org/ders/d.en/if.html
Sad that you adopted one of C's worst features. Why? Can you get rid of it and mandate the bracing every block?Maybe there are bigger issues with C? But that's a different discussion. I want to know why you copied something as simultaneously horrendous and useless as unbraced blocks? If you just didn't think it through and that's the way languages syntactically similar to C have always done it, ok. I'm sure I've made worse mistakes. But please call it one way or the other.
As far as correctness and safety goes, this is still true. It's difficult to scale systems-level programming to large teams. C++ gives the opportunity for more explicit semantics and more aggressive compile-time checks. C can scale well and can be used safely, but you need to do a lot more through convention (always call xyz_Create and xyz_Destroy in pairs!) and through runtime checks (calls to assert, unit testing).
D, Rust, OCaml, and a few other projects are interesting in this space since they provide some of the same benefits as C++ with respect to correctness and safety. Some are plausibly better in theory, though I'm not aware of huge, say, Rust projects that approach the size of huge C++ ones.
Some of the structural advantages of C++ over C can be achieved in C by using generative programming for example and building in automatic mechanisms to ensure there are no memory leaks for example. In other words, the C++ approach to structuring programs is not the only way to achieve the benefits that that structuring implies. It's just really easy to do it that way.
I was just saying that full-blown code generation isn't merely writing in the same language but adopting a DSL as well, so we're not strictly comparing languages at that point.
Here's one way it could be done simply. Let's say you wanted to automate the process of memory allocation and deallocation. You would need a way to describe to the code generator the mepory requirments of your structure. For that you would need a description outside of C. But that description could be embedded into the comments of your code and your code generator be designed to parse those comments to determine what needed to be done.
Knuth also came up with the idea of Literate Programming in which the description of a program is embedded as Latex in the code. This could work in a similar way. So, while you would use a DSL, the description would be inline with your code so the authoring process would be integrated and not 2-stream.
Sure it was for Bjarne, that is why he created it in first place.
After being forced to re-write his thesis from Simula to BCPL, he swore never having to deal with such low level languages again.
C with Classes was his solution to not having to write C directly, after he got his job at AT&T.
Can you name a correctness and safety benefit that C++ has that these programming languages do not?
And, on a pedantic level, C++ competitors can't provide exactly the same benefits of C++ because they took different approaches.
What's a key design difference among these languages? Well, C++ can mostly just #include a C header file and go with it. The other languages provide FFI mechanisms, but they each require declarations of the FFI to match the compiled C code. So theoretically there's a little more room for errors in that translation, though I doubt that's a big concern on the whole. Each of those languages have more mature module systems, which should more than make up for keeping FFI interfaces in sync with C headers.
It was a very bad choice to choose a program based on Glib for this kind of experiment.
Not really, with ABI issues and compiler incompatibility widely used C++ libs are either header-only, or have an "extern C" version of the public API. Id say C++ makes reuse much harder.
2) Being "header-only" is no impediment to code reuse.
I'm no expert on C++ and I've been considering using it for several projects.
An important thing for my needs is being able to define classes in one shared object and create new subtypes of those classes in another, possibly defining overrides on virtual methods and such.
A good friend of mine has said similar things as you - that the ABI issue has not been a major obstacle for some time.
And yet, as much as I search, I still find the same-old advice: Don't use STL types in your interfaces or throw exceptions across module boundaries.
If all the compilers used for a given platform follow the same ABI, would using a separate and specific STL implementation (say, STLport) instead alleviate that particular issue?
Sorry if this question seems a bit rambley but I'd really love to find out how to use C++ in the way I've mentioned.
STL, I guess, is more used on Linux. I would advise against trying to use portable libraries and instead using libraries designed for the platform you are targeting.
Having said that, a good portable UI library is the open source WxWidgets which is accessible through C++ for OSX, Linux, Windows
If you want to distribute dynamic-link binaries for windows, use MSVC.
If you want to distribute dynamic-link binaries for OS X, use Xcode.
If you want to distribute dynamic-link binaries for linux, you are SOL regardless of whether or not you are using C++, but if you use the same compiler and flags that the latest LTS version of Ubuntu uses, then it will work on Ubuntu, and will be made to work anywhere that Steam works.
It used to be that there were at least two C++ compilers for each *nix (typically GNU and something cfront based), so ABI was a much bigger deal.
When "Modern C++ Design" came out, famously none of the compilers could correctly compile all of the sample code. Since then things are much better; not that all compilers are bug-free of course, but they are sufficiently good enough that if you report a bug, you can expect it to be fixed.
[EDIT]
"Don't use STL Types in your interfaces" is not advice I've heard in like 15 years; I more often hear "If you're using a C array instead of a Vector, you're doing it wrong"
"Don't throw exceptions across module boundaries" seems similarly odd. Unless your constructors are inlined, no modern code-base will follow that rule because RAII relies so strongly on exceptions.
There are coding styles that are opposed to exceptions as part of an external interface, but that's due to exceptions not being checked as part of the type system, and is not what I would call a majority opinion.
To clarify "module boundaries", I mean "separate shared objects."
As for Linux, I'm not too concerned with creating a single binary that works for all distributions.
I'm more concerned with someone being able to build a set of shared libraries on their distribution of choice and those shared libraries being able to interact naturally regardless of which compiler s/he uses to build each of them.
Say, LibA is built using LLVM. LibB is built using G++ and LibC is built using ICC.
LibA defines several classes. LibB creates some subtypes. LibC instantiates types from both LibA and LibB.
All the functions present in LibA, LibB, LibC make use of STL types such as std::string, std::vector, etc. Some may throw exceptions, whatever.
With respect to MSVC, I've read that compatibility between Debug and Release builds is kind of suspect, especially if you're using STL types. Not to mention differences in MSVC version. Is this still a concern?
Sorry, but this is an unreasonable standard. Literally no language, including C supports this. With C it only works inasmuch as the C compiler authors work really hard to make it works, and even then it sometimes breaks (if your compiler inlines a call to malloc, and you free a pointer compiled with a different C Compiler that inlined a different malloc implementation, it can break horribly. Yes I've seen this happen.)
Some languages support cross-version linking (or whatever the language's equivalent of "linking" is), but I'm not aware of any that specify a complete ABI for unrelated implementations to support. IPC libraries do typically support this though.
[edit]
I don't want to go on a shared-library rant, but I am fairly strongly opposed to them (except perhaps in cases like how nixos manages it). You can take a statically linked binary from 1997 and run it unmodified on your linux machine today. It is a virtual guarantee that any dynamically-linked binary more than 2 years old will not work correctly. Linus puts a huge amount of effort into backwards compatibility, and it is completely destroyed by dynamic linking.
And yet... microsoft releases a new compiler every two years or so, and not every library you use is going to update at the same time. This is a huge frustration for a lot of people.
I write c++ professionally and I've seen people waste weeks on these things, and most the libraries we wrote had plain c interfaces because being able to use other languages to call into the code was important and c++ is a nightmare with that.
VS 6.0 is getting very hard to source legally these days, and I wish MS made it easier to get.
As far as having high-level languages call directly into C++, yes that's quite a pain (nearly impossible without something like https://github.com/rpav/c2ffi). Note also that calling into non-C ABI functions in any language is hard (and most HLLs don't support anything like extern "C" to make it easy).
2) Header only libraries are horrible for compile times, especially heavily templated ones (and if you use c++ generics it basically has to be a header library). The reason boost is banned from a lot of cpp projects isn't because the library is bad, it's because of compile time.
The organization of functions into objects makes it a lot easier to understand systems software specially if it's very large. Also the ability to subclass means you can take the base functionality of "template" classes provided by a library and subclass them to extend them with what you need. This was not as easy with C where you had to rely on sample code for this purpose. The Petzold book had a ton of sample code.
The C++ version uses many memory allocations. Using allocators some in the C++ program would certainly cut down on the number of allocations. It would also be interesting to see if doing so also improved performance.
Similarly, it would be interesting to see if using the C++17 string_view (or the gsl version if C++17 isn't available to you) instead of `const string &` parameters affected performance.
Finally. I see that in most (all?) cases, objects are returned by value, not returned through reference parameters or pointers. It's interesting to see that that choice didn't compare poorly to a C implementation.
C has relatively simple header files, they usually contain structs, function signatures, and a bunch of macros. They're easy to parse and apply by comparison, plus don't tend to be as deeply nested.
If C++ ever adopts the Pascal-style "module" extensions that have been kicking around in various proposals compile times could shrink by several orders of magnitude.
I'm skeptical. Modules don't avoid the need for template instantiation.
The motivation for modules in C++ is similar to that of developing a Binary AST for Javascript, discussed on HN recently.
Turbo C++ was never as "turbo" as Turbo Pascal.
You can get Visual C++ with support for them.
clang has support, but the old style ones
Currently they are already a TS targeted for C++20.
Really? And I thought that this is why C and C++ headers are typically wrapped in #ifndef-#define-#endif block, so they only produce whitespace after preprocessing on second inclusion.
That is, including a, b, c is not necessarily the same as a, c, b or b, a, c. This is not true with proper modules, they're order invariant, and as such you can make a ton of optimiztions.
I wrote the same program in two environments once: C++ with stdlib and boost, and C++ purely using QT abstractions. The second one already compiled in 1/5th of the time. I guess that's because with QT the headers are small because most implementations are hidden behind pointers (PIMPL), while with boost you often pull in lots of code through headers and compile dozens of specializations of similar types to avoid indirection costs.
This is interesting. Are they using modern C++ and making use of moves and perfect forwarding? Or are they just throwing std::strings around and doing millions of copies (e.g. remember std::vector must support copyconstructable, so copy constructors & operator=) in the process? That would explain the allocations in C++ being higher perhaps, particularly if they're using the "wrong" containers. Why not sure unique_ptr or shared_ptr?
It is worth remembering that move constructors and assignment operators only get used in very specific places and you have to ensure that any constructors you write yourself are explicitly noexcept.
You can't compare performance If one program doesn't free memory, which obviously "saves" time. Valgrind can tell you where non-freed heap blocks have been allocated and a fix should not be complicated.
In theory, theory and practice are the same. In practice, they aren't.
There's been various attempts at pre-compiling the headers over the years, but the results have always been, for various reasons, less than perfect.
Well, need is a strong word. You might not need better correctness while still coming out way ahead by using C++.
http://www.geeksforgeeks.org/virtual-functions-and-runtime-p...
struct A { int x; };
struct B : A { int y; };
is the same as if you had written struct B { int x; int y; };
Public/private/protected inheritance and access control do not add overhead. It's literally only if you opt in by typing `virtual` do you get class hierarchy overhead.That's interesting. So does the compiler just put the functions in different parts of the vtable to remember the access control rules. There's no such thing as a free lunch and you're adding information here -- has to be stored somewhere.
No, it's not. Calling an ordinary class member function in C++ has exactly the same overhead as calling a function in C. Even virtual functions in C++ are not the same as putting function pointers in a C struct (they live in a separate data structure called the vtable).
> the protection mechanisms and the class heirarchy
All C++ protection mechanisms occur at compile time and have no runtime overhead. Non-virtual inheritance hierarchies have the same overhead as C struct composition (because under the covers the memory layout is the same).
Modern compilers make the same optimization for function pointers which are only ever set to one value, too.
If you're using multiple inheritance then an object can have multiple vtable pointers, but again which one you need to use is known at compile-time based on which class the virtual function you're calling is declared within, and the type of pointer you have.
Once you have the vtable you then have to locate the function pointer for the function you're calling. Again, this is usually a compile-time constant offset from the start of the vtable. This ceases to be true when you have 'virtual inheritance' (not to be confused with virtual functions), when another indirection to find this function pointer is required.
Here are some examples:
You'll notice that the get_square() function, which returns a member function pointer to the virtual square function, doesn't even return any memory addresses, just metadata and an offset
Except that the poster above you shown that the opposite was true. The exact same program was written multiple times with C structs and with C++ classes and,
"Except for the non-inline unbundled monostate in C++, every non-inline C++ implementation outperformed every non-inline C implementation."
...as long as they're not virtual functions, this is correct. Add a virtual table and this is less correct (but the optimizer may still make it correct if it can prove the types match).