Implementing a class with void*
web.eecs.utk.edu
web.eecs.utk.edu
class C {
protected:
struct Impl;
std::unique_ptr<Impl> mImpl;
};
And then define C::Impl in the .cpp file. // Forward declaration of Impl. What does Impl do?
// You're not allowed to know.
class Impl;
class A {
protected:
Impl* mImpl;
};
I've been using this trick but for a different reason - to reduce the number of #include statements in header files that are included a lot themselves.These are the kinds of refactorings you often need to justify with hours of arguments and approvals and religious fights and code reviews at work, but can do in an afternoon in a private project just because it makes the code nicer.
These days I wish there was something similarly easy to speed up Webpack builds ...
A much better alternative than precompiled headers.
Mostly, you just want to decouple the interface of the class from the implementation details.
In that case do just that - define an interface and implement it elsewhere.
You can always just static link all binaries.
Not always you can't. There are many reasons to use dynamically linked libraries all of which are applicable whether you're using PIMPL or not.
What's an example of this (assuming you don't care about binary size)?
I don't have the spec handy but fairly certain that v-table implementation is compiler specific and while it may work it isn't guaranteed. However if you declare the function symbols as dynamic then you can leverage the linker to dynamically resolve the right symbols with the matching opaque data and achieve binary compatibility(assuming you use the same compiler or a compatible compiler ABI and all other caveats around C++ binary compatibility).
But yes, I don't think anybody claims is a free lunch. It is annoying, verbose, repetitive, but often necessary to keep the cose base complexity under check.
I might be misremembering the "for virtual functions" part, but the inability of C to call Impl's constructor directly in the presence of a derived class can sometimes force you to delay some of the initialization until after Impl is constructed. ("Necessary" might be too strong here, in that you could find some other workaround too.)
> often necessary to keep the cose base complexity under check.
Hmm... I'm not sure I agree. A pimpl-like idiom can be necessary for solving a few very specific problems, which are explained in [1] better than I can here: (a) ABI stability, (b) slow compilation, and (c) exception safety. There are other niche cases I can think of (e.g. "I need fast/atomic swapping like a reference type, but copying like a value type"), and even those might have better solutions, but none of them really has anything to do with code complexity... they're either domain requirements you either have or don't have, like in (a)/(c)/(d), or they're workarounds for slow toolchains, as in (b). But unless you have requirements/constraints like these, I have a hard time recalling any common situations where pimpl would be the best solution, especially if it's for taming complexity.
i see you haven't heard of the occult powers of going to peek into the source file, copying the impl struct definition in your own source and going for a big bad reinterpret_cast
struct deleter_t { void operator()(impl_t *); };
using impl_unique_ptr_t = std::unique_ptr<impl_t, deleter_t>;
You can place the implementation of the deleter alongside the impl.Maybe? I can still think of ways to have ABI problems in the implementation of class C.
It's true that there are fewer ABI problems to worry about, though.
Yes which is why GP is speaking in terms of allowance. Using this pattern you can retain ABI compatibility, that doesn’t mean you do and it’s otherwise a free for all.
In practice you get a strong amount of compatibility between Clang and GCC at the compiler level (since Clang basically copies GCC's ABI). But you also get the same fudging at the stdlib level with std::unique_ptr because it's header-only and is a zero overhead abstraction for a raw pointer (i.e. at the ABI level a normal unique_ptr is the same as a pointer except you can't move it around in a register).
It absolutely is a de facto industry standard way to program C++, and has a name: PIMPL (Pointer to IMPLementation). It has that name, because it's famous.
It's probably less fashionable in newer code bases; probably someone whose head is up in C++20 will probably scoff at this, and certainly at any version where the secret is hidden by void *.
It provides a good way to wrap C API's in C++.
And, speaking of that, the technique is basically the spiritual equivalent of what happens in many a FFI module in languages other than C++ too, where some C handle is represented as an opaque foreign pointer, which is wrapped in some object native to the language.
As others have said, there's no reason whatsoever to use void* here. Just declare the struct without defining it in the header. Then only define it in the C file. C is fine with pointers to structs which are merely declared but not defined.
There is almost no reason whatsoever to use void * anywhere other than to write a declaration that is compatible with another one which uses it.
(A pointer to any object is better implemented as a typedef for unsigned char *. This requires a cast in both directions, thus it is safer. At the same time, it is more convenient when you actually want to work with the memory as such: you have bytewise arithmetic and dereferencing.)
> C is fine with pointers to structs which are merely declared but not defined.
Particularly if the secret object is a struct/class then this certainly improves the code (and when the secret isn't such a thing, it can probably be made into one: e.g. a secret array of integers can probably just be a a struct containing an array, possibly a flexible one). Numerous unsafe casts are thereby eliminated.
But it makes no material difference to its organization or semantics. It's still the same PIMPL pattern.
Well, what if the actual class is a template? No matter, you can derive a class from a template, as in:
header: class Impl;
implementation:
class Impl: public std::vector<SomeType> { ... };
People who don't do this and use void* wind up with an implementation that has lots of casts in it, and it's easier to introduce bugs, especially once you get to the point where you have two or more hidden implementations.
Declaring a private and opaque forward decl is very different to VOID* all the things.
Maybe it is meant to serve notice that the method is archaic, and that better ways to achieve the same thing will be presented later. But then one might just as well code the throwaway examples in C and use a damn C compiler. It's archaic even for C. There is just no merit in storing a void* that will then be cast to some arbitrary other known type every time it is used.
The only really valid use for a void pointer is when it is being passed through a subsystem and comes back to somebody who knows what it is, such as when an abstract handler function is being registered along with a pointer to some context that will be passed to the handler; this is extremely common in C code.
In C++, you can just take a pointer to a type T with virtual member f() you promise to call. The caller supplies a pointer to T2 derived from T and with its own f(). This, too, feels a bit archaic, but is at least not actively silly.
See `OSMetaClassDefineReservedUnused` in `OSObject.cpp`, here: https://github.com/apple/darwin-xnu/blob/main/libkern/c++/OS...
The disadvantage of the method proposed in the article is that, now everything needs a pointer indirection.
A solution that fixes the drawback is used by lz4's library implementation: instead of storing a void star, store a char[] array of same size as the real struct. (of course, now you have to manually make sure the struct size are in sync. It's more error prone, but still not that bad).
The implementation should probably provide you with std::aligned_storage which may involve some magic to handle some of those concerns, although merely the easiest ones, so I would say probably still don't use that either, unless you are already quite a C++ expert and/or are prepared to dig into the standard with no clear response about what you are attempting to do is even formally possible (the implementers/standardizers do not even know some things they make impossible for quite a long time, see for example the insanity of std::launder, or if you want to loose your mind forever the semantic of pointer provenance analysis that compilers are maybe already using to "optimize" but that what the semantic should even be is still being debated.)
I think they have a std::launder thing exactly for this purpose of "safely casting an array of bytes into an object".
However, in this particular case (of using char array to hide real implementation), the implementation resides in another translation unit, so I don't think anything is going to break if LTO is not enabled. With LTO I have no idea..
But then I force myself to find a second reason for why the program will run correctly, and unfortunately nowadays it is more and more being strictly-conforming. Relearning std::launder, TBAA, pointer provenance, etc. every time is way too consuming. I'm forced to give-up on programmer optimization and hope for the compiler to be really up to its mythical promises (and this yet: without LTO; too dangerous...)
I don't think in this specific case there is an UB involved, but I'm not language standard lawyer so I'm not sure. I feel the standard's specification on what is allowed to reinterpret_cast and what isn't is arcane (or at least far from straightforward to understand).
I've had to use this technique in the past and therefore dived into the standard for quite a while. I don't recall encountering any UB concerns.
But, yes, you do technically need std::launder to get it directly via the array.
Do you have a source for this? IIRC char and std::byte have a specific aliasing exception. I.e. char* and std::byte are allowed to alias anything.
You obviously aren't allowed to modify the char or std::byte array (because that would violate the struct/class's aliasing rule).
Additionally, implementers are in completely agreement here that this works. There are zero standardization/implementation concerns with this method, and I would highly advise against scaring users away from it when necessary.
(and that's exactly my point: the standard is hard to access for normal devs like me, so I can guess "some reinterpret_cast hack seems ok", but never be 100% certain until some standard expert confirms)
Can't you just use a static assert in the implementation file?
class Impl;
This isn't just for hiding, it also gives you faster compilation.
That seems unlikely. What are you basing this on?
class Foo;
to
#include "Foo.h"
and I would not consider using a void pointer for anything other than an allocator.
Edit: This is interesting from the case of "beginning of the calendar" - coincidentally just read this from a reddit thread (linked):
> Let me start with a quote you may hear in a lot of history classrooms: "Jesus Christ has been born 7 years Before Christ", sounds a bit weird doesn't it. Yep, historically the idea was to mark his birth as a dividing point, however, there are a bit of problem of determining precisely what year we're trying to set as first one. To keep the meaning intact, we would have to move dates if we find better information about exact year of his birth, not convenient at all.
> When we say that we are in year 2012 CE, then we don't care that year 1 CE should be some specific event (there isn't year 0 in common notation BTW), we just need all agree on the same starting point. We leave the date intact because it's widespread, so it's more convenient not to change the date, the same way we still use non-decimal hours, minutes, seconds. If f.e. Anno Mundi (from Creation of the World) system remained in use in Europe with the agreement on the same date (Latin and Greek scholars disagreed on Biblical age of Earth), we would use it as Common Era.
https://www.reddit.com/r/TrueAtheism/comments/14nuuw/i_think...
This is interesting to me from a wild out there idea that feels similar to 'coordinate space for storing data.' Time as a coordinate space for simulations in a way.