In practice, everything you're likely to come across will be little endian nowadays, and the ABI you're using will most likely order your struct from top to bottom in memory, so they will look the same most of the time. However, it's still technically not portable.
The internet?
struct X { type alias : numeric_value = false; ... }
I love the power of C++ but there is _so much_ to the language. I'm sure there would also be some template meta programming solution even if this syntax was available.
Sign-extension of 1-bit fields also messes people up all the time, but that’s “just” an easy-to-fix bug.
1. Addressing happens at the byte level, not the bit level, so a type can't begin on any bit. You'd have to do your own addressing, such as in std::bitset.
2. For now, there's no reflection in the language, so you can't really assign names to members in a general way (hello, preprocessor). A solution might be to index by type; something like the following:
struct BitA{}; struct BitB{};
using ExampleBitFields = BitFields<BitA, BitB>;
bool checkBitA(const ExampleBitFields& x) { return x[BitA{}]; }
However, this has a lot of downsides.
It is architecture/compiler dependent though. This is explicitly acknowledged in the rationale document:
“Since some existing implementations, in the interest of enhanced access time, leave internal holes larger than absolutely necessary, it is not clear that a portable deterministic method can be given for traversing a structure field by field.”
For example:
struct X {
int defaulted = 1; // 1 is the default member initializer
};Ah, that would make more sense. The combination of bitfield and default initializer syntax does look especially odd, since both features are rarely used IME.
>I couldn't see any reference to anything similar in what you linked
Indeed, there is no way to specify default values for a C struct at definition time.
It is c++20 apparently https://godbolt.org/z/qvso544dr
<source>(6): error C7582: 't': default member initializers for bit-fields requires at least '/std:c++20'
Nice to see they also improved such "legacy" stuffSay you have, for instance (using C notation)
struct {
unsigned one : 8;
unsigned two : 8;
};
The fields are supposed to be represented in memory in the same order they are declared, so one is the first byte and two is the second byte. This should have the same representation as if I declared it as two uint8_t fields. If I type pun it and load it into a register as a uint16_t then it depends on the hardware whether the low and high bytes are one and two or two and one.It gets more tricky when you consider arbitrary bit widths.
struct {
unsigned u4 : 4;
unsigned uC : 12;
};
If the fields are allocated in order, is u4 the low bits of the first byte or the high bits? If you require it to be the low bits, then it works ok on a little endian machine, but on a big endian machine the uC field ends up split, so the 16 bit view looks like: CCCC4444CCCCCCCCI'm not sure I accept "consistency with a bytewise view of memory" as a well-defined, reasonable concept. I do expect to give a list of bit widths, and get a field that has these in consecutive order. Why would it randomly do weird things on an 8-bit boundary?
> If the fields are allocated in order, is u4 the low bits of the first byte or the high bits?
It's the low bits on LE, and the high bits on BE.
> If you require it to be the low bits, then it works ok on a little endian machine, but on a big endian machine the uC field ends up split
That's why the direction is defined to match the endianness; you get a consecutive chain of bits in either case.
> so the 16 bit view looks like: CCCC4444CCCCCCCC
It's 4444CCCCCCCCCCCC on BE, and CCCCCCCCCCCC4444 on LE. If you need something else, it's no longer a question of defining an ABI-consistent structure, but rather expressing a representation of an externally given constraint.
> It's the low bits on LE, and the high bits on BE.
You are advocating for the current rule in C. As I said, C's rule implies 8 bit fields will be in different orders in memory on machines of different endianness, which makes it very difficult to use bitfields to get exact control over memory layout in a portable manner.
Your binary isn't going to be portable. The only reason your memory layout should be is if you intend to serialize it. But if you're going that extra step, you _need_ to convert it to a platform-independent format regardless -- otherwise, not even your ints deserialize correctly.
This entire topic disaggregates into 2 distinct categories: deterministic packing for architecture ABIs, which needs to be consistent but can be arbitrary. And representing externally defined structures, which is a matter of exact representation capabilities.
Unless you define a single 'right' bit order and then swizzle/unswizzle every value being written to or read from a packed struct, but then that's becoming more of a serialize/unserialize which is a different thing.
MSVC does things simply, gcc does not
#pragma pack(1)
struct S {
char c:1;
char d:1 __attribute__((aligned(2)));
char e:1;
};
_Static_assert(sizeof(struct S) == 1, "wrong size");E.g. https://www.nntp.perl.org/group/perl.perl5.porters/2008/02/m... for tricks.
Wrong. N1256 (ISO C99 spec), for instance, explicitly states in 6.7.2.1.10:
The order of allocation of bit-fields within a unit
(high-order to low-order or low-order to high-order)
is implementation-defined.
Look it up.The word "non-portable" is being thrown here too cheaply. By the definition used here, technically every C program that puts two integers consecutively on a struct is non portable, either. Not just due to alignment but due to endianess, etc.
endianness: how an array of bytes is interpreted as an integer/how an integer is layed out in memory as an array of bytes.
bitfield allocations: how subsequent bitfields are allocated within an integer, typically starting from least significant bit to most significant, or the other way around.
In this particular case, the two are related, because the LSBit-first bitfield allocations can spill over between bytes, giving LSByte-first endianness as well.
Similarly bitfields mean you end up with "ints" that are actually 3 bits wide and so on.
While the layout is indeed implementation dependent, pragmatically if you stick to using ints the layouts are portable as far as I can tell. Just like the size of ints is implementation dependent, but is reliably 32 bits on 32 and 64 bit machines.
One major difference appears to be that C bitfields memory layout is compiler-dependant. The other major difference is Zig's arbitrary-bit-width integer types just leading to less footguns I would speculate
This is news to me, how? AFAIK they'll only eventually make it into C23 with _BitInt(N).
...and apart from building a C++ wrapper class of course, but how would this pack with data outside the class - like the 4 + 28 bits example in the blog post.
Also IIRC when I tinkered with C/C++ bitfields, some compilers (at least MSVC I think?) didn't properly pack the bits (e.g. a single bit would be padded to a full byte, or they couldn't agree on a common size of the containing integer - e.g. one compiler packing <8 number of bits into an uint8_t, and another into a uint32_t). In the end C/C++ bitfields weren't all that useful for the use case described in the blog post, at least if portability across the three big compilers is needed (gcc, clang, msvc).
You probably want static asserts for sizes in your code if you are trying to optimize your struct paddings
[0] section "Notes" in https://en.cppreference.com/w/cpp/language/bit_field
Some C decisions really confuses me. What would be the point of this one?
There's #pragma pack for that.
BitInt looks neat. It sounds kind of like a bit array which we use quite a lot (which is as you suggest a class wrapper over an array of uint32_t templated on a size).
We use bitfields quite a bit across clang and MSVC targeting mobile, PC, and console, and haven't had any problems as far as I know.
My experience is C compilers have ways packing and defining the order of bitfields and structs.
What is this awful advice. Only convert to big-endian where legacy demands it.
Endianness is really perfectly named: a meaningless difference that generations of people fight holy wars over.
It was a flip of a coin choice.
The only thing I don't usually run into is the very topic of this article: bit fields.
The compilers often barely document exactly how they lay things out.
In the embedded world you often have to deal with vendor specific toolchains, and with their finicky compilers, you'd be surprised the weirdness you run into when using bitfields.
You will learn to question everything not specifically defined by the c standard.
Just one thing that Zig improves: in C if you need the bit offsets of the fields you'll still need lots of defines with the offsets/masks and such. In Zig it can be extracted from the packed struct at comptime