... but that's what a bit field is?
... but that's what a bit field is?
A bit field (in C/C++) is a weird object type that can only exist in a structure or union type, which kind of acts like an underlying regular integral type except for those situations where it does not.
For an example of why compilers might have issues compiling bit fields properly (although this requires C++, since C's ternary operator works on rvalues, not lvalues):
struct A { int x: 3; int y: 5 } a;
(choice ? a.x : a.y) = val;
Enjoy making that codegen work properly. choice ? ((a.x = val),a.x) : ((a.y = val),a.y);
The trick with a compiler is to rewrite complex constructions into simpler equivalent ones, then the code gen is much simpler and more reliable.For example, in the D compiler the `while` loop doesn't survive the semantic pass, as it gets rewritten into a `for` loop. The `for` loop then gets rewritten into `if` and `goto` statements. The code generator only needs to learn about `if` and `goto`.
You throw away structured control flow in favour of unstructured control flow?
I would have thought it'd be the other way around and you'd be trying to recover nice clean structured control flow from raw concepts like goto.
So, in short, structured control flow--or at least a sufficient subset of such structured control flow--can be easily recovered from low-level information, and you're generally not going to lose much information going down to that level. LLVM even has a way to attach metadata to loops despite not having any dedicated loop construct.
[1] You only need two of these concepts: the third falls out from the definition of the other two.
That's right. It sounds counter-intuitive, but it works great. You wind up with a collection of blocks of code connected by edges. Then, you can use graph theory to work magic on them in a general, correct way. Data flow analysis is based on this.
One thing the graph math does is enable the reconstruction of loops out of the blocks and edges - so you can write loops any way you please, and the compiler will figure it all out and apply general algorithms to it (like loop rotation, loop unrolling, etc.).
But hey, if Graal works for you, great!
The compiler's job is to emit instructions that do precisely the things the code says to do, as efficiently as possible. There is no need for chips to understand the purpose of the code; they just need to do what it says. They don't get confused.
Which requires a high-level understanding of the program and its control flow.
The `for` loop construct does not offer the compiler any more understanding of the code than one made from `goto`s. They contain the same information, and are interchangeable. The latter, however, is more amenable to applying mathematical algorithms to. The former is more amenable to human understanding.
With structured control-flow I can do things like reason about the level of nesting of the loop that I'm in. I can peel a loop iteration by literally saying 'take this loop here - copy the body of it out'.
For context here's the kind of structure we use https://chrisseaton.com/truffleruby/basic-graal-graphs/#loop....
The first is lvalues. In compiler jargon, an lvalue is a kind of object that can have a value stored to it. And you can usually represent it as the address of some memory location [1]. Of course, bitfields break this representation: you need to know what the bit offset and bit size of the field you're storing is (as well as the signedness).
The next level of complexity is the conditional operator. This means that, when conditional operators yield lvalues [2], you now end up in a situation where the lvalue now has a conditional bit offset and bit size within the address. Or maybe one leg of the expression returns a bit-field and the other leg returns a regular int lvalue. Imagine how complex your datastructure needs to be to represent an lvalue during this code generation phase.
[1] Not all lvalues need to have memory locations. But if you're writing a C compiler, it's an easy first approximation to give every variable, even those marked register, some memory location and rely on an optimization pass to convert stack memory locations into register locations, rather than keeping track of this information when the frontend does code generation.
[2] As mentioned elsewhere, conditional operators in C do not yield lvalues. But conditional operators in C++ do.
If you must use bit fields, make them unsigned. Bugs love to hide under signed bit fields.
Unsigned bitfields are a nice way to get modular arithmetic with n bits without syntactic clutter.
Appear to be. Are, when all the stars align. Are not in fact, often enough that you are issued a red warning you may ignore if you are insulated from all consequences.
It also doesn’t seem like something that would come up very often. I can’t think of the last time I conditionally stored to one of two struct fields, if I ever have.
The much more normal case would be:
val = choice ? a.x : a.y;
That one seems pretty straightforward from a codegen perspective.The broader point is that bitfields are actually weird little objects that look a lot like regular objects in many, but not all, contexts. And it's very easy from a language design or implementation perspective to forget to account for the possibility that you're dealing with a weird little object. This leads to underspecified language specifications and compilers that crash if you do something weird (but legal) such as virtually inherit from a struct containing a bitfield as its last member.
[1] So challenging, in fact, that Clang gives an error message "cannot compile this conditional operator yet". It does work in g++, icx, and MSVC though.
Of course, if you go reach for C's standard "fun with lvalue" operations, you can get some crazy nonsense. What machine code should you generate here [1]:
struct A { int x : 5; volatile _Atomic int y: 3; } a;
a.y++;
I will note that the intersection of volatile and bitfields has been another fruitful area of compiler bugs [2] historically speaking. While C++ does provide better what-the-ever-living-fuck moments for bitfields, C has had its fair share of issues with bitfields.[1] Whether or not you can make a bitfield _Atomic in C is implementation-defined, so it's possible that someone writes a C implementation where this is legal. I will note that, in a rare display of sanity, all C compilers I can test do in fact sensibly reject _Atomic bitfields, but for the purposes of argument, assume that someone has one where it's permitted, since it is allowable by the standard.
[2] Or programmer bugs blamed on the compiler. This is the intersection of two areas that are notorious for underspecification to begin with, and combined with the general tendency of programmers to expect C compilers to be a thin veneer over assembly, makes it awfully difficult to figure out which behavior is language-intended.
The underlying problem has to do with whether the IR has first-class concept of arbitrary lvalue or whether the frontend has to convert lvalues that get passed around to some pointer-like thing.
It might look irrelevant for discussion of low-level AOT compilers, but it is also interesting to compare how this is implemented in dynamic/“scripting” runtimes and how the choice of underlying implementation of the concept of “lvalue”/“place” influences the user visible language. Somewhat notably first draft of Common Lisp had something akin to first-class lvalues and the final standard replaced all that with significantly simpler mechanism that purely relies on macros.
He is describing a trivial difference between C and C++ that is not the problem you are being warned about.
struct foo {
char a : 4;
char b : 4;
};
Is a in the high-order 4 bits, or the lower 4 bits? Both choices are allowed, so it's up to the compiler and makes the code non-portable.64-bit Linux distros and the BSDs follow the convention once set by the "C ABI for Itanium".
In that, bitfields are grouped in declaration order into container words of the same width as the bitfield's type (char, int, etc.). Bitfields don't span multiple container words, and container words don't overlap. On little-endian platforms, bitfields are packed LSB first, but on big-endian platforms they are packed MSB first within their container word. Alignment rules apply only to the container words.
If the instructions emitted and the instructions implemented both happen to match that, on every chip your code must run on, you got lucky.
If you want to produce same sequence of bytes regardless of underlying platform, then you have to do it by hand with uint8_t[] buffers and explicit shifts and masks. Casting pointer to struct to char* and writing it somewhere is inherently non-portable and this gas nothing to do with bitfields and nothing to do with things like __attributte__((packed)), although both of these things are useful when you want to do that and understand the (non-)portability implications.
You know where the bits are within a single word. But if you have a struct with multiple fields, it’s not safe to rely on the exact memory layout even if it doesn’t have any bitfields.
If you need to represent a very specific memory layout, it’s not just bitfields you need to avoid, it’s structs in general.
Conversely, if you don’t need to guarantee a specific layout, bitfields are fine to use, and could be a useful optimisation hint for the compiler.
Say I have a window manager, and I want to attach a bunch of boolean flags to each window object (isVisible, isMaximized, etc). I don’t need to serialize them to disk. It’s highly preferable that they should be efficiently bit-packed, but not strictly essential.
The conservative way to implement that would be bit-shifts and masking (either manually or via a macro). But implementing it with bitfields would be a lot easier and less error-prone, and would work just as well. What problems do you see with the bitfield approach?
If it works on your particular compiler release, on your particular CPU chip stepping, that tells you nothing about the next compiler over and the next chip over.
amd64 and arm64, compiled with gcc or clang, you are unlikely to run into these problems. But code tends to get around.
If so, I think that’s overly paranoid. The examples that are being given here are baroque usage that would immediately stand out in a code review - memory-mapped registers, conditional lvalues, volatile and atomic fields.
The point I wanted to make is that simple straightforward usage of bitfields, like the example I gave, works fine on any platform you’re likely to encounter.
There’s plenty of widely-used code out there that uses bitfields. I just did a code search to check that (the particular example I was thinking of comes from iOS) and found some in Clang - funnily enough, in its representation of lvalues!
x = foo.a
is simpler than x = (foo & FOO_MASK_A) >> FOO_SHIFT_A
and for assignments, the difference is even bigger: foo.a = x
is much better than foo = (foo &~ FOO_MASK_A) | ((a << FOO_SHIFT_A) & FOO_MASK_A)The more frequent perceived use for bit-fields (in the situation where they actually work) is to pack into a serialized data format, such that memory or a data stream can be accessed elsewhere. In that case, "the compiler can do whatever it wants with your data packing" is pretty useless, since your "elsewhere" might have a different compiler that does a totally different thing.
And as for the second part: anything that writes sizeof(struct foo) bytes of struct foo is inherently non-portable. If you portably want to (de)serialize something you want to write the thing explicitly, very often the compiler will optimize it to more direct implementation. (And well, this is only portable to platforms where CHAR_BITS == 8)
Anything that affects the actual instructions executed on the actual chip they're executed on may make what works here not work there.
Optimization that does not affect instructions is no optimization at all. Bitfields are an extremely fragile part of implementations. Trust it at your own risk.
Ladies and gentlemen, this thought is why we now consider 8GB of ram to be a "weak device".
No, no no no no, 1000 times no. Every situation is a low ram situation. Every!
Edit: I know it's hard to read a whole sentence at once, but I made that same point directly up there too.
Hopes, prayers, and a single version of a single compiler being involved.
A result of people avoiding declaring bit fields in serious use cases has been that compiler vendors didn't worry too much about bitfield codegen bugs.
Probably Gcc and Clang are OK on x86, by now. But that does not carry to, e.g., obscure microcontrollers. Heaven help you if your bit field members are supposed to correspond to hardware register sub-fields.
void bar(struct y *s, unsigned int foo) {
s->c = (s->c & 0xf0) | foo;
}Not always. Switch your example to AARCH64 and check out the BFI instruction.