Packed structs in Zig make bit/flag sets trivial
devlog.hexops.com
devlog.hexops.com
Doesn't this give you if alpha OR blue is set?
Errors like this are another reason syntactical sugar for readability is important.
if((mask & WGPUColorWriteMask_Alpha) && (mask & WGPUColorWriteMask_Blue)) { //… }
This is the least confusing form I’ve seen that doesn’t require a function, macro, custom operator, etc. static bool all_bits_set(i32 value, i32 mask) { return (value & mask) == mask; }
Here I'm assuming 32-bit values. In C, with relatively little support for generics, you can consider making multiple versions, possibly using _Generic (note, I haven't evaluated the sanity of using _Generic).Alternatively you can use a #define. However, you need to use "mask" twice, so that gets tricky - either it requires care to keep the corresponding expression at the call site side-effect free. Or the macro needs to be written using compiler extensions like statement expressions and typeof() variable declarations à la Linux kernel.
#define all_bits_set(value, mask) ({__typeof__(value) v = (value), m = (mask); (v & m) == m; })
I personally would go for the extension aboveor a single 64 bit function, but here's the _Generic version: static bool all_bits_set32(i32 value, i32 mask) { return (value & mask) == mask; }
static bool all_bits_set64(i64 value, i64 mask) { return (value & mask) == mask; }
#define all_bits_set(value, mask) _Generic((value), int32_t: all_bits_set32, int64_t: all_bits_set64)(value, mask) if ((mask & (WGPUColorWriteMask_Alpha|WGPUColorWriteMask_Blue)) == (WGPUColorWriteMask_Alpha|WGPUColorWriteMask_Blue)) { ... }
...which is quite a mouthful. if (mask &== WGPUColorWriteMask_Alpha|WGPUColorWriteMask_Blue) { … }
I find only two very minor problems with it: firstly, that it’s not commutative (that is, a &== b is not equivalent to b &== a). Secondly, that it can be confused with the bitwise-and assignment operator &= (a &= b being equivalent to a = a & b for singly-evaluated lvalue a).I’d then add |== for consistency and because it is conceivably useful. ^== is tempting, but since a ^== b would be just another spelling of a == 0, I’d skip it.
But then, continuing that thinking, if you have '&==', you'd also need '&!=' -- that looks really confusing.
Can't we turn that into a '==0' test somehow? Maybe '^&', because '((a ^ b) & b) == 0' is equivalent to '(a & b) == b'. And, as somehow said: '~&' also works: '(~a & b) == 0' is also equivalent to '(a & b) == b'.
if ((a ~& b) == 0) { ... }
Wait -- we can swap that into '&~' and we're back in C. However, it reverses the logical order, with the test bits first: if ((b & ~a) == 0) { ... }
if (((WGPUColorWriteMask_Alpha|WGPUColorWriteMask_Blue) & ~mask) == 0) {
...
}
So we're back -- this is standard C now. But I find it incomprehensible. if (!(b & ~a)) { … }
if (!((WGPUColorWriteMask_Alpha|WGPUColorWriteMask_Blue) & ~mask)) { … }
if (!(~a & b)) { … }
if (!(~mask & (WGPUColorWriteMask_Alpha|WGPUColorWriteMask_Blue))) { … }
Actually, that latter makes a little more intuitive sense to me because of how ! and ~ are both negation¹, so they kinda cancel out and leave just the masking. Kinda.—⁂—
¹ Fun fact: Rust uses ! for both logical and bitwise negation, backed by core::ops::Not, with bool → bool and {integer} → {integer}, since bool is a proper type and there’s no boolean coercion anywhere—so you would have to stick with `== 0` in Rust, though in practice you’d probably go all typey with the bitflags crate’s macro to generate good types with some handy extra methods, and write `mask.contains(WgpuColorWriteMask::Alpha | WgpuColorWriteMask::Blue)`.
if (mask.contains(WGPUColorWriteMask_Alpha) || mask.contains(WGPUColorWriteMask_Blue)) { ... } // To check for either flag:
if mask.contains(WgpuColorWriteMask::Alpha) || mask.contains(WgpuColorWriteMask::Blue) { … }
if mask.intersects(WgpuColorWriteMask::Alpha | WgpuColorWriteMask::Blue) { … }
// To check for both flags:
if mask.contains(WgpuColorWriteMask::Alpha | WgpuColorWriteMask::Blue) { … }Why is that a problem? That is, if a has additional bits set compared to b, then ((a & b) == b) != ((b & a) == a), no?
For equality comparisons, there are two opposing conventions: actual == expected (the more popular, in my experience), and expected == actual (by no means rare). Having &== and |== be order-sensitive (though there’s absolutely no question in my mind about what the ordering should be) is mildly unfortunate.
It’s very minor.
It's a separate operator after all, one I also miss a lot...
if (!(~mask & (WGPUColorWriteMask_Alpha|WGPUColorWriteMask_Blue))) { ... }
Not quite as long but perhaps less readable. if (all_of(mask, WGPUColorWriteMask_Alpha | WGPUColorWriteMask_Blue))
with all_of() being a #define. Likewise none_of(), any_of().No need for special operators.
if (@popCount(mask & (WGPUColorWriteMask_Alpha|WGPUColorWriteMask_Blue)) == 2)
I do not know Zig, so the syntax might not be right. I did check to see that it has popcount [1].If it has some concise way to flip all the bits, then this would be another possibility that isn't too verbose, but might raise other objections. Let fmask be mask with all the bits flipped (how would one do that in Zig?).
if ((fmask & (WGPUColorWriteMask_Alpha|WGPUColorWriteMask_Blue)) == 0)
[1] https://ziglang.org/documentation/master/#popCount if (mask.alpha and mask.blue) { if (mask & WGPUColorWriteMask_Alpha & WGPUColorWriteMask_Blue) {
// alpha and blue are set..
}If WGPUColorWriteMask_Alpha and WGPUColorWriteMask_Blue doesn't share bits, isn't this garanteed 100% to be false?
// cannot convert i (variable of type int) to type bool
bool(i)
youd need to use a function: func to_bool(i int) bool { return i != 0 }https://developer.mozilla.org/docs/Glossary/Truthy
and I fully support that decision. If you want to use a boolean, you need to be explicit about it.
The compiler isn't going to make that function, it's going to optimize back to a cast to boolean. Why make the poor shmoe user type it out?
If you want to call it arbitrary, it's been arbitrated decades ago, but Boolean logic is much older than computers and works as it does for a reason. I suspect you know that.
People keep telling me this koolaid is delicious but I just don't see it.
edit: oh, i think i'm wrong, nevermind.
Correct usage would be if you want both flags.
(flag & (WGPUColorWriteMask_Alpha|WGPUColorWriteMask_Blue) == (WGPUColorWriteMask_Alpha|WGPUColorWriteMask_Blue) if WGPUColorWriteMask.Alpha in mask and WGPUColorWriteMask.Blue in mask: ...In practice, everything you're likely to come across will be little endian nowadays, and the ABI you're using will most likely order your struct from top to bottom in memory, so they will look the same most of the time. However, it's still technically not portable.
The internet?
struct X { type alias : numeric_value = false; ... }
I love the power of C++ but there is _so much_ to the language. I'm sure there would also be some template meta programming solution even if this syntax was available.
Sign-extension of 1-bit fields also messes people up all the time, but that’s “just” an easy-to-fix bug.
1. Addressing happens at the byte level, not the bit level, so a type can't begin on any bit. You'd have to do your own addressing, such as in std::bitset.
2. For now, there's no reflection in the language, so you can't really assign names to members in a general way (hello, preprocessor). A solution might be to index by type; something like the following:
struct BitA{}; struct BitB{};
using ExampleBitFields = BitFields<BitA, BitB>;
bool checkBitA(const ExampleBitFields& x) { return x[BitA{}]; }
However, this has a lot of downsides.
<source>(6): error C7582: 't': default member initializers for bit-fields requires at least '/std:c++20'
Nice to see they also improved such "legacy" stuffIt is c++20 apparently https://godbolt.org/z/qvso544dr
It is architecture/compiler dependent though. This is explicitly acknowledged in the rationale document:
“Since some existing implementations, in the interest of enhanced access time, leave internal holes larger than absolutely necessary, it is not clear that a portable deterministic method can be given for traversing a structure field by field.”
For example:
struct X {
int defaulted = 1; // 1 is the default member initializer
};Ah, that would make more sense. The combination of bitfield and default initializer syntax does look especially odd, since both features are rarely used IME.
>I couldn't see any reference to anything similar in what you linked
Indeed, there is no way to specify default values for a C struct at definition time.
Say you have, for instance (using C notation)
struct {
unsigned one : 8;
unsigned two : 8;
};
The fields are supposed to be represented in memory in the same order they are declared, so one is the first byte and two is the second byte. This should have the same representation as if I declared it as two uint8_t fields. If I type pun it and load it into a register as a uint16_t then it depends on the hardware whether the low and high bytes are one and two or two and one.It gets more tricky when you consider arbitrary bit widths.
struct {
unsigned u4 : 4;
unsigned uC : 12;
};
If the fields are allocated in order, is u4 the low bits of the first byte or the high bits? If you require it to be the low bits, then it works ok on a little endian machine, but on a big endian machine the uC field ends up split, so the 16 bit view looks like: CCCC4444CCCCCCCCI'm not sure I accept "consistency with a bytewise view of memory" as a well-defined, reasonable concept. I do expect to give a list of bit widths, and get a field that has these in consecutive order. Why would it randomly do weird things on an 8-bit boundary?
> If the fields are allocated in order, is u4 the low bits of the first byte or the high bits?
It's the low bits on LE, and the high bits on BE.
> If you require it to be the low bits, then it works ok on a little endian machine, but on a big endian machine the uC field ends up split
That's why the direction is defined to match the endianness; you get a consecutive chain of bits in either case.
> so the 16 bit view looks like: CCCC4444CCCCCCCC
It's 4444CCCCCCCCCCCC on BE, and CCCCCCCCCCCC4444 on LE. If you need something else, it's no longer a question of defining an ABI-consistent structure, but rather expressing a representation of an externally given constraint.
> It's the low bits on LE, and the high bits on BE.
You are advocating for the current rule in C. As I said, C's rule implies 8 bit fields will be in different orders in memory on machines of different endianness, which makes it very difficult to use bitfields to get exact control over memory layout in a portable manner.
This entire topic disaggregates into 2 distinct categories: deterministic packing for architecture ABIs, which needs to be consistent but can be arbitrary. And representing externally defined structures, which is a matter of exact representation capabilities.
Your binary isn't going to be portable. The only reason your memory layout should be is if you intend to serialize it. But if you're going that extra step, you _need_ to convert it to a platform-independent format regardless -- otherwise, not even your ints deserialize correctly.
Unless you define a single 'right' bit order and then swizzle/unswizzle every value being written to or read from a packed struct, but then that's becoming more of a serialize/unserialize which is a different thing.
E.g. https://www.nntp.perl.org/group/perl.perl5.porters/2008/02/m... for tricks.
Wrong. N1256 (ISO C99 spec), for instance, explicitly states in 6.7.2.1.10:
The order of allocation of bit-fields within a unit
(high-order to low-order or low-order to high-order)
is implementation-defined.
Look it up.The word "non-portable" is being thrown here too cheaply. By the definition used here, technically every C program that puts two integers consecutively on a struct is non portable, either. Not just due to alignment but due to endianess, etc.
#pragma pack(1)
struct S {
char c:1;
char d:1 __attribute__((aligned(2)));
char e:1;
};
_Static_assert(sizeof(struct S) == 1, "wrong size");MSVC does things simply, gcc does not
Similarly bitfields mean you end up with "ints" that are actually 3 bits wide and so on.
While the layout is indeed implementation dependent, pragmatically if you stick to using ints the layouts are portable as far as I can tell. Just like the size of ints is implementation dependent, but is reliably 32 bits on 32 and 64 bit machines.
endianness: how an array of bytes is interpreted as an integer/how an integer is layed out in memory as an array of bytes.
bitfield allocations: how subsequent bitfields are allocated within an integer, typically starting from least significant bit to most significant, or the other way around.
In this particular case, the two are related, because the LSBit-first bitfield allocations can spill over between bytes, giving LSByte-first endianness as well.
One major difference appears to be that C bitfields memory layout is compiler-dependant. The other major difference is Zig's arbitrary-bit-width integer types just leading to less footguns I would speculate
This is news to me, how? AFAIK they'll only eventually make it into C23 with _BitInt(N).
...and apart from building a C++ wrapper class of course, but how would this pack with data outside the class - like the 4 + 28 bits example in the blog post.
Also IIRC when I tinkered with C/C++ bitfields, some compilers (at least MSVC I think?) didn't properly pack the bits (e.g. a single bit would be padded to a full byte, or they couldn't agree on a common size of the containing integer - e.g. one compiler packing <8 number of bits into an uint8_t, and another into a uint32_t). In the end C/C++ bitfields weren't all that useful for the use case described in the blog post, at least if portability across the three big compilers is needed (gcc, clang, msvc).
BitInt looks neat. It sounds kind of like a bit array which we use quite a lot (which is as you suggest a class wrapper over an array of uint32_t templated on a size).
We use bitfields quite a bit across clang and MSVC targeting mobile, PC, and console, and haven't had any problems as far as I know.
You probably want static asserts for sizes in your code if you are trying to optimize your struct paddings
[0] section "Notes" in https://en.cppreference.com/w/cpp/language/bit_field
Some C decisions really confuses me. What would be the point of this one?
There's #pragma pack for that.
My experience is C compilers have ways packing and defining the order of bitfields and structs.
The only thing I don't usually run into is the very topic of this article: bit fields.
The compilers often barely document exactly how they lay things out.
In the embedded world you often have to deal with vendor specific toolchains, and with their finicky compilers, you'd be surprised the weirdness you run into when using bitfields.
You will learn to question everything not specifically defined by the c standard.
What is this awful advice. Only convert to big-endian where legacy demands it.
Endianness is really perfectly named: a meaningless difference that generations of people fight holy wars over.
It was a flip of a coin choice.
Just one thing that Zig improves: in C if you need the bit offsets of the fields you'll still need lots of defines with the offsets/masks and such. In Zig it can be extracted from the packed struct at comptime
bitfield ColorWriteMaskFlags : u32 {
red 1 bool,
green 1 bool,
blue 1 bool,
alpha 1 bool,
};
The number after the field name is the bit size. An (offset, size) pair can be used instead (offsets start at 0 when not explicit). After that can be nothing (the bit field is an unsigned number), or the word 'bool' (the bit field is a boolean) or the word 'signed' (the bit field is a signed number in two's complement). The raw value of the bitfield can always be accessed with 'foo.#raw'.EDIT: There's also no restriction to how multiple fields can overlap, as long as they all fit within the backing type.
Also what happens to the padding (bits not covered by subfields)? How’s that going to look when shoved into a file or over a socket? how does the langage handle overflow (more bits in the bit fields than there are in the parent field)?
So red is the LSB because you decided it was the LSB. That is not a by-definition thing.
> Endianness is not a concern because you have to specify the underlying integer type, and bit endianness is not a thing
That’s not actually true. There are formats which process bytes LSB to MSB, and formats which process them MSB to LSB. E.g. git’s offset-encoding, the leading byte is a bitmap of continuation bytes, bit 0 indicates whether byte 7 is present.
Both are perfectly justifiable, one is offset-based, while the other is visualisation-based as bytes are usually represented MSB-first (as binary numbers, essentially).
> but when reading a field only its specified bits are read so unspecified bits doesn't change that result. When writing to a field, the source value is truncated to the field size, so you never end up writing to other bits.
I’m quite confused by “bitfield” do you mean the container field (the one that’s actually defined by the `bitfield` keyword) or the contained sub-fields?
By 'field' I meant the contained sub-fields
Maybe Zig can handle this case with unions though, I haven't tried this yet (this would require that unions can work on the 'bit level' in packed structs).
One thing that is a bit noisy is that you have to specify the bit index when you name the individual bits, as in (from that article):
static let secondDay = ShippingOptions(rawValue: 1 << 1)
Upside from that is that it makes it clear what bit each value specifies, and won’t easily accidentally change them when you reorder definitions, or insert or remove them. I can see why they made that choice.I also guess OptionSet could have been implemented in Swift in a third party library, whereas this Zig feature cannot.
I also think/guess neither language guarantees multiple bits would get read or written in one instruction. If your hardware needs that, you probably have to go down a level, or look at the disassembly of your code to check what your compiler did.
Making a Zig BitSet would probably be doable, for those cases where bitfields are overkill.
https://www.scattered-thoughts.net/writing/mmio-in-zig/
It gets tricky if you need more control over loads and stores, or if there are different address spaces.
> this would require that unions can work on the 'bit level' in packed structs
Just as Zig has packed structs, it also has packed unions. So that part shouldn't be an issue.
const std = @import("std");
const expect = std.testing.expect;
test {
const group1: P = .{ .a = true, .b = true };
const group2: P = .{ .c = true };
const group3 = @bitCast(P, @bitCast(u4, group1) | @bitCast(u4, group2));
try expect(group3.a);
try expect(group3.b);
try expect(group3.c);
try expect(!group3.d);
}
const P = packed struct {
a: bool = false,
b: bool = false,
c: bool = false,
d: bool = false,
};These are not standard, so you need some preprocessor magic to choose the right thing. And so on...
We can also do better in those other languages, too. For example, in Rust, I can use a crate like `bitfield` which gives me a macro with which I can write
bitfield! {
pub struct Color(u32);
red, set_red: 0;
green, set_green: 1;
blue, set_blue: 2;
alpha, set_alpha: 3;
}
Don't get me wrong: it's cool that functionality like this is built-in in Zig, since having to rely on third-party functionality for something like this is not always what you want. But Zig is not, as this article implies, uniquely capable of expressing this kind of thing.But yeah, sometimes it's better to have options. If it's common functionality though, there will likely be 1000 different implementations of it that all just slightly differ [0]. Perhaps it were better for that effort to be put into making the standard better.
I don't think there's a universally correct answer by any means, but for something so common as bitflags, I think I personally lean towards having a standard. Replacing an implementation wholesale feels like it should be reserved as a last resort.
Either way, I think mature pieces of software (languages especially) strive to provide a good upgrade path. Inevitably, the designers made something that doesn't match current needs. Even if it's just that "current needs" changed around them.
[0]: And if we subscribe to Sturgeon's Law, 90% of those are crap, anyway. Though they might not appear so on the surface...
That's the luxury of a standard build system: essential but rarely used features can be left out of the core language / lib because adding them back in is just a crate import away.
If the language supports implementing a feature externally then it’s a good thing, as it allows getting wide experience with the feature without saddling the language with it, and if the semantics are fine and it’s in wide-spread use, then nothing precludes adding it to the language later on.
It’s much easier to add a feature to a language than to remove it.
Many years ago I wrote a rust program to decode some game save data and it looks like what you’d expect.
https://github.com/aconbere/monster-hunter/blob/master/src/o...
Syntactic sugar neve hurts.
Note that the main difference between packed structs and regular structs is not the dense bit packing, rather that regular structs are allowed to reorder the fields however they wish in memory. The compiler is free to optimize your struct. It's not free to do so in other languages (like C for example). Thus you get a way to define structs exactly, with bit level precision, to build complex protocols where you can decide what every bit means.