Unaligned accesses in C/C++: what, why and solutions to do it properly
blog.quarkslab.com
blog.quarkslab.com
In particular, an 8 byte aligned uint8_t also can't be dereferenced via a uint64_t. It isn't UB because the alignment is wrong, it's UB by definition and also the alignment might be wrong.
This is a woeful state of affairs propagated by an ill founded and widespread confidence that the C and C++ standards probably aren't as hostile to type aliasing as they actually are.
Similarly, it is perfectly safe to cast a correctly aligned uint8_t* to uint64_t* and dereference it as long as you originally stored an uint64_t into it (or the compiler can't prove you didn't). Remember aliasing UB is always on derefencing with the wrong dynamic type, not on casting pointers.
But today, uint8_t is in practice always a char type. You can static_assert if you want to be completely sure. Or just use std::byte.
You can cast any pointer you like to any other pointer. When you dereference it, if the declared type is not the type of the underlying object, bad times for you.
An interface which takes a uint8_t* and immediately invokes UB if you pass it anything other than a pointer to a uint64_t can indeed be written, but one should expect people to pass it things other than a cast uint64_t*.
Further, the magic aliasing char doesn't work in both directions. If this code was changed to take a char instead of a uint8_t it would be exactly as undefined as it is currently. You can aim a char* into a uint64_t and deref the char, but you can't aim a uint64_t* into a char and deref the uint64_t.*
Edit: That's what I remember but I can't find evidence of it. The docs say this is ABI dependent so perhaps the above statement is true on some specific ABI (if I haven't misremembered it entirely). But I can't track down wheere a relevant ABI would be documented.
https://gcc.gnu.org/onlinedocs/gcc-13.2.0/gcc/Architecture-i...
Edit 2: this header has a typedef of uint8_t to unsigned char, but it's part of glibc rather than gcc
https://android.googlesource.com/platform/prebuilts/gcc/linu...
uint64_t x=0;
uint8_t * p = (uint8_t*)&x;
uint64_t y = *(uint64_t*)p;
It is not the casting that is the problem, it is the dereferencing. And as long as that matches the dynamic type, it is ok.But C and C++ are explicitly intended for this kind of situation. You know enough about how types are represented in memory to fiddle with their internals, and you have sufficient reason to do so? Go for it. Just don't expect the type checking to save you, because it won't.
However that's not the setup. C++ in particular, and increasingly C copying from it, optimises assuming you never do anything weird with pointers. This makes application code that doesn't do such things faster in a separate compilation world. It makes system code that does do such things undefined, forcing either compiler flags (fno-strict-aliasing and friends, aka writing in a different language that looks like the nominal one) or taking care around separate compilation to make sure the compiler never sees both sides of the boundary at the same time.
There exist open source C compilers that aren't GCC or Clang, such as TCC and Zig, but to my knowledge they all produce really, really bad assembly. There existed proprietary C compilers in the 1980s and 1990s that were written by one or two people that still produced solid assembly, so it's possible, just no one has done it, so Clang and GCC are the only free and open source options for decent assembly for C. Those two also have the downside of being C++ compilers, making them more complex than they need to be.
Maybe another option would be to improve an already existing open source C compiler's codegen than to write a new one from scratch again, assuming they'd agree with the above ideals.
I like C as a notation for programming computers. I don't like the slight misstep -> UB -> everything is ruined philosophy of WG14.
(edit: it should be possible to emit LLVM IR with the right metadata to avoid the optimiser mangling it. In practice there would be a long tail of accidentally assuming C++ semantics that aren't justified by the IR, where bug fixes for those would be valid upstream. If one wanted to take on the probably thankless task of trying to remove ISO C/C++ assumptions from llvm.)
You can mutate a vtable without UB. A method may call the destructor on its this pointer and use placement-new to create a new object in place. The fact that any method of a class may do this combined with the fact that the placement-new object might be a more derived version of the destroyed object (so an existing Foo* continues to be valid) means that compiler can't cache vtable lookups for consecutive method calls, making nearly any optimization of virtual function calls impossible because you don't know the type or called function, unless you see the object being constructed (when the vptr is assigned) and inline each called function in turn.
(At some point C++ added a rule that basically reads "you're allowed to cache the vptr, if the accesses were written using the same pointer variable name". This doesn't work well for optimizing compilers because they'll quickly fold two equal values into a single variable in their internal languages and lose track of whether the user wrote two distinct variable names or not.)
Yep. Honestly, while I get the historical reasons for strict aliasing, I wish it hadn't been created. Just make -fno-strict-aliasing the default and normalize the abundant use of restrict, maybe shorten it to res or rest just like we use const instead of typing out constant. These implicit rules that usually, but not always, do what you want can be a bigger footgun than having to do things explicitly.
Those aren't constants†, they're immutable variables which are quite different. Because of the "as if" rule and provenance in some cases you get the same benefit as with a real constant, but that's not what this is.
† In C or C++. In some other languages const does, logically, get you an actual constant.
In a language where the const keyword gets you an actual constant this condition wouldn't arise.
Strict aliasing, IOW, is as a practical matter mostly about when and to what extent compilers (and sometimes processors) can optimize code through automatic reordering, much like wrt parallelism and memory barriers. In this sense, the role of strict aliasing is more important than it ever has been, and while technically "merely" a matter of performance (correctness could be obtained by completely abstaining from reordering), the performance of modern processors plummets precipitously when implicit hardware parallelism can't be fully leveraged.
Also, I'm not sure how true your argument is in practice. Modern compilers have other, more sophisticated ways to prove that things don't alias. A number of projects compile with -fno-strict-aliasing without using restrict very often, for example the Linux kernel, and they don't seem to suffer much from it. Linus has this to say about aliasing in kernel code:
> In x86, I doubt _any_ amount of alias analysis makes a huge difference (as long as the compiler at least doesn't think that local variable spills can alias with anything else). Not enough registers, and generally pretty aggressively OoO (with alias analysis in hardware) makes for a much less sensitive platform.
class Integer {
// uses unsigned integers for all operations
// but converts to signed when you need output
};
class Pointer {
// uses void* and memcpy for all loads or stores
};
Put all of the correct operator overloading on those to get the ergonomics we had in the 90s.Then we can safely ignore all the undefined behavior which has been so eagerly embraced by the compiler writers, and we can go back to treating the hardware as though it is sane. (Which the hardware is desperately trying to fake being anyways)
After that, we can create some sane container types which don't offer up dangling references or pointers, and Bob's your uncle.
But more importantly, I don't ever want to worry about stumbling into undefined behavior, and compiler extensions like -ftrapv or -fwrapv are not part of the standard. I can't count on those being there when I switch compilers, and if I write a library for others, I can't count on the application developer using the flags I need.
If you try casting your pointer to (__m128i *) and dereferencing, sometimes the compiler will optimize it correctly (to PMOVSXBD with a memory operand, which really only loads 4 bytes), and sometimes (for example, when optimizations are disabled) it will emit MOVDQA, which does a 16-byte aligned load. Tools like UBSan will also (correctly) complain about it.
Eventually __mm_loadu_epi32() was added, but it was released in a broken state in the initial gcc implementation: https://gcc.gnu.org/bugzilla/show_bug.cgi?id=99754 so you can't rely on it. There is __mm_load_ss(), but gcc's implementation derefs a (float *), so it still requires 4-byte alignment: https://gcc.gnu.org/bugzilla/show_bug.cgi?id=84508
The workarounds described in the article (combined with _mm_cvtsi32_si128()) are the only reliable solution I have found. Various compilers still generate an extra MOVD instruction instead of using a memory operand, and reportedly ICC will even do an extra load into a general-purpose register before moving the value to a SIMD register: https://stackoverflow.com/questions/72837929/mm-loadu-si32-n...
Those extra moves are cheap on x86, so it is not the end of the world, but the whole situation is less than ideal.
I'm a little surprised about the icc thing, since it has/had a reputation for really good Intel x86 codegen.
It's also fun how intrinsics are mixed between ISA extension levels with confusingly similar names that make it really easy to use them in the wrong code path. Arithmetic shift right immediate for int16 (_mm_srai_epi16) and int32 (_mm_srai_epi32) are SSE2. But _mm_srai_epi64 for int64 is AVX512.
One thing that Intel does do better is that they have a standardized, OS-independent way of testing for ISA extensions through CPUID. ARM doesn't and you are at the mercy of the OS to provide you APIs to test whether, for instance, Crypto, CRC32, CAS, and UDOT instructions are available.
#define unaligned(P) (((struct { typeof(*(P)) x; } __attribute__((packed))*)(P))->x)
[0] https://github.com/torvalds/linux/blob/7b9e664beb237d90bc600... uint64_t load64(uint8_t const *b) {
union {
uint64_t u64;
uint8_t u8[8];
} u = {.u8 = {
b[0], b[1], b[2], b[3], b[4], b[5], b[6], b[7]
}};
return u.u64;
}
Type punning through a union is explicitly supported by the C standard, insofar as we do not evaluate uninitialized memory or trap representations. We are fine here, as uint64_t has no trap representations, and must have exactly the size of 8 uint8_ts, as those types have no padding bits.GCC appears to compile this function to a plain unaligned load at optimization level 2, when supported, e.g. on arm64 (see https://godbolt.org/z/frEox67sG).
(And you computer can convert endianness in a cycle or two. There is very little excuse for optimizing based on “native endian”.)
edit: about the only good thing I can say about “native endian” is that it maybe justified in cases where the value never leaves the machine. Hash tables in memory come to mind, and the article seems to be about that.
(I'm not that familiar with this problem so I figured I would take this opportunity on a seemingly related problem to get more perspective.)
Edit: And if the buffer to convert is very large, one could first check the alignment of the pointer (with alignof(T)) in C++
See: https://github.com/scanmem/scanmem/blob/c6045a8677f37a51b976...
Edit: I mean on x86, where it's possible. I couldn't even find any intrinsics for the corresponding lock instructions. And yes I'm aware of the performance penalty.
Usually the answer is to ignore the atomic abstractions of C++ and use the GCC style atomic intrinsics (they take a void* instead of some atomic qualified thing) but I'm not totally confident they'll do the right thing if the target memory has less than natural alignment.
Beware crossing cache lines with your operation. I'm not sure what the x64 instructions do in such a case.
I'm saying I don't see any such intrinsics for this.
I do believe x86 and x64 LOCK works correctly across cache lines.
I can't tell if this is a joke or not, haha.