GCC always assumes aligned pointer accesses
trust-in-soft.com
trust-in-soft.com
Coming next: how modern programmers can't imagine a different word size or byte order. After that: different memory-ordering models. That one's really fun.
https://www.kernel.org/doc/Documentation/memory-barriers.txt
Amusingly, it treats dependency issues as obsolete, but I worked on a processor as late as 2009 that had all of the weak-memory-ordering issues of the Alpha (because it was the same design team). We uncovered lots of missing memory barriers in the Linux kernel. Good times.
In about the era you described, the Linux kernel community did go around throwing rmb()s and wmb()s everywhere, and those did translate to hardware fences. (And in general, it is not safe to elide a hardware fence on x86 despite relatively strong memory ordering.) This is what I believed you were talking about; maybe I misunderstood: apologies.
Putting explicit compiler barriers in code is still not quite the modern model for relaxed atomic consumers. You use the abstract acquire/release load/store pseudo-functions, and they do the right thing depending on implementation. The compiler barriers in the relaxed atomics model are an internal implementation detail, not the API.
x86 is just an example where the cheapest way to do a release-semantics store is a plain store. I never said all the world was x86 — if it were, abstract acquire/release relaxed atomics would be kind of pointless.
Lets say you have a packed structure with a type of
struct MyPackedStruct {
uint8_t byteValue;
uint32_t intValues[5];
}
Normally a compiler would add 3 extra bytes between the uint8_t value and the uint32_t array to keep alignment of the int array, but that won't be the case if you force it to be packed. This results in uint32 values that span two word boundaries. If you access one of those array values in code directly, the compiler is smart enough to perform 2 separate memory reads and combine the resulting value so you don't have to really think about it.But if you do something like this:
uint32_t *intValues = myPacktedStruct.intValues;
The compiler allows this, but the resulting intValues pointer loses all packed awareness, and trying to dereference that pointer will result in an unaligned memory access exception.Moral of the story - only use packed data types when serializing/deserializing a protocol stream. Avoid using packed data types 'at rest', because it can cause subtle issues like this. The downside is that it results in extra parsing work when converting between packed and unpacked types.
uint32_t [[aligned(1)]]*intValues = myPackedStruct.intValues;
Would be nice if we also got e.g. float/intNua_t for unaligned types. Any code that gets broken by this change was broken to begin with.I believe that situation is impossible in standard C.
C is not a high-level assembly language, as pointed out in the article.
So just because you know that the target CPU supports e.g. unaligned accesses does not make them suddenly valid C.
The C standard requires that pointers generally be created from referencing a valid, existing object. A misaligned integer is not a valid "object", thus the compiler may assume that all pointers to ints are aligned (to 4 bytes.)
https://en.cppreference.com/w/cpp/language/pointer
Every value of pointer type is one of the following:
- a pointer to an object or function (in which case the pointer is said to point to the object or function), or
- a pointer past the end of an object, or
- the null pointer value for that type, or
- an invalid pointer value.
Relatedly, `malloc` is required to return a pointer with the largest alignment of any type ("suitably aligned to hold an object".)Edit/Add: that's C++ in the link above, but it's the same in C. You can get a pointer from either a valid static/stack object, or malloc, which is "max" aligned. Either way it's required to point to a valid object, which includes correct alignment. If you implement your own memory management / allocator, it's also your responsibility to meet these same alignment requirements.
But, yes. Edited to clarify.
FWIW this isn't a new problem either. If you wrote some C code in the 90ies using misaligned pointers, and then tried porting it to, say, a DEC Alpha, you'd get a SIGBUS in your face on first dereferencing such a pointer.
https://medium.com/@iLevex/the-curious-case-of-unaligned-acc...
> Strictly speaking, the function f invokes Undefined Behavior when it computes [the misaligned pointer]
But he thinks it is still a problem with GCC, in that GCC should not be allowed to take advantage of this particular UB. He explains his view much better in the bug report:
> GCC assumes that pointers must be aligned as part of its optimizations, even if the ISA does not force it (for instance, x86-64 without the vector instructions). The present feature wish as for an option to make it not make this assumption.
> Since the late 1990s, GCC has been adding optimizations based on undefined behavior, and “breaking” existing C programs that used to “work” by relying on the assumption that since they were compiled for architecture X, they would be fine. The reasonable developers have been kept happy by giving them options to preserve the old behavior. These options are -fno-strict-aliasing, -fwrapv, ... and I think there should be another one.
And note that this slowdown would have to apply to every pointer operation in your program because GCC can't know whether a pointer is aligned at compile-time. I'm sure many more people would complain about significant performance slowdowns in a GCC update than about unaligned memory accesses being something you need to do with great care in C.
No, because I think that on x86 the instructions are the same for aligned and unaligned up to, and including, 64 bits (quadword). The only slowdown that would occur is due to missed compiler optimizations, and the author proposed disabling them as an option.
I'm pretty sure this this is not true for modern x64, which is one of the most common desktop architectures: https://lemire.me/blog/2012/05/31/data-alignment-for-speed-m...
Which current architectures are you thinking of? Do you have comparable benchmarks showing the slowdown?
I am not saying or thinking that there is a problem with GCC. I do think that GCC and Clang would be more useful with an option to make them not assume that every pointer is aligned if the target architecture does not impose this, but that's not the same thing as saying there is something wrong with GCC.
The message of the post, rather than “something is wrong with GCC”, is, “Beware. You might think that this is okay to do in your C programs, but it is not and here is why.”
Also before I post something like this, I need a confirmation that the behavior is intended and not accidental. It has happened to me before that I was about to document that GCC had an agressive behavior with respect to a kind of optimization (while remaining arguably in line with the intent of the standard, even if the word of the standard was in this case ambiguous enough to be interpreted any which way), and my co-author and I had to use a “missed optimization” ticket on GCC's bugzilla in order to have them confirm that GCC was doing the thing in question on purpose. GCC's developers, seeing the bug report, changed the behavior to remove the optimization entirely instead: https://gcc.gnu.org/ml/gcc/2016-11/msg00111.html
Coming back to the example at hand, if I had phrased the ticket as “GCC shouldn't optimize this”, it would have been closed instantly as “well it's UB”. I hoped for a more interesting search for a trade-off that would satisfy everyone, from people who just want legacy C code to keep working with new compilers to people who want programs to run as fast as possible if I phrased it this way.
(And yes, you have to ask in the bugzilla if you need some sort of official answer for this kind of thing. If you ask on a mailing list, you'll get a “no that was UB from the start” answer from someone you have never heard of who is in fact a power user who subscribed to the mailing list, and whose opinion, while useful, should not be assumed to be that of the compiler developers.)
If you aren't already using all the sanitizers that come with your {CLang, GCC} compiler, you should! They are great!
UBSan detects everything that can be detected without metadata. It would be its job to find this, since this is a simple mask to apply and test at each pointer access.
UBSan cannot detect if memory is initialized or if a pointer is valid, because these questions cannot be answered locally, looking only at the instruction doing the access. You need metadata for this. The sanitizers that maintain the metadata to answer these questions are respectively MSan and ASan. Their heavy instrumentations are incompatible, so you can only use one at a time.
I'm wondering though, does UBSAN catch this?
It is common to reference memory mapped peripherals by casting an integer to a struct pointer. In that case only a human can certify the validity of the object.
Say you were trying to talk to three UART registers at 0xA000, 0xA004, and 0xA008. It's legal to access:
*(volatile uint32_t *)(0xA000);
*(volatile uint32_t *)(0xA004);
*(volatile uint32_t *)(0xA008);
The struct equivalent of this would be something like: struct UART_Peripheral {
volatile uint32_t control;
volatile uint32_t txdata;
volatile uint32_t rxdata;
};
struct UART_Peripheral *uart0 = 0xA000;
uart0->control;
uart0->txdata;
uart0->txdata;
As long as your structure is packed/aligned correctly, this is perfectly valid C. You can enforce the structure packing with an __attribute__ ((__packed__)) if using GCC or Clang. I think MSVC has a pragma for it.I know that the C language spec allows for structure padding, but I'm not sure if the concept is applicable if every member of the structure is the native word-size of the target CPU.
But regardless of whether the spec specifies this expicitly or not, I've never seen a compiler insert padding in-between struct members that are already word-aligned.
So the code might be very slightly questionable if written in pure C with no packing pragmas/attributes - but even then it would probably work on almost every target. And if used with the right pragma/attribute for your compiler, it's a perfectly valid construct. And it's also a common one - especially in embedded work or hardware drivers.
> It is common to reference memory mapped peripherals by casting an integer to a struct pointer.
If you check the C standard you'll see that C guarantees nothing about what happens when you do this. I work with a compliant C implementation where this would not work.
Which implementations?
GCC? The compiler used for the Linux kernel? Why do you think that?
https://gcc.gnu.org/onlinedocs/gcc/Arrays-and-pointers-imple... does not clarify what they do in this case. So it's not defined by the language nor the compiler documentation. How do we know what it does?
Do you check the disassembly? What if it compiles one way one day and another the next because something else changed in your code?
If you write C code to do this and compile it with the most common C compiler, and it works, then you're just lucky! Nobody is guaranteeing it will do anything at all!
It's much better to do it in assembly, where you know it's doing what you want.
You can find 100's of these tutorials online. One example: https://developer.arm.com/tools-and-software/embedded/legacy...
Section 6.3.2.3 of the C standard indicates it is implementation defined: "An integer may be converted to any pointer type. Except as previously specified, the result is implementation-defined, might not be correctly aligned, might not point to an entity of the referenced type, and might be a trap representation." (See http://www.open-std.org/jtc1/sc22/wg14/www/docs/n1570.pdf )
I don't know who you're quoting there, but it isn't me! I said C "does not support" it. And it doesn't!
> Section 6.3.2.3 of the C standard indicates it is implementation defined
Yes... and how does the implementation tells it defines it?
If you check the GCC documentation you'll find they don't commit to doing anything when you cast an integer to a pointer and then use it. They deliberately leave it undefined themselves. It could do absolutely anything. And not necessarily what you think it should based on existing code and online tutorials.
If I was reviewing your code that did this, and I asked you "how do you know that this will work as you expect", you'd have nothing to stand on to back it up! You'd just be able to say "well everyone assumes it works". That's not good engineering!
I know it's a common pattern to use. But who is supporting that it will work? Not the C language spec. Not the compiler documentation. It's not supported by anyone. Hence, "does not support it".
Make sense now?
You can look at the ARM example I included to see how one implementation defines it. It that correct or incorrect? It's implementation defined and supported by compiler vendors. Maybe not yours.
If you can't find out if it's supported for your environment, you shouldn't do it. I agree. But to claim you can't do this with C in general is silly. You clearly can.
Nothing in the OpenGL specs says the buffers passed to it (think it was glBufferData) has to be 16byte aligned, but with NVIDIA it would crash with an alignment exception if they were not. The memory allocater in the language I used, Delphi, only used 4 byte alignment (32bit).
They might have fixed this, this was well over a decade ago.
I believe this behavior is actually specified in the standard, actually, in the same section that defines the aliasing rules.
C was portable assembly back in the early 1980s. The result was that it was crushed in performance by Fortran for scientific codes, and most of the difference had to do with assumptions the compiler could make about aliasing (or the lack of it). This was fixed by the first ANSI C standard (finalized in 1989).
Yet I still see people who weren't even born then, or were young children then, talk as if C is supposed to be "portable assembly". It isn't.
Calling it "fixed", as if it were a problem of the language rather than programmers' expectations, is a rather slanted view of the history.
I make sense of the C standard for a living (this is literally my day job) and I do not see what clause of the C standard you are referring to. It would be very useful to me to know which clause you are referring to, and I would be eternally thankful.
It's true that fancy tricks like LTO can make that kind of behavior visible to the optimizer, but the standard predates that kind of magic by several decades.
On an architecture where "long long" is 8 bytes, but only has a 4 byte alignment requirement, the compiler knows this, and knows that 2 pointers of type "long long * " may overlap. Therefore it would disable the optimization. (Except the pointers still aren't supposed to overlap, but...)
Torvalds rant: https://www.yodaiken.com/2018/06/07/torvalds-on-aliasing
edit: I misread the parent as wondering about a compiler flag
Many programs rely on misaligned access, and features like pragma packed do too. I doubt they’ll be able to break all of that legacy code.
It's already broken. You're just not seeing it on x86. (Not sure how tolerant ARM is.)
I work in network software engineering, i.e. PowerPC used to be a common architecture. Those never accepted misaligned pointers (albeit you could of course trap the misalignment exception and emulate the access, or print a backtrace to identify the location.)
Really this is a problem caused by the "common case" (x86) being permissive in what it accepts, creating fallout further down the line if you ever need the code to run on something less permissive.
[Edit: actually, oops, it's not PowerPC; those are relatively permissive too. Pretty sure DEC Alpha and m68k are restrictive.]
Alignment can make caching simpler: an aligned load/store will never cross a cache boundary. So some architectures will have faster aligned access. This is the case for RISC-V. It's not so much that unaligned access is being punished there, as aligned access being optimized
It is, indirectly, since pointers are required to point to valid objects, and valid objects are required to be correctly aligned.
Plus on a some 32-bit ISA, a long long and a double only need to be aligned to 32-bit boundaries, so I note that in the made-up C rules that you are referring to, “basic type” is not very well defined.
> I believe this behavior is actually specified in the standard, actually,
> in the same section that defines the aliasing rules.
The strict aliasing rules are here: https://port70.net/~nsz/c/c11/n1570.html#6.5p7
Go ahead and point to the rule that says that “basic types” cannot overlap with themselves.
Edit: this is the text I was remembering, from 6.5.16.1 ("Simple Assignment"): "If the value being stored in an object is read from another object that overlaps in any way the storage of the first object, then the overlap shall be exact and the two objects shall have qualified or unqualified versions of a compatible type; otherwise, the behavior is undefined.".
That pretty much matches exactly what I was saying: compilers are free to assume that basic types don't overlap, because if they do then any generated code will be undefined behavior anyway.
It does not apply to “lvalue = 1;” or to “lvalue = 2;”, which are the two relevant assignments in the example in the article.
For context, I think I made it clear in the article that the program being discussed is UB, and therefore that the compiler is not to blame. But since I wrote this article, I have had people telling me “The complaint isn't about alignment at all, it's that the optimizer assumes that two pointers to the same basic type cannot overlap in memory”.
My reply to this specific sentence is:
No. You are wrong. There are no words in the standard that say that “basic types cannot overlap in memory”. There is not even a notion of “basic type”. There are clauses about pointer alignment, that are explicitly cited in the article, and there are clauses about strict aliasing, that are shown in the article not to be the reason for GCC optimizing the program by using -fno-strict-aliasing. There are no rules about “basic types not overlapping” in the C standard. You only think there are. Or please cite them. (6.5.16.1 is a rule about assignment, it only applies for the code pattern lvalue1 = lvalue2;)
The section on "simple" assignments doesn't say that the rvalue must be an lvalue expression syntactically . I think it applies very well to "*p = 1;", which is the statement in the linked code. What am I missing?
> There are no words in the standard that say that “basic types cannot overlap in memory”.
I don't believe I said there were. I said the standard expressly allowed the optimization in the linked article. And as far as I can see, absent a clearer explanation for why that section doesn't apply, it does.
You are interpreting the C standard as if it were a philosophy text. It contains a rule that says that in very precise circumstances (for an assignment from one to the other) objects must not partially overlap, and you are claiming that it means that “two pointers to the same basic type cannot overlap in memory”. The clause does not say that, sorry. The clause applies to the objects that are on one side and the other of an assignment.
> I said the standard expressly allowed the optimization in the linked article.
I hope that the article makes it clear that the standard expressly allows the optimization. Specific, explicit rules, cited in the article, about pointer alignment, allow the optimization.
For this reason, I, “violently” as you say, disagree with the sentence “The complaint isn't about alignment at all, it's that the optimizer assumes that two pointers to the same basic type cannot overlap in memory”. This sentence gets it all wrong. It is about alignment; it is not about “pointers to basic types”, whatever that is, not being allowed to overlap in memory; they are allowed to overlap for large enough “basic types” because it is about alignment, not overlap:
https://gcc.godbolt.org/z/ZAMkeH
You could argue that GCC 9.3 only missed the optimization in the example in this Compiler Explorer link for some other reason and that absence of optimization doesn't mean that p and q cannot overlap. This would be correct, this aspect is one of the difficulties in studying the rules that these compilers implement. However, what I am saying is that if you reported this missed optimization to GCC developers, they would tell you that GCC can't optimize the function f because p and q can overlap. There is no clause in the C standard that prevent them to (apart from strict aliasing rules, but I used the option to tell the compiler I didn't want it to take advantage of these ones).
(Please do not bother them with this, or if you do, at least leave me out of it; I have nothing better to do than to write this because it's the week-end but they have better things to do.)
> No. You are wrong. There are no words in the standard that say that “basic types cannot overlap in memory”.
§ J.2, Undefined Behavior
An object is assigned to an inexactly overlapping object or to an exactly overlapping object with incompatible type (6.5.16.1).
> There is not even a notion of “basic type”.
"Object."
You repeatedly (in this thread, and on your blog) express that you don't really understand "strict" (ISO standard) aliasing rules, and that seems to be the case.
You keep quoting this clause as if it applied to any of the assignments in the program being discussed.
It doesn't.
That clause says that in an assignment of the form “lvalue1 = lvalue2;”, there must only be exact overlap or no overlap between lvalue1 and lvalue2. This does not apply to assignments of the form “lvalue = 1;” or “lvalue = 2;” which are the interesting assignments in the program being discussed.
Objects are not “basic types” for the original sentence that claimed that “basic types cannot overlap in memory”. Objects overlap in memory all the time.
> You repeatedly (in this thread, and on your blog) express that you don't really understand "strict" (ISO standard) aliasing rules, and that seems to be the case.
If you say so. I'm not the one who thinks that “* p” and “1” overlap.
> It doesn't.
You keep asserting that 6.5.16.1 is not relevant, as if it makes it so; but it doesn't. It's your opinion; the assertions are not persuasive.
void f(void) {
char *t = malloc(1 + sizeof(int));
if (!t) abort();
int *fp = (int*)t;
int *fq = (int*)(t+1);
h(fp, fq);
int h(int *p, int *q){
*p = 1;
*q = 1;
return *p;
}
Please explain to me why you continue to believe that is not an object being assigned to an inexactly overlapping object or to an exactly overlapping object with incompatible type?The clause says:
“If the value being stored in an object is read from another object that overlaps in any way the storage of the first object, then the overlap shall be exact and the two objects shall have qualified or unqualified versions of a compatible type; otherwise, the behavior is undefined.”
Under “6.5.16.1 Simple assignment”, so this describes a rule about assignment.
Which assignment in the program are you claiming stores in an object a value read from another object that overlaps in any way the storage of the first object?
If it's meant to be in a certain way, it is because it would not only simplify the implementation, but also - "I hope you know what you are doing."
Of course, the point of this is that if you feel strongly against it, I'd rather you submit a patch / use a fork where it shows benefit. If enough people require this behaviour, they may enable it in the tree. They're looking for more contributors, not less.
Is there an example where things break on misaligned access to non-overlapping objects?
That's also why this rule is in the C standard to begin with. Misaligned accesses need different / multiple instructions on some of these architectures, making them significantly more expensive. So the decision was made in favor of assuming things are aligned.