The Byte Order Fiasco
justine.lol
justine.lol
One of the few reasons I ever even reached to C is the ability to slurp in data and reinterpret it as a struct, or the ability to reason in which registers things will show up and mix in some `asm` with my C.
I think there should really be a dialect of C(++) where the machine model is exactly the physical machine. That doesn't mean the compiler can't do optimizations, but it shouldn't do things like prove code as UB and fold everything to a no-op. (Like when you defensively compare a pointer to NULL that according to spec must not be NULL, but practically could be...)
`-fno-strict-overflow -fno-strict-aliasing -fno-delete-null-pointer-checks` gets you halfway there, but it would really only be viable if you had a blessed `-std=high-level-assembler` or `-std=friendly-c` flag.
Sounds great, until you have to rewrite all your software to go from x86-64 to ARM
However for the case in hand, it would suffice to just write the key routines in Assembly, not everything.
I don't understand what you mean by that. The direct equivalent of what? Endianess is not part of the type system in C so I'm not sure I follow.
> I think there should really be a dialect of C(++) where the machine model is exactly the physical machine.
Linus agrees with you here, and I disagree with both of you. Some UBs could certainly be relaxed, but as a rule I want my code to be portable and for the compiler to have enough leeway to correctly optimize my code for different targets without having to tweak my code.
I want strict aliasing and I want the compiler to delete extraneous NULL pointer checks. Strict overflow I'm willing to concede, at the very least the standard should mandate wrap-on-overflow ever for signed integers IMO.
Portability is still plenty relevant.
Right, which is why the kind of UB pedantry in the linked article is hurting and not helping. Cranky old man perspective here:
Folks: the fact that compilers will routinely exploit edge cases in undefined behavior in the language specification to miscompile obvious idiomatic code is a terrible bug in the compilers. Period. And we should address that by fixing the compilers, potentially by amending the spec if feasible.
But instead the community wants to all look smart by showing how much they understand about "UB" with blog posts and (worse) drive-by submissions to open source projects (with passive agressive sneers about code quality), so nothing gets better.
Seriously: don't tell people to shift and mask. Don't pontificate over compiler flags. Stop the masturbatory use of ubsan (though the tool itself is great). And start submitting bugs against the toolchain to get this fixed.
Shifts and ors really is the sanest and simplest way to express "assembling an integer from bytes". Masking is _a_ way to deal with the current C spec which has silly promotion rules. Unsigned everything is more fundamental than signed.
* Undefined behavior --- behavior, upon use of a nonportable or
erroneous program construct, of erroneous data, or of
indeterminately-valued objects, for which the Standard imposes no
requirements. Permissible undefined behavior ranges from ignoring the
situation completely with unpredictable results, to behaving during
translation or program execution in a documented manner characteristic
of the environment (with or without the issuance of a diagnostic
message), to terminating a translation or execution (with the issuance
of a diagnostic message).
In the past compilers "behaved during translation or program execution in a documented manner characteristic of the environment" and now they've decided to "ignore the situation completely with unpredictable results". So yes what gcc and clang are doing is hostile and dangerous, but it's legal. https://justine.lol/undefined.png So let's fix our code. The blog post is intended to help people do that.No; I say we force the compiler writers to fix their idiotic assumptions instead of bending over backwards to please what's essentially a tiny minority. There's a lot more programmers who are not compiler writers.
The standard is really a minimum bar to meet, and what's not defined by it is left to the discretion of the implementers, who should be doing their best to follow the "spirit of C", which ultimately means behaving sanely. "But the standard allows it" should never be a valid argument --- the standard allows a lot of other things, not all of which make sense.
A related rant by Linus Torvalds: https://bugzilla.redhat.com/show_bug.cgi?id=638477#c129
Now obviously there are lots of counter-examples for that. You can probably list ten in a minute. But it should be the guiding philosophy of compiler optimizations. If the programmer wrote some code, it shouldn't just be removed. If the program would be faster without that code, the programmer should be the one responsible for deciding whether the code gets removed or not.
As far as I understand it, they do neither. Transforming an AST to any level of target code is not done by handcrafted recipes, but instead is feeded into efficient abstract solvers which have these assumptions as an operational detail. E.g.:
p = &x;
if (p != &x) foo(); // optimized out
is not much different from if (p == NULL) foo(); // optimized out
printf("%c", *p);
No assumption here is idiotic, cause no single human was involved, it’s just a class of constraints, which alone to separate properly you’ll have to scratch your head extensively (imagine telling a logic system that p is both 0 and not-0 when 0-test is “explicit” and asking it to normally operate). Compiler writers do not format disks just to punish your UBs. Of course you can write a boring compiler that emits opcodes at face expr value, without most UBs being a problem. Plenty of these, why not just take one?Compiler writers do not format disks just to punish your UBs.
IMHO if the compiler exploiting UB is leading to counterintuitive behaviour that's making it harder to use the language, the compiler is the one that needs fixing, regardless of whether the standard allows it. "But we wrote the compiler so it can't be fixed" just feels like a "but the AI did it, not me" excuse.
If the compiler doesn't know if foo may modify p, then it can't remove the call. Even if it can prove that foo does not modify p, it still can't remove the call: foo may still have some other side-effects that matter (like not returning --- either longjmp()'ing elsewhere or perhaps printing an error message about p being null and exiting?), so it won't even get to the null dereference.
As a programmer, if I write code like that, I either intend for foo to be doing something to p to make it non-null, or if it doesn't for whatever reason, then it will actually dereference the null and whatever happens when that's attempted on the particular platform, happens. One of the fundamental principles of C is "trust the programmer". In other words, by trying to be "helpful" and second-guessing the intent of the code while making assumptions about UB, the compiler has completely broken the expectations of the programmer. This is why assumptions based on UB are stupid.
The standard allows this, but the whole intent of UB is not so compiler-writers can play language-lawyer and abuse programmers; things it leaves undefined are usually because existing and possible future implementations vary so widely that they didn't even try to consider or enumerate the possibilities (unlike with "implementation-defined").
I'm more of two minds about that other step, where the compiler goes like, "here in the printf call the p will be dereferenced, so it surely is non-null, so we silently optimize that other thing out where we consider the possibility of it being null".
Also @joshuamorton, couldn't the compiler at least print a warning that it removed code based on an assumption that was inferred by the compiler? I really don't know a lot about those abstract logic solver approaches, but it feels like it should be easy to do.
That would dump a ton of warnings from various macro/meta routines, which real-world C is usually peppered with. Not that it’s particularly hard to do (at the very least compilers know which lines are missing from debug info alone).
int main() {
char *p;
p = mmap(0, 65536, PROT_READ | PROT_WRITE, MAP_PRIVATE | MAP_ANONYMOUS | MAP_FIXED, -1, 0);
// ...
return __builtin_popcountl((uintptr_t)p);
}
Or you do this: void ContinueOnError(int sig, siginfo_t *si, ucontext_t *ctx) {
xed_decoded_inst_zero_set_mode(&xedd, XED_MACHINE_MODE_LONG_64);
xed_instruction_length_decode(&xedd, (void *)ctx->uc_mcontext.rip, 15);
ctx->uc_mcontext.rip += xedd.length;
}
int main() {
signal(SIGSEGV, ContinueOnError);
volatile long *x = NULL;
printf("*NULL = %ld\n", *x);
}Yes, the assumption that p is non-null is idiotic. Also, the implicit assumption that foo will always return.
> no single human was involved
Humans implemented the compilers that use the spec adversarially and humans lobby the standards committee to not fix the bugs
> Of course you can write a boring compiler that emits opcodes at face expr value, without most UBs being a problem. Plenty of these, why not just take one
The majority of optimizations are harmless and useful, only a handful are idiotic and harmful. I want a compiler that has the good optimizations and not the bad ones.
Which results in undefined behavior according to the C ISO standard.
Quote:
“2 All declarations that refer to the same object or function shall have compatible type; otherwise, the behavior is undefined.”
From: http://www.open-std.org/jtc1/sc22/wg14/www/docs/n1256.pdf 6.2.7
struct whatever p;
fread(p, sizeof(p), 1, fp); union reinterpret {
char raw[100];
struct myStruct interpreted;
} example;
read(fd, &example.raw)
struct myStruct dest = interpreted;
This is standard-compliant C code, and it is a common way of reading IP addresses from packets, for example. struct myStruct example;
read(fd, &example, sizeof(example));
That "should present no problem unless binary data written by one implementation are read by another" quoth ANSI X3.159-1988. One example of a time where I've used that, is when storing intermediary build artifacts. Those artifacts only exist on the host machine. If the binary that writes/reads those artifacts gets recompiled, then the Makefile will invalidate the artifacts so they're regenerated. Since flags like -mstructure-size-boundary=n do exist and ABI breakages have happened with structs in the past.People like to say „C is close to the metal“. Really not true at all anymore.
A language that does the same things regardless of endianness would not have pointer arithmetic. That is not ASM and not C.
Which given its heritage, that is what PDP-11 C used to be, after all BCPL origin was as minimal language required to bootstrap CPL, nothing else.
Actually, I think TI has a macro Assembler with a C like syntax, just cannot recall the name any longer.
Practically speaking, common compilers have intrinsics for bswap. The memcpy function can be thought of as an intrinsic for unaligned load/store.
(I'm not sure how to answer the question... what do you mean, "when?")
You know the byte order of the data. But the tricky part is, what is the byte order of the platform?
It will always be correct, but you can't just assume that the compiler will optimize the shifts into a byteswap instructions. If you look at the article you will see that it tires to no-true-scotsman that concern away by talking about a "good modern compiler".
https://github.com/libsdl-org/SDL/blob/9dc97afa7190aca5bdf92...
[1]: https://developer.arm.com/documentation/dui0489/h/arm-and-th...
uint32_t swap32(uint32_t x) { ... }
#if __BYTE_ORDER__ == __ORDER_LITTLE_ENDIAN__
uint32_t swap32be(uint32_t x) { return swap32(x); }
#elif __BYTE_ORDER__ == __ORDER_BIG_ENDIAN__
uint32_t swap32be(uint32_t x) { return x; }
#else
#error "Unknown endian"
#endif
You can make the preprocessor condition broader if you care about more compilers and more platforms. Yes, I'm making assumptions about which platforms you want to target... which is fine. No, I don't care about your PDP-11, nor about dynamically changing your endian at runtime. Nearly any problem in C can be made arbitrarily difficult if you care about sufficiently bizarre platforms, or ask that people write code that is correct on any theoretical conforming C implementation. So we pick some platforms to support.The above code is fairly simple. You can separate the part where you care about unaligned memory access and the part where you care about endian.
Some irrelevant details left out above.
#define READ32BE(p) bswap_32(*(uint32_t *)(p))
Which as you correctly state in the article, is incorrect code. We agree about this. I proposed an alternate solution, where the READ32BE would be like this: uint32_t read32be(const void *ptr) {
uint32_t x;
memcpy(&x, ptr, sizeof(x));
return swap32be(x); // Nop on big-endian.
}
What I like about this is that it breaks the problem down into two parts: reading unaligned data and converting byte order. The reason for this is, sometimes, you need a half of that. Some wire formats have alignment guarantees, and if you know that the alignment guarantees are compatible with your platform, you can just read the data into a buffer and then (optionally) swap the bytes in place.Just to give an example... not too long ago I was working with legacy code that was written for MIPS. Unaligned access does not work on MIPS, so the code was already carefully written to avoid that. All I had to do was make sure that the data types were sized (e.g. replace "long" with "int32_t") and then go through and byte swap everything.
struct Something {
int32_t x, y;
char name[16];
};
void Something_Swap(struct Something *p) {
p->x = swap32be(p->x);
p->y = swap32be(p->y);
}
So it's nice to have a function like swap32be(), and "you don't have to mask and shift" I would say is true, it just depends on which compilers you want to support. I would say that a key part of being a C programmer is making a conscious decision about which compilers you want to support.Yes, I'm aware that structs are not a great way to serialize data in general, but sometimes they're damn convenient.
What I think the language is missing is a way to clearly write "this might be unaligned and/or wrong endianness, handle that". (Sometimes compilers provide intrinsics for this sort of gap, as they do with popcount and count-leading-zeroes; sometimes they recognize common open-coded idioms. But proper standardised support would be nicer.)
I agree about bsf/bsr/popcnt. I wish ASCII had more punctuation marks because those operations are as fundamental as xor/and/or/shl/shr/sar.
The idea of C that "just" does a straightforward machine translation breaks down almost immediately. For example, you'd want `int` to just overflow instead of being UB. But then it turns out indexing `arr[i]` can't use 64-bit memory addressing modes, because they don't overflow like a 32-bit int does. With UB it doesn't matter, but a "straightforward C" would emit unnecessary separate 32-bit mul/shift instructions.
https://gist.github.com/rygorous/e0f055bfb74e3d5f0af20690759...
The value of compiler optimization isn't the same thing as the value of having extensive undefined behaviour in a programming language.
Rust and Ada perform about the same as C, but lack C's many footguns.
> indexing `arr[i]` can't use 64-bit memory addressing modes
What do you mean here?
x = *(y + z);
where y and z are both 64-bit integers. If I had int arr[1000];
initialize(&arr);
int i = read_int();
int x = arr[i];
print(x);
then to get x I'd need to do something like, tmp = i * 4;
tmp1 = (uint64_t)tmp;
x = *(arr + tmp1);
Which, since i is signed, can't just be a cheap shift, and then needs to be upcasted to a uint64_t (which is cheap, at least).UB doesn't just mean the compiler can treat it as a no-op. It means the compiler can do whatever it likes and still be compliant with the spec.
From the POV of someone consulting the spec, if something results in UB, what it means is: "Don't look here for documentation, look in the documentation of your compiler!".
Many compilers prefer to do a no-op because it is the cheapest thing to do.
u32::from_le_byte(bytes) // u32 from 4 bytes, little endian
u32::from_be_byte(bytes) // u32 from 4 bytes, big endian
u32::to_le_bytes(num) // u32 to 4 bytes, little endian
u32::to_be_bytes(num) // u32 to 4 bytes, big endian
This was very useful to me recently as I had to write the marshaling and un-marshaling for a game networking format with hundreds of messages. With primitives like this, you can see what's going on.Big endian is dead for game developers.
Copy entire arrays of structs onto the wire without fear!
(Just #pragma pack them first)
Maybe your program isn't a game.
Maybe you have to deal a server that uses Power, or an embedded system that uses PowerPC (or ARM or MIPS in big-endian mode).
Maybe you're running on an older architecture (SPARC, PowerPC, 68K.)
Maybe you have to deal with a pre-defined data format (e.g. TCP/IP packet headers) that uses big-endian byte ordering for some of its components.
I had to read through excessively-clever C++ code that did the same thing to figure out what conversions were happening, then re-express it in Rust. I'm re-implementing a legacy mess that people are afraid to work on. As it happens, in this message system, some items, mainly packet sequence numbers, are big-endian, because they were following what IP and UDP do, and everything else is little endian.
I know how to do this with shifts and masks, and I've done things like that when programming in assembly. That was a long time ago. There's been progress in how to write programs.
But what's the argument for using them in new game code?
The code will not be running anywhere that has big-endianness. No current platform a game could run on uses it, and I can't imagine a scenario where a new platform would come into existance and use it either.
If you insist on using network-byte order anyway, then you have to do an extra bswap op for each bit of data you send. Sure, the cost of that is super minor and probably not worth worrying about.
But the bigger cost is that you can't just send whole structures at a time. You have to individually serialise each thing. Now you have to have a whole serialisation concept. You have to have some way of enumerating all the fields. You have to walk all the structures. What a pain.
If you want to send a thing over the network, just send it.
when we ported little endian x86 Linux to the big endian mainframe we sprinkled hton/ntoh all over the place, happily so. they are the way to go and they should be implemented properly, not be replaced by a homegrown version.
all that said, I'm surprised 64bit htonll and ntohll are not standard yet. anybody knows why?
The original listing was written by “Jennifer — please call me if you have troubles”, an undergraduate from MIT. It was hand-assembled machine code, in a neat hand in a big blue binder. That code ran non-stop except for a few hurricanes from 1988 until 2008; bug-free as far as I could tell. Jennifer last-name-unknown, you were my idol & my demon!
I swore off programming for nearly a year after that.
fn from_be(bytes: [u8; 4]) -> u32 {
(bytes[0] as u32) << 24
| (bytes[1] as u32) << 16
| (bytes[2] as u32) << 8
| (bytes[3] as u32) << 0
}
It's direct, to the point, and does exactly what it says on the tin because all pertinent behaviour is defined. The way Rust's corelib implements it is to transmute the array into the integer, then call the bswap intrinsic if the bytes need swapping(detected at compile time).For reference, check how Linux does it. https://elixir.bootlin.com/linux/latest/source/include/linux...
> Because unsigned char in C expressions gets type promoted to the signed type int.
> So if we say 0x80<<24 it overwrites the sign bit, which is an undefined behavior
Does the u8 type protect against this?
The following is perfectly well defined in C++, despite looking like almost the same as the original unsafe C:
#include <boost/endian.hpp>
#include <cstdio>
using namespace boost::endian;
unsigned char b[5] = {0x80,0x01,0x02,0x03,0x04};
int main() {
uint32_t x = *((big_uint32_t*)(b+1));
printf("%08x\n", x);
}
Note that I deliberately misaligned the pointer by adding 1.https://gcc.godbolt.org/z/5416oefjx
[Edit] Fun twist: the above code doesn't work where the intermediate variable x is removed because printf itself is not type safe, so no type conversion (which is when the bswap is deferred to) happens. In pure C++ when using a type safe formatting function (like fmt or iostreams) this wouldn't happen. printf will let you throw any garbage in to it. tl;dr outside embedded use cases writing C in 2021 is fucking nuts.
This example is demonstrating:
- First class treatment of user (or library) defined types
- Operator overloading
- The fact that it produces fast machine code. Try changing big_uint32_t to regular uint32_t to see how this changes. When you use the later ubsan will introduce a trap for runtime checks, but it doesn't need to in this case.
For instance I'm not familiar with this boost library so I'd have a lot of trouble piecing out what your snippet does, especially since there's no explicit function call besides the printf.
Personally if we're going the OOP route I'd much prefer something like Rust's `var.to_be()`, `var.to_le` etc... At least it's very explicit.
My hot take is that operator overloading should only ever be used for mathematical operators (multiplying vectors etc...), everything else is almost invariably a bad idea.
Boost.Endian actually lets you pick between arithmetic and buffer types.
'big_uint32_buf_t' is a buffer type that requires you to call .value() or do a conversion to an integral type. It does not support arithmetic operations.
'big_uint32_t' is an arithmetic type, and supports all the arithmetic operators.
There are also variants of both endian suffixed '_at' for when you know you have aligned access.
If I have a 3 byte big endian integer can I access it easily in rust without resorting to shifts?
In C++ I could probably create a fairly convincing big_uint24_t type and use it in a packed struct and there would be no inconsistencies with how it's used with respect to the more common varieties
I'm not sure about a 3 byte big endian integer. I mean, that's going to compile down to some combination of shifting and masking operations anyway, isn't it? I suspect that if you have some oddball binary format that needs, this it will be possible to write some code to marshal it, that compiles down to the best possible asm. Godbolt is your friend here :)
[1]: https://rust-lang.github.io/rfcs/0445-extension-trait-conven...
I think there's no need for explicit shifts. You need to memcpy anyway to deal with alignment issues, so you may as well just copy in to the last 3 bytes of a zero-initialized, big endian, 32bit uint.
https://gcc.godbolt.org/z/9qGqh6M1E
And I think we're on the same page, it should be possible to get similar results in Rust.
If you need to change byte orders, you should use library to achieve that.
This is why ubsan is silent and not even injecting a check in to the compiled code.
You can check the alignment constraints with static_assert (something else you can't do in standard C): https://gcc.godbolt.org/z/KTcf9ax6r
Is also has _Generic() so you can roll up a family of endianness conversion functions and safely change types without blowing up somewhere else with a hardcoded conversion routine.
With regard to teaching C++ specifically I tend to agree with this talk:
CppCon 2015 - Kate Gregory “Stop Teaching C": https://www.youtube.com/watch?v=YnWhqhNdYyk
The biggest problem with C++ in industry is that people tend to write "C/C++" when it deserves to be recognized as a language in its own right.
Apparently the first year students at my university didn't had any issue going from Standard Pascal to C++, in the mid-90's.
Proper C++ was taught using our string, vector and collection classes, given that we were still a couple of years away from ISO C++ being fully defined.
C style programming with low level tricks were only introduced later as advanced topics.
Apparently thousands of students managed to get going the remaining 5 years of the degree.
Just like no Python newbie is able to master Python 3.9 full language set, standard library, numpy, pandas, django,...
The only subjects that went full into Java were distributed computing and compiler design.
And during the last 20 years they already went back into their decision.
I should note that languages like Prolog, ML and Smalltalk were part of the learning subjects as well.
Assembly was part of electronic subjects where design of a pseudo CPU was also part of the themes. So we had our own pseudo Assembly, x86 and MIPS.
Where ? I learned algorithms in C and C++ (and also a bit in Caml and LISP) and I was in university 2011-2014
printf("%08x\n", *((uint32_t*)(b)));
to your example and you'll see that it produces UB as well. The reason there is no UB with big_uint32_t probably is that that struct/class/whatever it is probably redefines its dereferencing operator to perform byte-wise reads.Godbolt example: https://gcc.godbolt.org/z/seWrb5cz7
Obviously if you write C and compile it as C++ you still end up with UB, because C++ aims for extreme levels of compatibility with C.
Of course your example solves both a) and b) by using big_uint32_t, and I agree that this is an interesting abstraction provided by Boost, but I think the takeaway "use C++ for low-level byte fiddling" is slightly misleading: Say I was a novice C++ programmer, saw your example of how C++ improves this but at the same time don't know that big_uint32_t solves the hassle of reading a word from an unaligned address for me. Now I use your pattern in my byte-fiddling code, but then I need to read a word in host endianness. What do I do? Right, I remember the HN post and write *((uint32_t*)(b+1)) (without the big_, because I don't need that!). And then I unintentionally introduced UB. In other words, big_uint32_t is a little "magic" in this case, as it suggests a similarity to uint32_t which does not actually exist.
To be honest, I don't think the byte-wise reading is in any way inappropriate in this case: If you're trying to read a word in non-native byte order from an unaligned access, it is perfectly fine to be very explicit about what you're doing in my opinion. There also is nothing unsafe about doing this as long as you follow certain guidelines, as mentioned elsewhere in this thread.
I still think being able to define a type that models what you're doing is incredibly valuable because as long as you don't step outside your type system you get so much for free.
Conceptual consistency is a good thing, but there is a generally higher cognitive load to using C++ over C. I've used both C++ and C professionally, and I've gone deeper with type safety and metaprogramming than most folk. I've mostly used C for the last few years, and I don't feel like I'm missing anything. It's still possible to write hard-to-misuse code by coming up with abstractions that play to the language's strengths.
Operator overloading in particular is something I've refined my opinion on over the years. My current thought is that it's best not to use operators in user/application defined APIs, and should be reserved for implementing language defined "standard" APIs like the STL. Instead, it's better to use functions with names that unambiguously describe their purpose.
uint32_t read_big_uint32(char *bytes);
Having a big_uint32_t type seems wrong to me conceptually. You should either deal with sequences of bytes with a defined endianness or with native 32-bit integers of indeterminate endianness (assuming that your code is intended to be endian neutral). Having some kind of halfway house just confuses things.If you're defining a struct to mirror a data structure from a device, protocol or file format then the language / type system should let you define the properties of the fields, not necessarily force you to introduce a parsing/decoding stage which could be more easily bypassed.
In practice, most projects (e.g. the Linux kernel or the socket interface) differentiate between host (indeterminate) byte order and a specific byte order (e.g. network byte order/big endian).
Also, I will say that when it comes to programming Arduinos and ESP8266/ESP32 chips, I still find that C is my go to despite things like Alia, MicroPython, etc. I think it’s possible that once Zig supports those devices fully that I might move over. But in the meantime I guess I’ll keep minding my off by one errors.
In my estimation, libraries like boost are way too big and way too clever and they create more problems than they solve. Also, they don't make me happy.
You're overfocusing on a "problem" that is almost completely irrelevant for most of programming. Big endian is rare to be found (almost no hardware to be found, but some file formats and networking APIs have big-endian data in them). Where you still meet it, you don't do endianness conversions willy-nilly. You have only a few lines in a huge project that should be concerned with it. Similar situation for dealing with aligned reads.
So, with boost you end up with a huge slow-compiling dependency to solve a problem using obscure implicit mechanisms that almost no-one understands or can even spot (I would never have guessed that your line above seems to handle misalignment or byte swapping).
This approach is typical for a large group of C++ programmers, who seem to like to optimize for short code snippets, cleverness, and/or pedantry.
The actual issue described in the post was the UB that is easy to hit when doing bit shifting, caused by the implicit conversions that are defined in C. While this is definitely an unhappy situation, it's easy enough to avoid this using plain C syntax (cast expression to unsigned before shifting), using not more code than the boost-type cast in your above code.
The fact that the UB is so easy to hit doesn't call for excessive abstraction, but simply a revisit of some of the UB defined in C, and how compiler writers exploit it.
(Anecdata: I've written a fair share of C code, while not compression or encryption algorithms, and personally I'm not sure I've ever hit one of the evil cases of UB. I've hit Segmentation faults or had Out-of-bounds accesses, sure, but personally I've never seen the language or compilers "haunt me".)
Re the anecdata at the end. Have you ever run your code through the sanitizers? I have. CVE-2016-2414 is one of my battle scars, and I consider myself a pretty good programmer who is aware of security implications.
Admittedly I'm not the type of person who spends his time fuzzing his own projects, so my statement was just to say that the kind of bugs that I hit by just testing my software casually are almost all of the very trivial kind - I've never experienced the feeling that the compiler "betrayed" me and introduced an obscure bug for something that looks like correct code.
I can't immediately see the problem in your CVE here [0], was that some kind of betrayal by compiler situation? Seems like strange things could happen if (end - start) underflows.
[0] https://android.googlesource.com/platform/frameworks/minikin...
Also, the fact that you can't see the problem is actually evidence of how insidious these problems are :)
The rules for this are arcane, and, while the solution suggested in OP is correct, it skates close to the edge, in that there are many similar idioms that are not ok. In particular, (p[1] << 8) & 0xff00, which is code I've written, is potentially UB (hence "mask, and then shift" as a mantra). I'd be surprised if anyone other than jart or someone who's been part of the C or C++ standards process can say why.
I've looked for a while now, but still can't see it, would you be willing to share?
> (p[1] << 8) & 0xff00
With p[1] being uint8_t? Because then I cannot imagine why, and also fail to see a reason to apply the 0xff00 mask here.
If this is for int8_t instead, the problem you are alluding to is sign extension? If p[1] gets promoted to an int in the negative range, (then its representation has the high order bit set), and shifting that to the left is UB.
(imo, both c and cpp are mainly advocated by people suffering from stockholm syndrome.)
nitpick: the 64bit versions are not fully available yet, htonll, ntohll
[0] https://gist.github.com/shafik/848ae25ee209f698763cffee272a5...
C++ has a multitude of its own pitfalls. Some of the C programmer hate for C++ is justified. After all, it's just C with a pre-processing stage in the end.
There's good reasons why many C projects never considered C++ but are already integrating the nascent Rust. I always hated low level programming until Rust made it just as easy and productive as high level stuff
And now we've settled on all machines being 8 bit bytes, and programmers no longer have to worry about such details?
Is it time to do the same for big endian machines? Is it time to accept that all machines that matter are little endian, and the extra effort keeping everything portable to big endian is no longer worth the mental effort?
A bit like the electron has a negative charge...
Little-endian is natural with casts because the address doesn't change, and it's the order in which addition takes place.
But more _natural_ is little endian because, well, it's just more straightforward to have the digits' magnitude be in ascending order (2^0, 2^1, 2^2, 2^3...) instead of putting it in reverse.
Plus you encounter less roadblocks in practice with little endian (e.g. address changes with casts) which is often a sign of good natural design
All human number systems I've ever seen write numbers out as big Endian (yes, even Roman numerals), so I'm really struggling to see how that wouldn't be considered natural.
But that's not what we're doing here, so it's not entirely relevant.
I wonder if we went big endian “by mistake” with Arabic numerals given that Arabic is written right to left.
Some ancient texts have “four and twenty” which is little endian.
We also add commas to large numbers to help with a human processing problem - you have to get to the end of the number to know what the first digit represents and then count backwards (groups of three help).
This is only true on a big endian serial connection (that is, one that, tautologically, sends the most-significant bit first). Offhand, I think most serial protocols are big endian, but by that logic, most CPUs are little endian, so that doesn't really help.
The thing that's actually useful about big endian is not that it's natural (as kangalioo points out, that's little endian) or that it's how humans write numbers (by that logic crap like BCD or decimal floats is a good idea), but that big endian preserves lexicographic order of fixed-width integers.
However endianness isn't just about supporting IBM. Modern compilers will literally break your code if you alias memory using a type wider than char. It's illegal per the standard. In the past compilers would simply not care and say, oh the architecture permits unaligned reads so we'll just let you do that. Not anymore. Modern GCC and Clang force your code to conform to the abstract standard definition rather than the local architecture definition.
It's also worth noting that people think x86 architecture permits unaligned reads but that's not entirely true. For example, you can't do unaligned read-ahead on C strings, because in extremely rare cases you might cross a page boundary that isn't defined and trigger a segfault.
But that's not a problem with an unaligned read but rather that you are reading more than you are allowed to. And in C even an aligned readahead is UB.
A better example might be SSE instructions which do have aligned variants that trap on unaligned pointers.
Vending machines have an internal protocol a little like I2C. We created a custom peripheral to bridge the machine to the web, based on a Raspberry Pi.
The protocol was defined by Coca Cola Japan in 1975 (in order to have optionality in their supply chain). It's still in use today. But because it was designed in Japan, with a need for wide characters, it assumes 9 bit bytes.
We couldn't find any way to get a Raspberry Pi to speak 9 bit bytes. The eventual solution was a custom shield that would read the bits, and reserialise to 8 bit bytes for the Pi to understand. And vice versa.
9 bit bytes. I grew up knowing that bytes had variable length, bit this was the first time I encountered it in the wild. This was 2015.
This is best solvable the closer to the device in question and in the simplest way possible.
To transmit a bit pattern 10010010 over a single pin channel, for example, you'd literally set the pin high, sleep for a some predetermined amount of time, set it low, sleep, set it low, sleep, set it high, etc.
But sometimes you can't use a UART: maybe you're working on a tiny embedded computer without one, or maybe you need to speak a weird 9-bit protocol a standard UART doesn't understand. In that case, you can make the CPU pump the serial line directly. It's inefficient (there's probably more interesting work the CPU could be doing) and it can be difficult to make the CPU pause for exactly the right amount of time (CPUs are normally designed to run as fast or as efficiently as possible, nothing in between), but it's possible and sometimes it's all you've got. That's bit-banging.
For Unicode, 32-bit bytes would be an incredibly wasteful in memory encoding.
Text generally uses a small fraction of memory and storage these days.
[1] at runtime, you could dynamically assign 'virtual' codepoints to grapheme clusters and get a fixed-length encoding for strings that way
I am thankful that almost all the Unicode text I see is rendered properly now, farewell the little boxes. Good job lots of people.
Only if you are naively operating in the Anglosphere / world where the most complex thing you have to handle is larger character sets. In reality, there's ligatures, diacritics, combining characters, RTL, nbsp, locales, and emoji (with skin tones!). Not to mention legacy encoding.
And no, it does not use a "small fraction of memory and storage" in a huge range of applications, to the point where some regions have transcoding proxies still.
"Anglosphere" would be just 7(&"8") bit ASCII, and it's the current situation where it takes quite a lot of skill and knowledge just to start learning how to properly deal with Unicode, because it's often not even taught !
IMHO 32-bit bytes would help tremendously with onboarding developers into Unicode, because it would force dumping ASCII-only as the starting point (and sadly, often ending point) for teaching how to deal with text.
And who can blame the teachers, Unicode is already hard enough without even having to deal with the difficulties coming from having to explain its multi-byte representation...
Last but not least : this would have forced standardization between the Unix world now on UTF-8 and the Windows world which is still stuck on UTF-16 (and Windows-1252 ?!?) for some of the core functions like filenames, which, for instance, still regularly results in issues working with files with non-ASCII filenames.
To be a bit more explicit: Unicode is a character encoding, to 20-and-a-half-bit 'bytes', that is variable-width in those 'bytes', even before considering how the 'bytes' are encoded to actual bytes. Eg "ψ̊" (greek small psi with ring above) is U+3C8 U+30A (two 'bytes').
> a single character can be made up of multiple code points.
It's really the other way round...
Unfortunately, the term “character“ alone is ambiguous because depending on the context it can refer to either code points or code units.
A code point is the atomic unit of the abstract Unicode encoding. By "abstract" I mean it is not an actual text encoding you can write to a file.
A code unit is the atomic unit of an actual text encoding, such as UTF-8, UTF-16LE or UTF-32LE (and their BE equivalents).
---
So to put it together a "user-perceived character" is made up of one or more "code points". When implemented in an application, each "code point" is encoded using one or more "code units".
And do you have examples of still widely used 8-bit sized data formats ?
And, by an interesting coincidence, with the arrival of "HDR", 8-bit per channel is slowly becoming obsolete (because insufficient). The next "step" is 10-bit per channel with 3 channels (hence "HDR10(+)"), and so should fit quite well in 32 bits ?
(However, it would seem that even Dolby's Perceptual Quantizer transfer function might need 12 bits per channel to avoid banding over the "HDR" Rec.2020/2100-sized color gamut..?)
And it's particularly important to have this property for text, because not only data is overwhelmingly stored as text (in importance, not by "weight"), but because computer programs themselves are written using text.
Sadly, I kind of gave up on getting a "HDR" display, at least for now, because :
- AFAIK neither Linux nor Windows have good enough "HDR" support yet. (MacOS supposedly does, but I'm not interested.)
- I'm happy enough with my HP LP2475w which I got for dirt cheap just before "HDR" became a thing. I consider the 1920x1200 resolution to be perfect for now (as a bonus I can manually scale various old resolutions like 800x600 to be pixel-perfect) - too many programs/OSes still have issues with auto-scaling programs on higher resolution screens (which would come with "HDR"). I'm also particularly fond of the 16:10 ratio, which seems to have gone extinct.
- Maybe I'll be able to run this monitor properly in wide gamuts (though with banding), or maybe even in some kind of ""HDR" compatibility mode", though it would seem that the current sellers of "HDR" screens aren't going to make that easy. I might be able to get a colorimeter soon to properly calibrate it.
I'm not sure what DICOM has to do with color reproduction quality ? Also it seems to be a quite a bit older standard than sRGB...
By definition, you can't "see the unseen". "Yellow" and "blue" are opponent "colors", so, by definition, a proper mixture of them is going to give you grey :
https://www.handprint.com/HP/WCL/color2.html#opponentfunctio...
Also, when talking about subtle color effects, you have to consider that personal variation might come into play (for instance red-green "colorblindness" is a spectrum).
But still, 'byte' refers to the smallest addressable unit of memory. There's just no point in arguing over its size...
However, since Unicode is limited to 21-bits by utf-16 encoding, a unicode code point will fit in a small integer.
[1] unless you use binaries, which is often a better choice.
Not to mention that bytes have nothing to do with unicode. Unicode codepoints can be encoded in many different ways: UTF8, UTF16, UTF32, etc.
These various ways to encode Unicode have quite a lot to do with bytes being 8-bit sized !
Anyway, it doesn't make much sense to define the size of a “byte“ as anything else then 8 bits, because that's the smallest adressable memory unit. If you need a 32 bit data type, just use one!
Again, bytes are not foremost about text. We habe to deal with all sorts of data, many of which is shorter than 32 bits.
You can always pick a larger data type for your type of work, but not the opposite.
https://news.ycombinator.com/item?id=27094663
https://news.ycombinator.com/item?id=27104860
(You'll also notice that caring about not wasting the 8th bit with ASCII has lead us into all sorts of issues... and why care so much about it when as soon as data density becomes important, we can use compression which AFAIK easily rids us of padding ?)
But again and again, all of this has nothing to do with the size of a byte.
BTW, are you aware that 8-bit Microcontrollers are still in widespread use and nowhere near of being discontinued?
Programming microcontrollers isn't considered to be "mandatory computer literacy" in college, while basic scripting, which involves understanding how text is encoded at the storage/memory level - is.
> mandatory computer literacy
I don't understand why you keep bringing up this phrase and ignore a huge part of real world computing. College students should simply learn how Unicode works. Are you seriously demanding that CPU designers should change their chip design instead?
http://pdp10.nocrew.org/docs/instruction-set/Byte.html
>In the PDP-10 a "byte" is some number of contiguous bits within one word. A byte pointer is a quantity (which occupies a whole word) which describes the location of a byte. There are three parts to the description of a byte: the word (i.e., address) in which the byte occurs, the position of the byte within the word, and the length of the byte.
>A byte pointer has the following format:
000000 000011 1 1 1111 112222222222333333
012345 678901 2 3 4567 890123456789012345
_________________________________________
| | | | | | |
| POS | SIZE |U|I| X | Y |
|______|______|_|_|____|__________________|
>POS is the byte position: the number of bits from the right end of the byte to the right end of the word.
SIZE is the byte size in bits.>The U field is ignored by the byte instructions.
>The I, X and Y fields are used, just as in an instruction, to compute an effective address which specifies the location of the word containing the byte.
"If you're not playing with 36 bits, you're not playing with a full DEC!" -DIGEX (Doug Humphrey)
Well, no, because it's not the case. SPARC is big-endian, and a bunch of IBM processors. ARM processors are mostly bi-endian.
> Is it time to do the same for big endian machines?
No. Not just because of their prevalence, but because there isn't a compelling reason why everything should be little- endian.
If you do use unsigned char, an alternative to masking would be performing the cast to uint32_t before instead of after the shift.
edit: For reference, this is what it would look like when implemented as a function instead of a macro:
static inline uint32_t read32be(const uint8_t *p)
{
return (uint32_t)p[0] << 24
| (uint32_t)p[1] << 16
| (uint32_t)p[2] << 8
| (uint32_t)p[3];
}EDIT: Reading the bug report [1], the actual cause for the missing ret is that the for loop will overflow, which is UB and causes clang to not emit any code for the function.
1) They define a char array (which defaults to signed char, as mentioned in the post), including the value 0x80 which can't be represented in char, resulting in a compiler warning (e.g. in GCC 11.1).
The mentioned reason against using unsigned char (that shifting 128 left by 24 places results in UB) is also misleading: I could not reproduce the UB when changing the array to unsigned char. Perhaps the author meant leaving the array defined as signed char, but casting the signed chars to unsigned before shifting. That indeed results in UB, but I don't see why you would define the array as signed in the first place.
2) The cause for the undefined behavior isn't the bswap_32, rather it's because they try reading an uint32_t value from a char array, where b[0] is not aligned on a word boundary.
There is no need at all do redefine bswap. The simple solution would be to use an unsigned char array instead of a char array and just reading the values byte-wise.
Of course C has its footguns and warts and so on, but there is no need to dramatize it this much in my opinion.
I've prepared a Godbolt example to better explain the arguments mentioned above: https://godbolt.org/z/Y1EWK6e17
Edit: To add to point 2) above: Another way to avoid the UB (in this specific case) would be to add __attribute__ ((aligned (4))) to the definition of b. In that case, even reading the array as a single uint32_t works as expected since the access is aligned to a word boundary.
Obviously, you can't expect any random (unsigned char) pointer to be aligned on a word boundary. Therefore, it is still necessary to read the uint32_t byte by byte.
No, that reasoning is correct. Integer promotions are performed on the operands of a shift expression, meaning the left operand will be promoted to signed int even if it starts out as unsigned char. Trying to shift a byte value with highest bit set by 24 will results in a value not representable as signed int, leading to UB.
Still, adding an explicit cast to the left operand seems to be enough to avoid this, e.g.:
uint32_t x = ((uint32_t)b[0]) << 24;
In summary, I think my point that using unsigned char would be appropriate in this case still stands.Indeed. See my other comment, https://news.ycombinator.com/item?id=27086482
A similar one is that signedness of char is machine dependent. It's typically signed on Intel and unsigned on ARM.
Sigh!
But for example bitmaps in BE are a huge source of bugs, as readers and writers need to agree on the size to use for memory operations.
"SIMD in a word" (e.g. doing strlen or strcmp with 32- or 64-bit memory accesses) might have mostly fallen out of fashion these days, but it's also more efficient in LE.
I used to like big endian more, but after deep investigation I now prefer little endian for any encoding schemes.
E.g. decimal example:
1.00/1.00 = 1.00
1.000/1.001 = 0.999000999000...
(adding one more bit changes the first bits of the outcome)For example, with ULEB128 [1], you just read 7 bits at a time, going higher and higher up the value you're reconstituting. If the value grows too big and you need to spill over to the next (such as with big integer implementations), you just fill the last bits of the old value, then put the remainder bits in the next value and continue on.
With a big endian encoding method (i.e. VLQ used in MIDI format), you start from the high bits and work your way down, which is fine until your value spills over. Because you only have the high bits decoded at the time of the spillover, you now have to start shifting bits along each of your already decoded big integer portions until you finally decode the lowest bit. This of course gets progressively slower as the bits and your big integer portions pile up.
Encoding is easier too, since you don't need to check if for example a uint64 integer value can be encoded in 1, 2, 3, 4, 5, 6, 7 or 8 bits. Just encode the low 8 bits, shift the source right by 8, repeat, until the source value is 0. Then backtrack to the as-yet-blank encoded length field in your message and stuff in how many bytes you encoded. You just got the length calculation for free. Use a scheme where you only encode up to 60 bit values, place the length field in the low 4 bits, and Robert's your father's brother!
For data that is right-heavy (i.e. the fully formed data always has real data on the right side and blank filler on the left - such as uint32 value 8 is actually 0x00000008), you want a little endian scheme. For data that is left-heavy, you want a big endian scheme. Since most of the data we deal with is right-heavy, little endian is the way to go.
You can see how this has influenced my encoding design in [2] [3] [4].
[1] https://en.wikipedia.org/wiki/LEB128
[2] https://github.com/kstenerud/concise-encoding/blob/master/cb...
[3] https://github.com/kstenerud/compact-float/blob/master/compa...
[4] https://github.com/kstenerud/compact-time/blob/master/compac...
7 6 5 4 3 2 1 0 15 14 13 12 11 10 9 8
instead of
15 to 0
This is because little endian is not how humans write numbers. For consistency with little endianness we would have to switch to writing "one hundred and twenty three" as
321
The only thing I'm aware of that's neat in little endian is that if you want the low byte (or word or whatever suffix) of a number stored at address a, then you can simply read a byte from exactly that address. Even if you don't know the size of the original number.
- Long addition is possible across very large integers by just adding the bytes and keeping track of the carry.
- Encoding variable sized integers is possible through an easy algorithm: set aside space in the encoded data for the size, then encode the low bits of the value, shift, repeat until value = 0. When done, store the number of bytes you wrote to the earlier length field. The length calculation comes for free.
- Decoding unaligned bits into big integers is easy because you just store the leftover bits in the next value of the bigint array and keep going. With big endian, you're going high bits to low bits, so once you pass to more than one element in the bigint array, you have to start shifting across multiple elements for every piece you decode from then on.
- Storing bit-encoded length fields into structs becomes trivial since it's always in the low bit, and you can just incrementally build the value low-to-high using the previously decoded length field. Super easy and quick decoding, without having to prepare specific sized destinations.
123 = 3x10^0 + 2x10^1 + 1x10^2
So if you were to go and label each digit in 123 with the power of 10 it represents, you end up with little endian ordering (eg the 3 has index 0 and the 1 has index 2). This is why little endian has always made more sense to me, personally.
I sometimes think about arithmetic in little endian, since addition always starts with the least significant digit, due to the right-to-left dependency of carrying.
Except lately I’ve been doing large additions big-endian style left-to-right, allowing intermediate “digits” with a value greater than 9, and doing the carry pass separately after the digit addition pass. It feels easier to me to think about addition this way, even though it’s a less efficient notation.
Long division and modulus are also big-endian operations. My favorite CS trick was learning how you can compute any arbitrarily sized number mod 7 in your head as fast as people are reading the digits of the number, from left to right. If you did it little-endian you’d have to remember the entire number, but in big endian you can forget each digit as soon as you use it.
Loss of big endian chips saddens me like the loss of underscores in var names in Go Lang. The homogeneity is worth something, thanks intel and camelCase, but the old order that passes away and is no more had the beauty of a new world.
https://en.wikipedia.org/wiki/Number#First_use_of_numbers
said a friend who also quips: "never trust a computer you can lift"
This is nonsense - many file formats are big endian.
POWER also still uses big endian though recently little endian POWER have gotten more popular
This is defined in C to be the order the fields are declared in.
This was not a problem for the old SPARC system, which naturally put everything in the correct order without any fuss, but one of the biggest sticking points in porting over to x64 was having to now manually pack all of that binary data. Using Ada, (what else!) of course.
> Could save a huge amount of time debugging when compilers or architecture changes.
I'm assuming we come from very different backgrounds, but it's not clear to me how switching compilers or architectures is so common that hardening code against it by default is appropriate. I would think that switching compilers or architectures is generally done very deliberately, so instrumenting code with UBsan for that transition would be the right thing to do?
I don't really want to have to take that yearly update to go through and review (and presumablu fix) all the UB that has managed to sneak in over the year. It would be better to have avoided putting it in.
https://blog.hboeck.de/archives/879-Safer-use-of-C-code-runn...
if (undefined_behavior) break_program()
If it did, it could easily report the undefined behavior. However, that's not how it works. Instead, the compiler has optimization rules that are only valid if the code contains no undefined behavior. If the code contains undefined behavior, the optimization rules change the result of the program. For example, this code: bool function(int x) {
return x + 1 > x;
}
Can be optimized to "return true". That is correct if x does not overflow, but if x overflows and wraps around the optimization changes the result of the program. In this case, it is acceptable according to the C/C++ standards for the optimization to assume x does not overflow, and hence this optimization is valid.The compiler could tell you for every instance of signed integer arithmetic that it is making assumptions about your program, and that the signed integer arithmetic could potentially overflow, but that doesn't seem particularly helpful.
In your example about the comparison of x + 1 vs x, I'm not sure that is a contraversial optimization. However this one, to me, is:
http://blog.llvm.org/2011/05/what-every-c-programmer-should-...
Here a diligent programmer is trying to do a null pointer check, but because dereferencing null is UB, then the optimizer removes the null pointer check. This is compile time UB that should be flagged to users.
Though the article only mentions bswap64 and mentioning __builtin_bswap64 would be a nice addition.
https://gcc.gnu.org/onlinedocs/gcc/Integer-Overflow-Builtins...
Or, on Linux and BSD systems at least, you can use the <endian.h> or <sys/endian.h> functions (https://linux.die.net/man/3/endian) and rely on the libc implementation to do the system/compiler detection for you and use an appropriate compiler builtin inside of an inline function instead of bothering to hack something together in your own code.
The article mentions those functions at the bottom, but strangely still recommends hacking up your own macros.
https://github.com/rustyrussell/ccan/blob/master/ccan/endian...
Fun fact: CD-ROM superblocks have both-endian fields. Each integer is stored twice in big and little endian format. I assume this was to allow underpowered 80s hardware which didn't have enough resource to do byte swapping.
> Now you don't need to use those APIs because you know the secret.
This sentiment seems problematic. The solution shouldn't be "we just have to educate the masses of C programmers on how to properly deal with endianness". That will never happen.
The solution should be "It's in the standard library. Go look there and don't think too hard." C is sufficiently low-level, and endianness problems sufficiently common, that I would expect that kind of routine to be available.
This is why people that memcpy structs right into the buf get such derision, even if it’s faster and written for a mono-Implementation of a language semantics. It is sloppy thought made manifest.
(Yes, the article mentions those, but they've been standard for decades).
The "Be reasonable, do it my way" approach does not work. Neither does the Esperanto approach of "let's all switch to yet a new language".
His bottom line conclusion being It is more important to agree upon an order than which order is agreed upon.
[0] https://www.rfc-editor.org/ien/ien137.txtC++ 20 is quite new so I would assume that very few people know this yet.
C and C++ obviously differ a lot, but by that phrase she clearly means “the part where then two languages overlap”. The C++ committee has been willing to break C compatibility in a few ways (not every valid C program is a valid C++ program), and this has been true for a while.
The C++ committee decided that everyone had figured this out by now and so made this breaking change.
That header is also available on Linux, but glibc (and compatible libraries) named it <endian.h> instead.
See: man 3 endian (https://linux.die.net/man/3/endian)
Of course it gets a bit hairier if the code is also supposed to run on other systems.
MacOS has OSSwapHostToLittleIntXX, OSSwapLittleToHostIntXX, OSSwapHostToBigIntXX and OSSwapBigToHostIntXX in <libkern/OSByteOrder.h>.
I'm not sure if Windows has something similar, or if it even supports running on big endian machines (if you know, please tell).
My solution for achieving some portability currently entails cobbling together a "compat.h" header that defines macros for the MacOS functions and including the right headers. Something like this:
https://github.com/AgentD/squashfs-tools-ng/blob/master/incl...
This is usually my go-to-solution for working with low level on-disk or on-the-wire binary data structures that demand a specific endianness. In C I use "load/store" style functions that memcpy the data from a buffer into a struct instance and do the endian swapping (or reverse for the store). The copying is also necessary because the struct in the buffer may not have proper alignment.
Technically, the giant macro of doom in the article takes care of all of this as well. But unlike the article, I would very much not recommend hacking up your own stuff if there are systems libraries readily available that take care of doing the same thing in an efficient manner.
In C++ code, all of this can of course be neatly stowed away in a special class with overloaded operators that transparently takes care of everything and "decays" into a single integer and exactly the above code after compilation, but is IMO somewhat cleaner to read and adds much needed type safety.
Please don't do that. Use battle-tested low-level routines. Unless your USP is "our software swaps bytes faster than the competition", you should not spend brain power on that.
Boost provides boost::endian which allows converting between native and big or little, which just does the right thing on all architectures and compilers and compiles down to a no-op or bswap instruction instruction. It's much better than writing (and testing!) your own giant pile macros and ifdefs to detect the compiler/architecture/OS, include the correct includes, and perform the correct conversions in the correct places.
static uint32_t load32_le(const uint8_t s[4])
{
return (uint32_t)s[0]
| ((uint32_t)s[1] << 8)
| ((uint32_t)s[2] << 16)
| ((uint32_t)s[3] << 24);
}
I start with unsigned char to begin with (well `uint8_t` to be precise, which has the advantage of not compiling at all if you happen to use a DSP that uses 32-bit chars). Then I convert those chars to unsigned 32-bit integers. Only then do I shift them. There is no need to mask anything here.Note that modern compilers translate this whole thing into a single unaligned load operation. Even better, I've noticed that using a macro instead of a function tends to make performance worse with modern compilers.
(uint32_t)x[0] << 24 | ...
Of course, this requires that x[0] be unsigned.Or just use the endian.h / sys/endian.h routines, which do the right thing (be32dec / be32enc / whatever). memcpy+swap is fine, and easier to get right than the author's giant expressions, but you might as well use the named routines that do exactly what you want already.
You're thinking of std::bit_cast. std::launder solves a different, much more obscure problem: https://miyuki.github.io/2016/10/21/std-launder.html
if constexpr (std::endian::native == std::endian::big) {
std::cout << "big-endian" << '\n';
}
else if constexpr (std::endian::native == std::endian::little) {
std::cout << "little-endian" << '\n';
}
else {
std::cout << "mixed-endian" << '\n';
}
Doesn't solve everything, but it's saner even if what you're writing is C-style low-level code.Just out of curiosity, I would be interested in learning why so many CPUs today are little-endian. Is it because it is cheaper / more efficient for processor implementations or is it because “the others do it, so we do it the same way”?
It simplifies certain instructions internally. Practically everything is little endian because x86 won.
> And if you think about a serial machine, you have to process all the addresses and data one-bit at a time, and the rational way to do that is: low-bit to high-bit because that’s the way that carry would propagate. So it means that [in] the jump instruction itself, the way the 14-bit address would be put in a serial machine is bit-backwards, as you look at it, because that’s the way you’d want to process it. Well, we were gonna built a byte-parallel machine, not bit-serial and our compromise (in the spirit of the customer and just for him), we put the bytes in backwards. We put the low- byte [first] and then the high-byte. This has since been dubbed “Little Endian” format and it’s sort of contrary to what you’d think would be natural. Well, we did it for Datapoint. As you’ll see, they never did use the [8008] chip and so it was in some sense “a mistake”, but that [Little Endian format] has lived on to the 8080 and 8086 and [is] one of the marks of this family.
Fascism vs. not-fascism, Stalinist Communism vs. Western Capitalism, Islamism vs. liberal democracy... I’m not sure “the existence of war around a divide in ideas proves that neither sides ideas are correct” is a particularly comfortable maxim to consider the ramifications of.
Well, sure, that it’s a trivial idea pretty much inherently means either that neither is right or (and this is very much not an exclusive or) being right doesn’t matter.
The problem with real cases is that people inside the conflict don’t believe the idea is trivial (conversely, to people outside rhe conflict—or caught in the middle—even the conflicts we think of as about foundational ideas seem like trivial or irrelevant differences.)
Of all those reasons, the only one I can make sense of is the "I can’t transparently widen fields after the fact!", and that one is way too niche to explain anything.
Another example: memory representation of pixels in GPUs which are swizzled to make computations efficient
There's no reason to, as there's no reason not to. It's basically irrelevant.
If carrier passing is so important, why can't you just mirror your transistors and operate on the same wires, but on the opposite order? Well, you can, and it's trivial. (And, by the way, carrier passing isn't important. High performances ALU pass carrier only though blocks, that can appear anywhere. And the wiring of those isn't even planar, so how you arrange them isn't a showstopper.)
Would it hurt anyone to define this undefined behavior and do exactly what the source code says?
But to answer the actual question: For C++20, integer types were revisited. It is now (finally) guaranteed that signed integers are two's complement, along with a list of other changes. See http://www.open-std.org/jtc1/sc22/wg21/docs/papers/2018/p090... also for how the committee voted on the individual issues.
Note in particular:
> The main change between [P0907r0] and the subsequent revision is to maintain undefined behavior when signed integer overflow occurs, instead of defining wrapping behavior. This direction was motivated by:
> - Performance concerns, whereby defining the behavior prevents optimizers from assuming that overflow never occurs;
> - Implementation leeway for tools such as sanitizers;
> - Data from Google suggesting that over 90% of all overflow is a bug, and defining wrapping behavior would not have solved the bug.
So yes, the committee very recently revisited this specific issue, and re-affirmed that signed integer overflow should be UB.
> Data from Google suggesting that over 90% of all overflow is a bug, and defining wrapping behavior would not have solved the bug.
Of all overflow? Including unsigned integers where the behavior is defined?