Endiannes and Polymorphism in Rust
thecodedmessage.com
thecodedmessage.com
https://gitlab.com/xmdr/ananke/-/blob/master/drafts/example....
So far, progress has been a bit slow. I still don’t have a working interpreter, not to mention compiler, because I find myself tweaking the syntax a lot. Perhaps later in december it will have finalized enough for me to dare write it out.
Endianness in bitfields is not straightforward though.
[0] https://www.adacore.com/gems/gem-140-bridging-the-endianness...
Adding type annotations encourages a coding style that makes many conversions at every point of use, instead of only once at the boundary.
- IBM long numbered bits from the most significant starting with 0 (think of 360s and the like, not PCs)
- there are modern chips with registers numbered the same way (a current family of multi-core ARM chips I'm currently working with number their GPIOs this way too)
So when people are choosing to write things to wire, big endian seems quite natural because what is natural in radix 10 is equally natural in radix 16 and equally natural in radix 2.
Performance. Byte-swapping "everything" isn't free. And there's already precedence for it, after all most binary formats use multiples of 8 bits for the same reason when storing numbers etc.
I am only using big endian for variable-length integers: https://www.rfc-editor.org/rfc/rfc8794.html#name-variable-si...
In that particular use case, allows to find length of the integer by only reading the first byte. This feature helps for use cases when you wanna skip fields instead of de-serializing them.
And, it doesn’t screw up bit patterns. A variable-length integer 0x808080 gonna have three 0x80 bytes somewhere in the serialized stream. This feature sometimes helps debugging things. A little endian equivalent would have them bit-shifted by the size of header.
For fixed-length integers, I only using little-endian in my file formats.
Little endian has too many advantages to be ignored, which is why all popular architectures and also newer protocols are now little endian.
I did a write-up about it here: https://www.technicalsourcery.net/posts/on-endianness/
The biggest mistake these days is protocols that use big endian as the serialization order even though most data gets processed on little endian systems.
# Little Endian Machine 87654321 87 65 43 21
# Big Endian Machine 87654321 21 43 65 87
# Big Endian Machine, word-wise 87654321 12 34 56 78
In a world where endianness concerns only words, there would only be 12345678 and 87654321, but not the word-wise mixing as in 21 43 65 87. I suppose this is because the individual adressable unit of memory is a byte, and byte streams can be interpreted in different ways (as int8, int16 etc)? And because bytes are usually opaque to the programmer into how they lay out their bits internally to represent a uint8?
It makes sense, but without a lot of context around it, it is very confusing initially.
This is all internally consistent. However, if you said on your big endian machine “write the 16 bit number 0xabcd to main memory” it would store “0xab 0xcd”. Then you say “write those 16 bits to I/O and OS I/O routines don’t understand the structure you’ve assigned to memory so they just copy that to the wire. This leaves you with “0xab 0xcd” on the wire. Then your little endian machine says “read 16 bytes from the wire” and gets “0xab 0xcd”. But then you’re smart and you go “I was expecting a 16bit number here - cpu read this as a 16 bit number” and the CPU goes “main memory stores little endian 16bit numbers so I see 16 bit number 0xcdab” (because 0xcd is the least significant byte”.
The smallest addressable unit of CPUs you encounter is bytes whatever size the CPU vendor chooses that is in terms of bits. So a CPU would never read a little endian or big endian byte because it’s simpler to just pump the bits directly instead of swizzling them (easier to debug, analyze etc).
You could design a CPU that stored numbers in bit-order little endian or big endian but you’d basically be incompatible with all software out there and no language provides support for dealing with that because you’d be very niche.
When dealing with low level hardware protocols bit endianness does become a thing.
Sure the concept is easy, but have you ever met someone who mixes left and right up? It can be quite catastrophic in traffic
There's no right, there's only left and the other left.