Possible reasons for 8-bit bytes
jvns.ca
jvns.ca
Having 6-bit bytes on a CDC was a terrific PITA! The byte size was a tradeoffs between saving MONEY (RAM) and the hassle of shift codes (070) used to get uppercase letters and rare symbols! Once semiconductor memory began to be available (2M words of 'ECS' - "extended core storage" - actually semiconductor memory - was added to our 1M byte memory in ~1978) computer architects could afford to burn the extra 2 bits in every word to make programming easier...
At about the same time microprocessors like the 8008 were starting to take off (1975). If the basic instruction could not support a 0-100 value it would be virtually useless! There was only 1 microprocessor that DID NOT use the 8-bit byte and that was the 12-bit intersil 6100 which copied the pdp-8 instruction set!
Also the invention of double precision floating point made 32-bit floating point okay. From the 40s till the 70s the most critical decision in computer architecture was the size of the floating point word: 36, 48, 52, 60 bits ... But 32 is clearly inadequate. But the idea that you could have a second larger floating point fpu that handled 32 AND 64-bit words made 32-bit floating point acceptable..
Also in the early 1970s text processing took off, partly from the invention of ASCII (1963), partly from 8-bit microprocessors, partly from a little known OS whose fundamental idea was that characters should be the only unit of I/O (Unix -father of Linux).
So why do we have 8-bit bytes? Thank you, Gordon Moore!
The 180 introduced a somewhat less wild, certainly more C friendly, 64-bit arch revision.
There was only 1 microprocessor that DID NOT use the 8-bit byte
Toshiba had a 12-bit single chip processor at one time I'm pretty sure you could make a similar claim about. More of a microcontroller for automotive that general purpose processor, tho.
From the beginning right up to the birth of microprocessors computer architects were still experimenting around these tradeoffs - around the late 70s it all jelled into 8 bit bytes and power of 2 word sizes
I think that 8 bits is a nice round number is you move to a memory model that is an array of bytes and an array of words - making everything a power of 2 in size makes the hardware simpler (much the same reason why bcd, and decimal arithmetic have largely disapeared)
It was a DSP, memory was still relatively expensive and MPEG needed 9-bit delta frames, however DRAMS with byte parity were available so it was a price point that made sense at the time
The C compiler used 9-bit bytes, just to make pointer arithmetic simpler.
It is extremely confusing. A "word" means different things depending on architecture, programming language, and probably a few other things that make up the context.
For example, in the x86 assembly world you usually call an (usually aligned) 16 bit value a "word", a 32 bit value a "double word", and a 64 bit value a "quad word", all because the original 8086 had 16 bit wide registers that could be alternatively be used as two 8 bit registers.
In other architectures or more generally other contexts, a word often relates to "how wide the bus is". But even then some of the old nomenclature bleeds into newer things that may have gotten wider busses, and it's not even always clear what bus is being referred to.
Basically, unless you operate within a certain shared context, the word "word" is fuzzy and ambiguous. Maybe not quite as much as the word "object", but it's definitely up there somewhere.
Because of backwards compatibility, and because CPUs are far more flexible about loading more memory sizes than they traditionally have been, the meaning of "word" doesn't really matter anymore. In the context of Win32 and x86 programming, a WORD is 16 bits, a DWORD ("double word") is 32 bits, and a QWORD ("quad word") is 64 bits. That's most likely where you'll see it these days.
CPU word size for individual CPU designs much easier to define. It's not really the bus size, though it was strongly correlated until caches and bus-multipliers became a thing (For example, the PowerPC 601 has 32bit words (and 32bit registers) despite having a 64bit bus)
Word size is the native operation size. The size of integer you should use in your code for the best performance. Which is why the 68000 has a 16bit word size despite having 32bit registers, because instructions for 32bit operations took a few cycles longer.
And AArch64 also uses qword to refer to 128 bit data types.
I assume you will find the same issue on the Window NT ports to MIPS and PowerPC.
Where you see dword as 64bit on those platforms is their documentation and assembly syntax. Most software in higher level languages (that aren't windows) avoided naming their type alias as WORD, DWORD and QWORD.
The other interesting thing about EBCDIC was that character codes with A–F in the lower nibble of the code were generally unused or reserved for uncommon characters. This restriction didn’t apply to the upper nibble (most notably the digits were encoded as F0–F9).
Modern z/Architecture (64-bit 360/370/390 successor) has instructions to convert numbers between EBCDIC, ASCII, Unicode, packed decimal, zoned decimal, and binary, and to convert Unicode character strings between UTF-8, UTF-16, and UTF-32.
In addition to conversion instructions, z/Architecture has full arithmetic support for binary and decimal integer operations and binary, decimal, and hexadecimal floating-point.
POWER6 and later instruction sets also include full support for decimal floating-point.
https://www.ibm.com/docs/en/cics-ts/5.4?topic=basics-24-bit-...
Wikipedia has a section on the design considerations of ASCII: https://en.wikipedia.org/wiki/ASCII#Design_considerations
Also the ones I used to program (PDP-6/10/20) had an 18-bit address space, which you may note is a CONS cell. In fact the PDP-6 (first installed in 1964) was designed with LISP in mind and several of its common instructions were LISP primitives (like CAR and CDR).
The formal definition of a byte is that it's the smallest addressable unit of memory. Think of a memory as a linear string of bits. A memory address points to a specific group of bits (say, 8 of them). If you add 1 to the address, the new address points to the group of bits immediately after the first group. The size of those bit groups is 1 byte.
In modern usage, "byte" has come to mean "a group of 8 bits", even in situations where there is no memory addressing. This is due to the overwhelming dominance of systems with 8-bit bytes. Another term for a group of 8 bits is "octet", which is used in e.g. the TCP standard.
Words are a bit fuzzier. One way to think of a word is that it's the largest number of bits acted on in a single operation without any special handling. The word size is typically the size of a CPU register or memory bus. x86 is a little weird with its register addressing, but if you look at an ARM Cortex-M you will see that its general-purpose CPU registers are 32 bits wide. There are instructions for working on smaller or larger units of data, but if you just do a generic MOV, LDR (load), or ADD instruction, you will act on 32 register bits. This is what it means for 32 bits to be the "natural" unit of data. So we say that an ARM Cortex-M is a 32-bit CPU, even though there are a few instructions that modify 64 bits (two registers) at once.
Some of the fuzziness in the definition comes from the fact that the sizes of the CPU registers, address space, and physical address bus can all be different. The original AMD64 CPUs had 64-bit registers, implemented a 48-bit address space, and brought out 40 address lines. x86-64 CPUs now have 256-bit SIMD instructions. "32-bit" and "64-bit" were also used as marketing terms, with the definitions stretched accordingly.
What it comes down to is that "word" is a very old term that is no longer quite as useful for describing CPUs. But memories also have word sizes, and here there is a concrete definition. The word size of a memory is the number of bits you can read or write at once -- that is, the number of data lines brought out from the memory IC.
(Note that a memory "word" is technically also a "byte" from the memory's point of view -- it's both the natural unit of data and the smallest addressable unit of data. CPU bytes are split out from the memory word by the memory bus or the CPU itself. Since computers are all about running software, we take the CPU's perspective when talking about byte size.)
In the ARM ARM, a word is 32 bits, because that was the Arm’s original word size. Other sizes are half words and double words.
It is a very context-sensitive term.
Ah, yes. That terminology is still used in the Windows registry, although Windows 10 seems to be limited to DWORD and QWORD. Probably dates back to the 286 or earlier. :-)
There were even bit addressable computers, but it didn't catch on :)
If it wasn't for text, there would be nothing "natural" about an 8-bit byte (but powers-of-two are natural in binary computers).
Cortex-M has bit addressability as a vendor option. They're mapped onto virtual bytes in a region of memory so the core doesn't have to handle them in a special way.
I found this LLVM forum thread discussing it interesting: https://groups.google.com/g/llvm-dev/c/s2yuELeQMA8
And related PR https://reviews.llvm.org/D61725
But what would we lose if we just got rid of the notion of bytes and just let every bit be addressable?
To start, we'd still be able to fit the entire address space into a 64-bit pointer. The maximum address space would merely be reduced from 16 exabytes to 2 exabytes.
I presume there's some efficiency reason why we can't address bits in the first place. How much does that still apply? I admit, I'd just rather live in a world where I don't have to think about alignment or padding ever again. :P
The article touches on this but having your addressable unit fit a single character is incredibly convenient. If you are manipulating text you will never worry about single bits in isolation. Ditto for mathematical operations, do you really have a need for numbers less than 255? It is a lot more convenient to think about memory locations as some reasonable unit that covers 99% of your computing use cases.
The Intel iAPX 432 did use bit-aligned instructions:
> https://en.wikipedia.org/w/index.php?title=Intel_iAPX_432&ol...
When ASCII was released in 1963[3], integrated circuits didn't exist.
Computing hardware was a build it from available, off the shelf parts endeavor.
Easy to repurpose readily available 4-wire[2] to operate 4 wire relay. (2 relays -> 8 bits).
-----
[0] http://ed-thelen.org/comp-hist/Reckoners-ch-4.html
[1] http://quadibloc.com/comp/cardint.htm
[2] : https://en.wikipedia.org/wiki/Four-wire_circuit
[3] : https://www.historyofinformation.com/detail.php?id=803
4 bits fits into a 16 pin DIP well and cascades well. It is no coincidence that the 4004 operates on 4 bit units (BCD was also a factor). The 8008 needed an 18-pin DIP and was still far too constrained.
So, your choice of unit is likely a multiple of 4.
The 8080 was indeed a 40-pin DIP. However, it still needed something like 3 other support chip to demultiplex everything.
It wasn't until the 8085 that you didn't need so much support circuitry.
People have invented new ones for ML (eg the Brain Float16), but even then some people have demonstrated training on int8 or even int4.
There isn't even consensus on how to map the state space onto the numberline - is linear (as in ints) or exponential (as in floats) better? Perhaps some entirely new mapping?
And obviously there could be different optimal numbersystems for different ML applications or different phases of training or inference.
I haven't thought about the possible values for that "class" field in a long time. It does seem interesting and extremely weird from today's perspective.
4 bits for 0..9 as in BCD is probably more convincing.
The justification was that on this particular platform, char* pointers were differently structured to int* pointers, because char* pointers had to reference a single byte and int* pointers didn't.
EDIT - I appear to have cut this story short. See my response to "wyldfire" for the rest. Sorry for causing confusion.
I must be missing some context or you have a typo. Probably most architectures I've ever worked with had `int *` refer to a register/word-sized value, and I've not yet worked with an architecture that had single-byte registers.
Decades ago I worked on a codebase that used void * everywhere and rampant casting of pointer types to and fro. It was a total nightmare - the compiler was completely out of the loop and runtime was the only place to find your bugs.
C/C++ compilers often by default add extra bytes if necessary to make sure everything's aligned. So if you have struct X { int a; char b; int c; char d; } and struct Y { int a; int b; char c; char d; } actually X takes up more memory than Y, because X needs 6 extra bytes to align the int fields to 32-bit boundaries (or 14 bytes to align to a 64-bit boundary) while Y only needs 2 bytes (or 6 bytes for 64-bit).
Meaning you can sometimes save significant amounts of memory in a C/C++ program by re-ordering struct fields [1].
The 'memory bus' is not architectural. Different microarchitectures implement things differently, but most high-performance microarchitectures these days have relatively efficient misaligned accesses.
Whether we classify the issue as "architectural" (whatever that means) is beside the point. Alignment has real effects on performance, and being aware of those effects is practically useful for working programmers.
I do agree with you about one thing: The unaligned access penalty is probably less on the x86 CPU's of the 2020's than it was, say, on a 486 from the mid-1990's.
But I'm talking about retrieving the 9th to 16th bit of a word, which is a little different. x86 does this just fine/quickly, because bytes are addressable.
Yes. Lots of little microcontrollers and older big machines have this "feature" and C compilers fix it for you.
There are nightmarish microcontrollers with Harvard architectures and standards-compliant C compilers that fix this up all behind the scenes for you. E.g. the 8051 is ubiquitous, and it has a Harvard architecture: there are separate buses/instructions to access program memory and normal data memory. The program memory is only word addressable, and the data memory is byte addressable.
So, a "pointer" in many C environments for 8051 says what bus the data is on and stashes in other bits what the byte address is, if applicable. And dereferencing the pointer involves a whole lot of conditional operations.
Then there's things like the PDP-10, where there's hardware support for doing fancy things with byte pointers, but the pointers still have a different format than word pointers (e.g. they stash the byte offset in the high bits, not the low bits).
The C standards makes relatively few demands upon pointers so that you can do interesting things if necessary for an architecture.
Digging deeper, this particular microcontroller was tuned for accessing 32 bits at a time. Accessing individual bytes needed extra bit-shuffling code to be added by the compiler.
Raymond's blog on the Alpha https://devblogs.microsoft.com/oldnewthing/20170816-00/?p=96...
But -- they are different. Architectures where they're treated the same are probably the exception. Depending on what you mean by "very different" - most architectures will emit different code for byte access versus word access.
Accessing an 8 bit byte from a pointer, the compiler would insert assembly code into the generated object code. The "normal" part of the pointer would be read, loading four characters into a 32 bit register. Two extra bits were squirreled away somewhere in the pointer and these would feed into a shift instruction so the requested byte would appear in the lowest-significant 8 bits of the register. Finally, an AND instruction would clear the top 24 bits.
Wild guess, but the OP might be talking about the Intel 8051. Single-byte registers, and depending on the C compiler (and there are a few of them) 8-bit int* pointing to the first 128/256 bytes of memory, but up to 64K of (much slower) memory is supported in different memory spaces with different instructions and a 16-bit register called DPTR (and some implementations have 2 DPTR registers). C support for these additional spaces is mostly via compiler extensions analogous but different from the old 8086 NEAR and FAR pointers. I'm obviously greatly simplifying and leaving out a ton of details.
Oh, yeah...on 8051 you need to support bit addressing as well, at least for the 16 bytes from 20h to 2Fh. It's an odd chip.
It’s embarrassing you’re even asking this. Obviously we wouldn’t use 11 bits per byte because that’s nonsense.
Why don’t we use 4 bits? Because that’s like saying you could only write 2 digits for decimal numbers and never more. That’s not enough digits for practical use. I don’t want to have to write every single number as 10 pairs of two digits when I could just use numbers of a practical length.
How about the PDP-11? Both the 360 and the PDP-11 heavily influenced microprocessor architecture.
https://en.wikipedia.org/wiki/8-N-1
* words
If you add a parity bit, (say '8E1'), that makes the total 9 bits. When you wrap that in the mark-space pair that makes a total of 11 bits.
Then if you have two stop-bits, as in (say '8O2') you have 8 data bits, one parity bit, one start-bit, and two stop-bits. That makes 12 bits altogether for the single data byte sent serially.
For instance a binary '3' is inexact, but a BCD '3' is exact.
That means that currency transactions tend to work better as long as you can restrict the number of significant digits to less than the number of digits used in the BCD software.
I used to use North Star BASIC back in the early 80s. North Star BASIC's default number of digits was 8, but they also supplied 10, 12 and 14 digit BASIC along with the default. You used whichever one was required for accuracy in your application.
Huh? Do you mean '0.3'?
Having 'cut my teeth' on North Star BASIC's BCD floating point representation of numbers ( "1 + 2 = 3"), with all numbers, even integers, stored as 5-byte 8-digit BCD floating-point numbers internally. It used to drive me nuts to use Lawrence Livermore Laboratories' BASIC which used binary floating-point. You'd get things like "1 + 2 = 2.99999998" (or similar, I'm working from a 40-years ago memory here).
LATER UPDATE: Sorry, I omitted the words 'floating point'. No wonder you said 'huh?'. It should have read 'Binary floating point '3' is inexact'.
There is one spot where binary floating point has trouble compared to decimal, and that's dividing by powers of 5 (or 10). If you divide by powers of 2 they both do well, and if you divide by any other number they both do badly. If you use whole numbers, they both do well.
Also even if you do want decimal, you don't want BCD. You want to use an encoding that stores 3 digits per 10 bits.
If you use 12 digit BCD to store 1/3, you get .333 333 333 333 000 000
If you use the same number of bits for binary, you get .333 333 333 333 333 925
That's true for the the original x86 instruction set. IA-32 has 32 bit word size and x86-64 has... you guessed it 64.
16 and 32 bit registers are still retained for compatibility reasons (just look at the instruction set!).
Essentially, enough stuff got written assume=ing that "word" was 16 bits (and double word, quad word, etc., had their obvious relationship) that even though the term had not previously been fixed, it would break the world to let it change, even as processors with larger word sizes (in the "natural unit of data" sense) became available, then popular, then dominant.
https://www.truenorthfloatingpoint.com/problem
Floating point arithmetic has its problems.
[1] Ariane 5 ROCKET, Flight 501
[2] Vancouver Stock Exchange
[3] PATRIOT MISSILE FAILURE
[4] The sinking of the Sleipner A offshore platform
[1] https://en.wikipedia.org/wiki/Ariane_flight_V88
[2] https://en.wikipedia.org /wiki/Vancouver_Stock_Exchange#Rounding_errors_on_its_Index_price
[3] https://www-users.cse.umn.edu/~arnold/disasters/patriot.html
As per the provided link, the Patriot missile error was 24-bit fixed point arithmetic, not floating point. Granted, a fixed-point representation in tenths of a second would have fixed this particular problem, as would have using a clock frequency that's a power of 1/2 (in Hz). Though, using a base 10 representation would have prevented this rounding error, it would also have reduced the time before overflow.
I think IEEE-754r decimal floating point is a huge step forward. In particular, I think there was a huge missed opportunity in defining open spreadsheet formats that decimal floating point option wasn't introduced.
However, binary floating point rounding is irrelevant to the Patriot fixed-point bug.
It's not reasonable to expect accountants and laypeople to understand binary floating point rounding. I've seen plenty of programmers make goofy rounding errors in financial models and trading systems. I've encountered a few developers who literally believed the least significant few bits of a floating point calculation are literally non-deterministic. (The best I can tell, they thought spilling/loading x87 80-bit floats from 64-bit stack-allocated storage resulted in whatever bits were already present in the low-order bits in the x87 registers.)
floating point cannot do that, its precision is based on powers of 2 (1/2, 1/4, 1/8, and so on). For small values (in the range 0-1), there are _so many_ values represented that the powers of 2 map pretty tightly to the powers of 10. But as you repeat calculations, or get into larger values (say, in the range 1,000,000 - 1,000,001), the floating points become more sparse and errors crop up even easier.
For example, using 32 bit floating point values, each consecutive floating point in the range 1,000,000 - 1,000,001 is 0.0625 away from the next.
jshell> Math.ulp((float)1_000_000)
$5 ==> 0.0625BCD is a completely different thing, instead of tightly encoding an integer you encode it digit by digit wasting some fraction of a bit each time but make conversion to and from decimal numbers much easier. But there is no advantage compared to a base ten fixed or floating point representation when it comes to representable numbers.
IEEE 754 also supports encoding decimal integers as binary, by converting the entire decimal integer to a single binary integer (i.e., not by storing each decimal digit as a separate binary number).
[1] https://en.wikipedia.org/wiki/Densely_packed_decimal
[2] https://files.openpower.foundation/s/dAYSdGzTfW4j2r2#page=22...
BCD is attractive to human beings programming computers to duplicate algorithms (generally financial ones) intended for other human beings to execute using arabic numerals. But it's not any more "accurate" (per transistor, it's actually less accurate due to the overhead).
The only real disadvantage for BCD is its not as quick as Floating point arithmetic, or bit swapping data types, but with todays faster processors, for most people I'd say the slower speed of BCD is a non issue.
Throw in other hardware issues, like bit swapping in non ECC memory and the chances of error's accumulate if not using BCD.
There are two reasons for BCD, (1) to avoid the cost of division for conversion to human readable representation as implied in the OP, (2) when used to represent floating point, to avoid "odd" representations in the human format resulting from the conversion (like 1/10 not shown as 0.1). (2) implies floating point.
Eben in floating point represented using BCD you'd have rounding errors when doing number calculations, that's independent of the conversion to human readable formats; so I don't see any reason to think that BCD would have avoided any disasters unless humans were involved. BCD or not is all about talking to humans, not to physics.
Parent didn't say it's a logical necessity, as in "avoid floating point ==> MUST use BCD".
Just casually mentioned that one reason BCD got popular to sidestep such issues in floating point.
(I'm not saying that's the reason, or that it's the best such option. It might even be historically untrue that this was the reason - just saying the parent's statements can and probably should be read like that).
If they just want to side step problems with floating point rounding targetting the physical world, they need to go with integers. Choosing BCD to represent those integers makes no sense at all for that purpose. All I sense is a conflation of issues.
Also, thinking about it from a different angle, avoiding issues with the physical world is one of properly calculating so that rounding errors become no issues. Choosing integers probably helps with that more in the sense that it is making the programmer aware. Integers are still discrete and you'll have rounding issues. Higher precision can hide risks from rounding errors becoming relevant, which is why f64 is often chosen over f32. Going with an explicit resolution and range will presumably (I'm not a specialist in this area) make issues more upfront. Maybe at the risk of missing some others (like with the Ariane rocket that blew up because of a range overflow on integer numbers -- Edit: that didn't happen on the integer numbers though, but when converting to them).
A BCD number representation helps over the binary representation when humans are involved who shouldn't be surprised by the machine having different rounding than what the human is used to from base 10. And maybe historically the cost of conversion. That's all. (Pocket calculators, and finance are the only areas I'm aware of where that matters.)
PS. danbruc (https://news.ycombinator.com/item?id=35057850) says it better than me.