The x86's Decimal Adjust after Addition (DAA) instruction
righto.com
righto.com
Interestingly, the John Mashey at MIPS defended the use of BCD instructions in what I think is PA-RISC (which does have a DCOR - Decimal Correct instruction). That's because they looked at their instruction traces, saw lots of use of BCD arithmetic as the processor was targeting COBOL among other languages, and implemented that hardware support as single cycle register to register ops, and kept the overall arch of a normal RISC. That sort of diffusion test where you actually go look at instruction traces to come up with single cycle accelerators for your expected use case, but architecturally don't comingle your load store pathway with your ALU ops is sort of the heart of RISC. HP simply had different goals for their RISC chip than MIPS.
https://yarchive.net/comp/bcd_instructions.html
That argument still pops up, for instance the "javascript instruction" in ARM. There's a lot about ARM that isn't pure RISC, but I argue that instruction is. They looked at instruction traces and found a single cycle (dispatch at least, it might be pipelined) ALU op that could accelerate the workloads that were important to them. In this case converting float to integer using x86 rounding modes rather than whatever modes are in the config register because Javascript expects x86 rounding which is different than default ARM rounding and switching back and forth is expensive. That instruction and the process used to come up with it are about as RISC as you can get, IMO.
</semi_unrelated_rant>
In an alternate universe where the IEEE-754r decimal formats were in the original IEEE-754 spec, presumably most of these decimal integer operations would be performed in the FPU. If they're common enough, a handful of specialized fast-path instructions could be implemented for cases where it's known a priori that the exponents of the arguments match.
Though, as sad is it is that we still use gallons in the US, it is handy that all of the subdivisions of a gallon down to the ounce are powers of two.
Alternatively, the Babylonians were more advanced than we generally give them credit, as their base-60 number system isn't intuitive, and yet has lots of small factors for handy fractions. Any fraction storable in a finite number of digits in base-2, base-10, or base-12 (including any integer power of 1/3) is storable in a finite number of digits in base-60. So, a base-60 calculation would result in fewer display issues due to rounding for common calculations in people's everyday lives. Storing one base-60 digit in 6 bits isn't terribly wasteful, though more wasteful than storing 3 base-10 digits in 10 bits.
add al, 90h
daa
add al, 40h
daa
I believe it faded into obscurity once the shorter 5-byte sequence using DAS instead was discovered: cmp al, 0ah
sbb al, 69h
dasCuriously, the article discusses design goals for microcomputers at a high level. The hex to ASCII algorithm was a random example of the gap between compilers and machine code. Allison had done a bunch of work with Knuth on decimal arithmetic, so he probably had this example at his fingertips.
If the integer is a fixed-width data type, say a 16-bit integer, a common algorithm was repeatedly subtracting 10000, 1000, 100, 10, etc., from the number. Not very efficient in terms of algorithm, but it can be implemented in a few lines of fast assembly. Still, if the number is frequently updated (perhaps a spreadsheet, or an on-screen score in a game), the fastest solution was storing it in BCD and avoiding the conversion. So native support of BCD arithmetic was not just a legacy due to desk calculators or compatibility, but also necessity.
Earlier discrete logic circuits and calculators used BCD for the same reason. Conversion from BCD to readable formats like a 7-segment display or ASCII output was simple and only requires simple combinational logic.
It's fortunate that computers are advanced enough today that we usually don't need to worry about the overhead of printf("%d", integer) anymore.
If you get it right, they spend quite some time trying to track down why their code is misbehaving. A bit of misdirection (e.g 'oh, that target machine was giving problems yesterday, maybe the hardware is failing') can really help :)
Separate BCD instruction would require to duplicate all instructions that are "compatible" with DAA, and there might not have been enough room in the instruction set for those.
The ABCD -(An),-(An) form is presumably for big-endian multi-byte BCD values, as the X flag participates.
> Even Intel's x86 processors are moving away from the DAA instruction; it generates an invalid opcode exception in x86-64 mode.
That's the first I've heard of that so I'm guessing no one complained.
0. https://github.com/rvalles/optromloader/
1. https://github.com/rvalles/optromloader/blob/master/optromlo...
It’s not that bad conceptually but in practice making sure all the flags are set right by the other instructions so DAA would work is a bit of a pain.
(As an aside, Google gives absolutely horrible results if you search for "8086 DAA instruction timing". At least half the ones on the first page seem to be for the 8085, there's also this article itself and the discussion on HN, neither of which have the information.)
The 8086 implements the instruction in microcode, but like all the arithmetic operations, the interesting stuff happens in the ALU. In particular, the ALU hardware has a DAA instruction that figures out what correction value is needed and adjusts the flags appropriately. Another microcode instruction performs the addition. So the microcode is just four instructions long.
Presumably AMD wanted to specify only the future-useful subset of the architecture to be usable in long mode so that in the future when people started shipping long mode only CPUs[1] they wouldn't have to bother with the junk.
[1] Which never happened, presumably because the transistor overhead of supporting real and 32 bit mode on the monster cores of the modern world was negligible. But you could totally ship a successful processor today that booted in long mode and lacked the ability to transition to real mode.
But you could totally ship a successful processor today that booted in long mode and lacked the ability to transition to real mode
I don't think so. I certainly wouldn't want one, and neither would a lot of other people. Intel learned that lesson twice --- once with the 80376, and again when it started making x86 smartphone SoCs. No one really wants x86 without the rest of the PC architecture, since then they might as well use something like ARM or even RISC-V instead.
1. What software would you want to run that used 286-style segmentation in a 64 bit space? That model had been abandoned long before AMD wrote the x86_64 spec.
2. What 386 protected mode (or unreal mode too, I guess) software are you needing to run such that you wouldn't buy a 64-bit-only CPU? Again, this has been entirely abandoned (with the sole exception of the SMP boot mechanism, which is defined to start in real mode) on modern systems, even at the firmware level.
> Fun fact: you can count to hexadecimal on the fingers of one hand.
> The way it works is that each of your four fingers has four well-defined places on them: the three joints, plus the tip. You use your thumb to point at one of these. The first finger represents 1-4, the second 5-8, the third 9-12, the fourth 13-16. You count up and down by moving your thumb.
(...)
> As a special bonus: if you point to the fleshy pads of the fingers instead of the joints, you can use the same system for base 12.
Please teach new programmers the number type stack first.
Also I think your answer is unnecessarily nitpicky. Sentence 'floats can't represent "simple" values like 0.1 (...)' is perfectly clear to everyone on HN, I hope. In fact, I disagree that changing "values" here to a more precise mathematical term would improve anything.
That's an interesting philosophical question. Does M_PI represent a rational number?
(M_PI isn't mentioned in my copy of the C spec, unless I just can't find it.)
But we have great native support for approximating real numbers instead. (Because computing is designed first and foremost to make numerical methods efficient.)
The exist so that trigonometry and logarithms work properly in numerical methods. (A vastly more important problem than getting '0.1' to print pretty on the screen.)
There is no such thing as a 'value' in mathematics, and fixating on 0.1 in particular is pointless.
tldr - everyone should learn the math behind this first.
The one thing you'd lose would be two's complement trickery. Addition and subtraction in BCD actually are fundamentally different algorithms, so that's something to worry about. But then signedness mismatches are themselves a historical footgun, so you could spin this away too, I think.
But... beyond that this could absolutely be done. And it's also true that a lot of traditional bugs with FPU/brain precision mismatch would be fixed in the process.
I don't know that it's an architecture I'd pick, but it's not something to reject out of hand either. It's worth discussing and not downvoting.
The 8008 came from the Datapoint 2200, a desktop computer built from a board of TTL chips. Datapoint asked Intel and Texas Instruments if they could replace this TTL processor with a single chip. Texas Instruments created the TMX 1795, the first 8-bit microprocessor. Intel created the 8008 shortly after. Both chips copied the instruction set and architecture of the 2200. Datapoint didn't like either chip and stuck with TTL. Texas Instruments couldn't find a customer for the TMX 1795 and abandoned it. Intel, on the other hand, marketed the 8008 as a general-purpose microprocessor, essentially creating the microprocessor industry.
The 8008 was cleaned up to create the popular 8080. Intel started the iAPX 432 as their flagship "micromainframe" processor, but it was extremely ambitious and fell behind the schedule. Intel created a stopgap 16-bit processor so they'd have something to sell until the iAPX 432 was ready: this was the 8086. They designed the 8086 to be able to run 8080 assembly after translation, so it inherited a lot from the 8080. As for the iAPX 432, it was late, a failure, and is mostly forgotten.
The point of this is that the x86 is descended from the Datapoint 2200 and has no connection to the Intel 4004.
I wouldn't feel safe about my posts having them hosted there...