>According to the datasheets, the 80186 can execute the same instruction in 3 cycles, and the 80286 can do it in 2.
I've disassembled the 80186 microcode, it's very similar to the 8086 but with some enhancements. Most importantly, multiply/divide and address calculations like [BX+SI+imm] are now handled in hardware rather than microcode. The "update flags" bit has another role when combined with a jump, it means the following µ-instruction will always be executed during the next cycle (like the branch delay slot on MIPS).
Sometimes a jump without this bit set seems to be needed, for no other reason but to delay execution. ALU with an immediate operand on the '186 uses this code:
0 Q -> tmpb 0 L16 1
1 M -> tmpa 1 XI tmpa NX
2 SIGMA -> M 4 none RNI F
Here the first line can load either 8 or 16 bits into tmpb, with a conditional jump to the next line (causing a 1 cycle delay) if the operand is 16 bits. For 8 bit operands it should be three cycles total, maybe when pipelined also for 16 bits?
A more interesting example is BOUND, which compares a register with a (signed) lower and upper bound in memory, in a somewhat convoluted way:
0 R -> tmpa 6 R DD,P2
1 OPR -> tmpb 1 SUBT tmpa
2 F -> tmpc 0 UNC 9 F
3 SIGMA -> tmpa 1 SHL tmpa F
4 CR -> OPR 5 UNC INT_N
5 R -> tmpb 6 R DD,P0
6 OPR -> tmpa 1 SUBT tmpa
7 4 CF1 none
8 SIGMA -> tmpa 1 SHL tmpa F
9 SIGMA -> none 0 OF 12
10 tmpc -> F 0 CY 4
11 0 UNC 13
12 tmpc -> F 0 NCY 4
13 0 NF1 5
14 4 none RNI
Step by step, the following happens:
(0) move the register operand into tmpa, read lower bound from memory
(1) move lower bound into tmpb, set ALU to subtract tmpa-tmpb
(2) preserve old flags in tmpc; execute next line and then goto line 9
(3) move ALU output into tmpa, shift it left and update flags
(9) if overflow, goto line 12
-> (12) if carry clear, goto line 4 (-> "out of bounds" exception); restore flags
else
-> (10) if carry set, goto line 4; restore flags
(11) goto line 13
(13) if internal flag clear (should be true), goto line 5
(5) move register operand into tmpb, read upper upper bound from memory
(6) move upper bound into tmpa, subtract tmpa-tmpb
(7) toggle internal flag
(8) move ALU output into tmpa, shift left & update flags
(9-12) same as before
(13) if internal flag clear (should be false), goto line 5
(14) run next instruction
While the x86 instruction set can do both signed and unsigned comparisons, the microcode - for the 80186 at least - seems to be limited to testing carry and overflow flags individually (sign bit xor overflow means a < b for signed integers). The internal flag is used as a 1-bit "loop counter" It is normally cleared at first, except when there is a REP prefix. Executing "REP BOUND" on the 80816 will only compare against the lower bound!
Apparently keeping the microcode short was more important than efficiency, or they would have used a few more instructions instead of jumping backwards.
Starting with the '286, the REP prefix would be handled during decoding, either ignored or going to a completely different microcode address.
Also on the 80186 but no other processor, AAM with an immediate operand of zero does not cause an exception (because the check for a zero divisor is still done in microcode, and AAM doesn't do it since the only documented variant is the one dividing by 10 decimal). The result is AH=FF, AL=unchanged.