RISC-V also has instructions that are two bytes long. The assembler automatically uses a 2 byte instruction instead of the corresponding four byte instruction if it can.
Examples of available two byte instructions:
Using any of the 32 registers:
- move or add one register to another
- shift left by a constant 0..63
- add a constant -32..+31 to a register
- load or store a register within 0..63 register sizes above the stack pointer
Using only the eight "most popular" registers (s0..s1, a0..a5):
- subtract, AND, OR, XOR two registers
- AND with a constant -32..+31
- logical or arithmetic shift right by a constant 1..63
- load or store a 32 or 64 bit value at an offset of 0..31 operand widths above the address in a register.
The programmer (or even compiler) doesn't need to remember which operations and registers are supported by 2 byte instructions, the assembler just selects them automatically when possible.
Of course you can decrease your code size if you do know what can use a 2 byte instruction and organise your code to maximise it.
I think the motivation was something more like: there's room in the instruction encoding and it's more flexible. Like you can do things like move a value from one register to another using an `add` instruction instead of a dedicated `move` by using `dest = source + zero`. Internally I think some CPUs were doing that sort of thing anyway to simplify the datapath, and the basic point of RISC was to expose the microarch in the instruction set.
You can see CISC archs hacking in the "xor REG" instruction as 'erase dependency in OoO hardware' to get around the dependency issues I talked about. Three address doesn't have that root issue because you can always specify a destination that wasn't a source.
Additionally, it's not a case of exposing the CISC internal datapath as RISC instructions. CISC vertical microinstructions tend to be two address as well. The horizontal microinstructions can't really be said to have clear source and destination enough to be either two or three address.
Note: "jalr zero, 1b" can also be written as "j 1b", "jalr zero, 0(ra)" can be written as "ret"
`j` and `ret` are so-called "pseudo instructions" [1], not compressed instructions.
Pseudo instructions are just shortcuts used in assembly language to pretend that some common operations really "exist" without the need to type (or display) the corresponding more complex (but actually existing) instructions. `nop` is a common pseudo instruction. RISC-V has no real `nop` instruction, but, instead, the "do nothing instruction" is canonically encoded as `addi x0, x0, 0`. The programmer can write a more understandable `nop`, and the assembler will write instead the binary code equivalent to `addi x0, x0, 0`.
The compressed instruction set (a.k.a "extension C"), instead, is a subset of the full [2] instruction set, in which a restricted combinations of operands are possible. The assembly (human readable) code of the compressed instruction set looks similar to that of the full instruction set (including pseudo instructions), but they are encoded as completely different binary sequences.
[1] https://github.com/riscv/riscv-asm-manual/blob/master/riscv-...
[2] https://riscv.org/wp-content/uploads/2019/06/riscv-spec.pdf#...
X86 has them too (because of course it does)