An 8 wide front-end for ARM64 needs to run a decoder every 4 bytes (so 8 in total), then combine the results. It's an entirely trivial affair.
For RISC-V, without any macro-op fusion, you first need a unit every 2-bytes that determines the instruction length using the first 2 bits (16 of these), then you need a very wide mux to expand out any compressed instructions (the last one can come from 7 or 8 different offsets) so you have 8x 32b instructions, then you need your 8 decoders. It's looking much more unpleasant because of that very wide muxing (8x 8:1 32-bit muxes is not a pleasant sight). Alternatively you could do it the brute force way and have 16 decoders followed by a 16->8 reduction layer (ignoring the decode results that were at bad offsets). The designer is going to have a hard time choosing between those two bad options (personally I would go with the latter I think).
Now let's throw in macro-op fusion, suddenly rather than have two possible lengths (16b or 32b) you can have 16b+16b, 32b+32b, 16b+32b, 32b+16b. Complexity just explodes. If anybody really tried to implement a 8 wide RISC-V front-end with pervasive macro-op fusion then it would be a truly nasty thing.