yes, i mostly agree with this, and rvc is especially competitive on this point
often the hardware engineers weren't being stupid. it is possible to some degree to trade off hardware speed and complexity against code size (the most obvious example is instruction alignment requirements), and architectures like arm64 (aarch64) have gone pretty far in the direction of higher speed at the expensive of larger code size, even if not quite as far as things like itanic and the original mips. if you're using such an architecture, bytecode can give you a large compactness advantage, so designers try not to use them where that counts. still, the modern world contains vastly more avr and pic16 chips than it does x86 and arm, and those architectures are not great for code density, especially for c code that tends to use a lot of ints
but thumb2 and rvc are pretty hard to beat. i think compactness-optimized stack bytecodes can do it, but only by a little. on this example (which is admittedly a toy example, but not one i cherry-picked to showcase the merits of stack machines) rvc gets to 18 bytes, though, as you can see from the thread, how to get there was far from obvious to me. the stack bytecode is probably 15 bytes, but it probably needs a procedure header declaring the number of arguments to get there, so 17 bytes. plus probably a 2-byte entry in a global procedure table, which rvc doesn't need. you can implement the bytecode interpreter so that it doesn't have any alignment requirements, while rvc requires 2-byte alignment, which you should expect to cost you about a byte every so often on average—though i'm not sure how to account for that in a case like this. one byte per subroutine? one byte per basic block? is that already included in the 18?
note, though, that for the first hours of this thread, rvc was stuck with considerably worse code size at 22 bytes. it wasn't until dzaima whipped out their clang that we realized that by adding a redundant down-counter register (and extra decrement instruction in the inner loop) we could trim it down to 18 bytes
i'm amused that you describe rx as 'an improved m68k', because my thought when i first looked it over (on january 18) was that it looked like renesas thought, 'what if we did the 80386 right?' but it's true that, for example, its addressing mode repertoire is a lot more m68k-ish. by my crude, easily gamable measures, it looks like it's one of the most popular five cpu architectures, which makes it surprising that i hadn't heard of it before this year. my notes are in the (now misnamed) file remaining-8-bit-micros.md in http://canonical.org/~kragen/sw/pavnotes2.git
with respect to your particular examples of non-complex expressions, a large fraction of `a = a + 1` and `n = n - 1` are loop counters, which can be handled with a one-byte or two-byte `loop` instruction, as the 8086 does (though its design choice to only handle down-counting toward zero reduces its usefulness for compiling idiomatic c code), and typically in `c = a + b` it can be arranged for one or two of the three operands to be on top of the stack, so it's typically about 2.5 bytes instead of 4, which is still worse than thumb but only slightly. so typically stack bytecodes with either one operation or one local-variable load or store per byte come in pretty close to equal to rvc-style instruction sets with one operation and two register operands per 16-bit parcel.
the mention of the ubiquitous pic and avr above reminds me that i should mention the other advantage of interpreters in general, including bytecode interpreters: they make harvard-architecture machines programmable. a substantial fraction of the pic and avr machines in the field have hundreds or thousands of bytes of ram, enough for a program that does very interesting things, and an interpreter makes it possible for them to run code from that ram. as stm32 and its clones have progressively squeezed out pic and avr, this has become less important, but if i'm not mistaken, when gigadevice made a risc-v version of their popular higher-frequency gd32f clones of the stm32 (the short-lived gd32vf) they made it harvard! (or quasi-harvard: like the newer avrs, you can read data from the flash with regular instructions, but as i understand the datasheet, you can't execute instructions from the sram—though https://offzone.moscow/upload/iblock/0a5/nad1d86e3ah3ayx38ue... and https://www.eevblog.com/forum/microcontrollers/risc-v-microc... say this is wrong)
some of my previous notes on this problem, in case it's a problem that interests you, are at https://dernocua.github.io/notes/c-stack-bytecode.html. in there, i took six example pieces of c code (from macpaint, the openbsd c library, the linux kernel, the implementation of php, an undergraduate cs exercise, and open-source flashlight firmware) and examined how to design a compact bytecode to encode them