* Function calls as the primary built-in control flow mechanism. This is so obviously better its insane, this should be the base unit of control flow in an ISA. No inventing incompatible "conventions" for calls, no manually saving & restoring registers (or need to eke out optimizations by avoiding using registers), no accidentally trampling on external state. Called functions receive parameters in a consistent location and returned values are output in a consistent location, other incidental state not part of the parameters is inaccessible, upon return the caller's state is automatically restored as if it had just executed a 1-cycle cpu instruction. All for an overhead of 1 cycle.
* Instructions and functions can have multiple return values. No stupid out params just to return both an error and value or a value that is bigger than one word. CPU instructions use multiple return values for things that obviously need them instead of overloading stateful flags (eww) or interrupts (heave).
* Unified address space & address translation pushed down to the memory controller. This solves so many problems it's not even funny. Cached writes to memory can unblock as soon as it hits L1, reading from uninitialized memory gives you a pre-zeroed page -- instantly -- not in 300 cycles, the previous two means that it's possible to allocate + write + read + deallocate a memory page where it never actually touches main memory and is entirely served by cpu cache hierarchy, ejects expensive TLB address translation out of the hot path of memory accesses instructions, memory protection-based access control can be done in parallel with the fetch instead of the fetch being delayed until translation is calculated. (A process should not assume its the only thing that exists and that its starting address is always the same, ASLR should be on by default. Arguing that per-program virtual memory is a security feature is like arguing that NAT is a security feature. Access and addressing can be separated securely.)
* Machine code is basically a directly encoded form of Static Single Assignment (SSA) which is typically a compiler Intermediate Representation (IR) middle step in the compilation process, but the Mill consumes SSA directly which means that a lot of data flow information is preserved during the lowering to machine code and doesn't have to be inferred or guessed dynamically by the CPU during execution.