ARM has consistently beat x86 in performance/watt at larger node sizes since the beginning. The first Archimedes had better floating point performance without a dedicated FPU than the then market-leading Compaq 386 WITH an 80387 FPU.
A lot of the extra performance of the M1 family has nothing to do with node, but with the fact the ARM ISA is much more amenable to a lot of optimizations that allow these chips to have surreally large reordering buffer, which, in turn, keep more of the execution ports busy at any given time, resulting in a very high ICP. Less silicon used to deal with a complicated ISA also leaves more space for caches, which are easier to manage (remember the more regular instructions), putting less stress on the main memory bus (which is insanely wide here, BTW). On top of that, the M1 family has some instructions that help make JavaScript code faster.
So, assume that Intel and AMD, when they get 5nm designs, will have to use more threads and cores to extract the same level of parallelism that the M1 does with an arm (no pun intended) tied behind its back.