That's really cool, Marat. Thanks for the additional info on the A15 and Swift.
It's a lot of work to optimize the assembly code to each ARM variant, but glad to know that Swift will generally run the same code at the same or faster speeds as the Cortex-A8.
The 3-cycle latency on simple ALU instructions is a bummer, but fortunately I use them sparingly for computation as compared to NEON. (They're great for pointer arithmetic and computing image row strides.)
The multiple issue of an ALU + LOAD is awesome. That would definitely help some of my routines.