But only Apple's chip has a large reordering buffer. ARM Neoverse V1 / N1 / N2 don't have it, no one else is doing it.
Apple made a bet and went very wide. I'm not 100% sure if that bet is worth the tradeoffs. I'm certain that if other companies thought that a larger reordering buffer was useful, they'd have done it.
I'll give credit to Apple for deciding that width still had places to grow. But its a very weird design. Despite all that width, Apple CPUs don't have SMT, so I'd expect that a lot of the performance is "wasted" with idle pipelines, and that SMT would really help out the design.
Like, who makes an 8-wide chip that supports only 1 thread? Apple but... no one else. IBM's 8-wide decode is on a SMT4 chip (4-threads per core).