Processors basically do this: read instructions, execute them out of order to fully use cpu resources, put the instructions back in order along with their results, and then write results to memory. This way the “final state” follows the order of instructions.
The “put back in order” step is handled by the re order buffer. It is basically an online sort algorithm.
Fully sorting instructions after execution is simple and correct, but in extreme cases may be suboptimal; perhaps it doesn’t matter if the final state is exactly in order or just close to that. I believe the article is saying the M1 may just validate states as they are written rather than fully reordering.
Basically the whole, "put it back in order" bits only have to happen around memory barriers because the visibility of writes isn't guaranteed otherwise.
Whether this matters given large write combining/writeback buffering that queues up early writes until the later writes have completed is one of those decade+ long arguments.
>… The M1 seems to use something other than an entirely conventional reorder buffer, which complicates measurements a bit. So these may or may not be accurate. (This paragraph previously said "it seems to use something along the lines of a validation buffer". I think the VB hypothesis has since been disproven. Various attempts to measure ROB size have yielded values 623, 853, and 2295 (see the previous link). My uninformed hypothesis is that this may imply a kind of distributed reorder buffer, where only structures that need to know about a given operation track them, and/or some kind of out-of-order retirement.)
https://github.com/dougallj/applecpu/commit/dc3c220f58f428b5...