none of these techniques except MLA are new
No one (publically) had really pushed any of these techniques far, especially not for such a big run.
the transformer was an entirely new architecture, very different step change than this
e: and alibaba
I think recomputing MLA and RMS on backprop is something few would have done.
Dispensing with tensor parallelism by kind of overlapping forward and backprop. That would not have been intuitive to me. (I do, however, make room for the possibility that I'm just not terribly good at this anymore.)
I don't know? I just think there's a lot of new takes in there.