Did private frontier models use sparse embedding and ngram first? The article claims sparse attention was copied from open weight but we can't know that. We could just as easily argue that OAI and ANT had these improvements for years and decided to slash their margins only now to stay competitive with open weight neoclouds.
Second, sparse attention is an old area of active research. Offloaded N-gram tables are the next big open weight technological leap.