The performance ceiling from the pre-training seems fairly high and they demonstrated impressive post-training improvements from Flash 3.6 -> Flash 3.8. If they can reproduce that in this model then this can be a good model for the next year. But the question is whether they can keep this up over coming years; they missed one pretraining cycle due to internal misallocation and it costed them several months of frontier competitions, and I still don't know if they addressed this structural problem.