Do any of the open weight models from smaller labs exist if they can't distill from the SoTA models that are throwing billions of dollars of compute into pretraining?
What does that even mean?
I'm just wondering if the smaller labs see the same velocity of advances without SOTA models to generate Terabytes of training data?