I can't comment on concrete sizes of future models, partly because it's not decided yet. But: Pre-training for Kolibri only started in August so there's a high chance continued post-training - which we plan to do - will yield some nice checkpoints. We believe this size and sparsity allows achieving great inference throughput at reasonable levels of intelligence.