You could also just continue pre-training of an existing foundation model. Would still be cheaper by not starting from zero.
The amount of accuracy while doing fine tuning or distillation is usually better than pre-training an existing model, not to mention the graph against the cost.