I think this is changing. Previously you had to train something massive off a huge data set. But now it's moving towards having a pipeline of pre-trained models that are trained on massive data sets and then smaller models that you train in house to tweak results from that pipeline. Any start-up should be able to get its hands on enough data to train a LoRA for example. There are good enough open source components to build a moat out of a good pipeline with one or two components in the pipeline trained in-house and the rest pretrained.