The difference in my experience comes from the data-set. If you have an unusual and proprietary dataset, then off-the-shelf models are only a starting point.
Ingesting non-public data that matches the format that will be used during ingestion during implementation so inference will be more accurate is my reason.
I train LoRa Models or finetune foundation models for the task at hand. I've not encountered the need to train a foundation model, yet.
Can you share some example descriptions?
Does fine tuning count? If so there isn't a model for the input I'm working with that I know of