This is an interesting observation that deserves to be addressed.
I think in practice, what happens is that there's no easy transactional boundary that can keep the data ownership and ML in separate firms. The theory of the firm [1] states that firms arise when transaction costs are less than the economic inefficiencies of centralized resource allocation. There are some pretty heavy transaction costs between doing data cleaning and data science in separate organizations:
1.) DRM for datasets isn't really a thing, since to explore, visualize, and train on them, you need access to the raw data, and then instead of your machine-learning function you can just pass the identity function to get the raw data. DRM for consumers always relies on a publisher whose incentive is to stay in business (by not breaking any laws) rather than to obtain the raw data.
2.) You don't know whether a cleaned data set will be useful for your ML application until you've inspected it, visualized it, run some statistics over it, etc. at which point you've done most of the work for setting up your models. That means there's a big risk premium for buying data, and potential buyers usually want samples & statistics before committing a lot of money.
3.) At the price that good datasets go for, you're usually dealing with enterprise sales, which involves commissioned salespeople, face-to-face meetings, expensive dinners out, etc.
4.) The type of labeling you need to do is often intimately connected with the usage of the ML model. What constitutes spam? What constitutes abuse? These are questions for your policy team, which the data-collection organization would have no visibility into and no way to set up a one-size-fits-all policy that all potential customers would be okay with.
5.) Software markets tend to be winner-take-all, which means you're dealing with a handful of customers, and one will likely become dominant and able to acquire you.
Instead of separate firms for data ownership, ML, and user interface, usually the service that makes the final product will end up collecting or buying the data outright, hire data-scientists to do the ML, and then offer a product. The end result, economically, is what you say: the money is in data hoarding. But it isn't apparent in prices of a sustainable customer ecosystem, it's apparent in acquisition prices for startups with data vs. salaries of data scientists.
[1] https://en.wikipedia.org/wiki/The_Nature_of_the_Firm