The data lifecycle is waaay overpopulated with Data Scientists who are not empowered or knowledgeable enough to work with product designers and engineers to do everything that empowers Data Science and ML.
We need more Data Engineers involved at time zero in projects to help:
1. Plan out what data should be produced/captured by the product
2. Instrument systems to actually generate data consistently and effectively
3. Build ETL pipelines and data management systems
4. Manage enterprise data sharing and resiliency
etc...
What ends up happening is you have a bunch of Data Scientists just handed a pg_dump or flat file from some ops team. That is typically missing data or poorly formatted and they spend 90% of their time cleaning it up then running some basic regression with numpy or whatever.
Need better understanding of the data lifecycle by organizations and investment in instrumentation and data management.