Data pipelines are a real problem though, and I'm very interested in what startups do with this space.
Data pipelines are a real problem though, and I'm very interested in what startups do with this space.
Can you please elaborate more, thanks.
"Static ETL" like running the same database load every day at 1:00am isn't a super challenging problem. Doing it across many tables with complex transformations and multiple steps easily can be. You really have to consider reliability, processing speed, failure methods and other problems that dont really arise until you hit a certain scale.
There's also the issue of what people want out of a Pipeline that's changing. If you want people to be ""data driven"", then that means they need easy access to potentially all of your company's data on an ad hoc basis. So now your boring ETL 1 am pipeline isnt really serving any of these new usecases.
How do you create flexible pipelines that can be created from any dataset on an ad hoc basis? This is where tools like Airflow or Prefect come in. Creating a platform that can create these types of Pipelines is a real problem.
And before you even ask yourself _how_ to process this data, you need to also ask _where_? If you want to do what I outlined above - making your data more accessible and easy to use - then you probably need to rework how you're storing your data. But Data Lakes (and others) are a whole topic in and of itself.