Just about every company these days has their data spread out all over the cloud: marketing data on Facebook and Google, social media data on Twitter and Snapchat, customer data on Salesforce, sales data on Shopify and Amazon, and so on. Most companies will either (a) hire a team of data engineers to collect and exploit this data or; (b) hire an expensive consulting firm to build an ETL pipeline, or; (c) let this data rot in the cloud. For the past 6 years, I've worked as a data engineer (where I became intimately familiar with Facebook and Salesforce APIs), and I'm confident that I can automate around 80% of my job.
It's clear that the value prop is astronomical: just one data engineer will run you at least 150k/yr and most of the work will involve maintaining API data pipelines. Having a "one-click" solution where one simply provides an API key and what data they'd like to warehouse (e.g. marketing data, social media data, customer data) and where (FTP, S3, Redshift, DynamoDB) would be invaluable to companies that want to make sure they exploit this treasure trove.
Some hard/interesting problems:
- API specs constantly change (Facebook, for example, has a quarterly update schedule)
- Inferring JSON schemas is hard
- Data integrity is hard (data types sometimes change willy-nilly)
- API rate limiting is tricky
- Resilience is hard
- Recovering old data (especially for certain services) might be impossible
Everyone is starting to become keenly aware that letting the data rot is starting to have a higher and higher opportunity cost. Not warehousing your own data is simply not a tenable option any more: the world’s most valuable resource is no longer oil, but data[1].
[1] https://www.economist.com/leaders/2017/05/06/the-worlds-most...