Ask HN: Why do you use a data lake instead of a data warehouse?
- Cost
- Supporting data science workloads
However, the emergence of data warehouses that separate compute from storage, like Snowflake, BigQuery and Redshift RA3, has drastically reduced the cost advantage for data lakes. And in my experience, data science workloads are generally compute-bound, so in principal you can just execute a SQL query from your data science environment (Spark, Python) against your data warehouse---the overhead of serializing and deserializing your data an extra time should be a tiny % of your overall runtime.
I suspect that the reason why some data scientist prefer data lakes is cultural rather than technical---basically, they have an irrational dislike of SQL and RDBMSs. However, I haven't done data science work myself in years, so maybe I'm just not getting it. Are there good reasons for using a data lake rather than Snowflake, BigQuery, or Redshift RA3? Please enlighten me!