What are some good resources that can help educate folks on these differences?
What are some good resources that can help educate folks on these differences?
- data warehouse: schema on write. you have to know the end form before you load it. breaks every time upstream changes (a lot, in this world)
- data lake: schema on read. load everything into S3 and deal with it later. Mongo for data platforms
- data lakehouse: something in between. store everything loosely like a lake, but have in-lakehouse processes present user-friendly transforms or views like a warehouse. Made possible by cheap storage (parquet on S3), reduces schema breakage by keeping both sides of the T in the same place
Most of the content that I've seen in this area is really high-level. I'm trying to write posts that are a bit more concrete with some code snippets / high level benchmarks, etc. Hopefully this will help.
covering:
- What’s a Data Lake and Why Do You Need One?
- What’s the Differences between a Data Lake, Data Warehouse, and Data Lakehouse
- Components of a Data Lake
- Data Lake Trends in the Market
- How to Turn Your Data Lake into a Lakehouse
The advantage is that various query engines will make it quack like a database, but you have a completely open interop layer that will let any combination of query engines (or just SDKs that implement the table format, or whatever) coexist. And in addition, you can feel good about "owning" your data and not being overtly locked in to Snowflake or Databricks.