I ask because, if I didn't know either word, the one would mean, to me, "tiny storage next to a big body of data" and the other would mean "a big body of data".
I ask because, if I didn't know either word, the one would mean, to me, "tiny storage next to a big body of data" and the other would mean "a big body of data".
Data Lakehouse involves adding things like the ability to query via SQL, the ability to update/insert/delete, transactions.
Where before people needed warehouses for BI and lakes for data science, they can now have only one approach.
It’s likely to be a big trend as data moves to this format and arrangement and the DBMS vendors like Snowflake have to play nicely with it.
This is all very interesting, and thank you for taking the time to explain. Any good starting points for someone who would like to know more?
It’s a technology independent pattern though.
It's not only the names, but something like 98% of the tools too, that suck.
1. Snowflake has always used blob stores + file data + metadata. Architecturally it’s actually always been very Lakehouse-y
2. Parquet and Iceberg should be equivalent in performance and features. It’s more than playing nicely - it’s more choose your own adventure where all things are equal.
Thank you for that. Do you have any suggestions on where one would start if they wanted to get a better idea and/or some experience using lakehouses?
Databricks (with Delta as the underpinning) seems to have lead the charge of lakehouse meaning, your data lake+file formats/helpers+compute==data lake+datawarehouse==lakehouse.
The latter seems to be the prevailing definition today with the former aging in place.