Apache Hudi vs. Delta Lake vs. Apache Iceberg Lakehouse Feature Comparison
onehouse.ai
onehouse.ai
Vendor published benchmarks are worthless.
Feature matrices are extremely easy to game depending on your choice of rows.
Feature matrices are fundamentally flawed for the reason that the GP gave.
Data lakes are better for ML / AI workloads, cheaper, more flexible, and separate compute from storage. With a data warehouse, you need to share compute with other users. With data lakes you can attach an arbitrary number of computational clusters to the data.
Data lakes were limited in many regards. They were easily corrupted (no schema enforcement), required slow file listings when reading data, and didn't support ACID transactions.
I'm on the Delta Lake team and will speak to some of the benefits of Delta Lake compared to data lakes:
* Delta Lake supports ACID transactions, so Delta tables are harder to corrupt. The transaction log makes it easy to time travel, version datasets, and rollback to earlier versions of your data.
* Delta Lake allows for schema enforcement & evolution
* Delta Lake makes it easy to compact small files (big data systems don't like an excessive number of small files)
* Delta Lake lets readers get files and skip files via the transaction log (much faster than a file listing). Z ORDERING the data makes reads even faster.
The Delta Lake protocol is implemented in a Scala library and exposed via PySpark, Scala Spark, and Java Spark bindings. This is the library most people think of when conceptualizing Delta Lake.
There is also a Delta Lake Java Standalone library that's used to build other readers like the Trino & Hive readers.
The Delta Rust project is another implementation of the Delta Lake protocol that is implemented in Rust. This library is accessible via Rust or Python bindings. Polars just added a Delta Lake reader with delta-rs and this library can also be used to easily read Delta Lakes into other DataFrames like pandas or Dask.
Lots of DataFrame users are struggling with data lakes / single data files. They don't have any data skipping capabilities (unless Parquet file footers are read), their datasets are easily corruptible, and they don't have any schema enforcement / schema evolution / data versioning / etc. I expect the data community to accelerate the shift to Lakehouse storage systems as they learn about all of these advantages.
But then again, data lake may simply be what a data warehouse is now called in marketspeak.
Also, I stopped paying attention when the treadmill of new frameworks became unbearable to track, is spark now settled as the standard of distributed "processing", as in mapreduce / distributed query / distributed batch / etc?
I get that performance can improve by unifying to a file format like parquet, but again that seems like a data warehouse. A data lake should be something over heterogenous sources with "drivers" or "adaptors" IMO, in particular because the restoration of the data inputs stays in the knowledge domain of the source production database maintainers.
Data lakes are typically CSV/JSON/ORC/Avro/Parquet files stored in a storage system (cloud storage like AWS S3 or HDFS). Data lakes are schema on read (the query engine gets the schema when reading the data).
A data warehouse is something like Redshift that bundles storage and compute. You have to buy the storage and compute as a single package. Data warehouses are schema on write. The schema is defined when the table is created.
And yes, I'd say that Spark is generally considered the "standard" distributed data processing engine these days although there are alternatives.
> Data lakes are better for ML / AI workloads, cheaper, more flexible, and separate compute from storage. With a data warehouse, you need to share compute with other users. With data lakes you can attach an arbitrary number of computational clusters to the data.
- I am not sure it's any cheaper than BQ or Snowflake storage.
- Modern CDW separates compute from storage.
- I am not sure what you mean by "you need to share compute with others". Why?
- You can attach an arbitrary number of "clusters" in BQ and Snowflake as well.
Additionally, modern CDW provides a very high level of abstraction and a very high level of manageability. Their time travel and compaction actually work, and their storage systems are continuously optimized for optimal performance.
nick ( a t ) nickkarpov.com
I cant paste images here but imo this table comparing the 3 formats is the big takeaway https://assets-global.website-files.com/6064b31ff49a2d31e049... (explained inline, we do cite onehouse heavily but we are independent of them)
It resembles previous trendy technologies that are mostly forgotten now, such as:
- Lambda architecture (based on a wrong assumption that you cannot have a real-time and historical layers in the same system);
- Multidimensional OLAP (based on a wrong assumption that you cannot do analytic queries directly on non-aggregated data);
- Big data (based on a wrong assumption that map-reduce is better than relational DBMS).
I'm exaggerating a little.
Disclaimer: I work on ClickHouse, and I'm a follower of every technology in the data processing area.
I understand why folks want options. At the end of the day, folks want an easy to use, ALWAYS CORRECT stable database, with minimal well-documented predictable knobs, correct distributed execution plan, no OOMs, separation of storage and compute, and standard SQL, and Clickhouse struggles with all of the above.
(co-founder of MotherDuck)
I am interested to learn more about your point of view, as well as tangentially the strategic vision of MotherDuck as a company.
(VP Support at ClickHouse)
- Stability. It OOMS, your CTO mentioned that last week.
- It is not correct. I believe your team is aware of cases in which your very own benchmarks revealed Clickhouse to be incorrect.
- Scale. The distributed plan is broken and I'm not sure Clickhouse even has shuffle.
- SQL. It is very non-standard.
- Knobs. Lots of knobs that are poorly documented. It's unclear which are mandatory. You have to restart for most.
Don't get me wrong, I love open source, and I love what Clickhouse has done. I am not a fan of overselling. There are problems with Clickhouse. Trying to sell it as a superset of the modern CDW is not doing users any favors.
> Stability. It OOMS, your CTO mentioned that last week.
I ran ClickHouse clusters for years with zero stability issues (even as a beginner at the time) at an extremely large volume video game studio with real-time needs. Using online materialized views, I was able to construct rollups of vital KPIs at millisecond level while maintaining multi-thousand QPS. Stability was never a concern of ours, and quite frankly, we were kind of blown away.
> Scale. The distributed plan is broken and I'm not sure Clickhouse even has shuffle.
First, I hate the word "broken" with zero explanation what you mean by this. Based on your language, I'm assuming you're just suggesting the distributed plans aren't as efficient as possible, a limitation that the engineers are not shy to admit.
> SQL. It is very non-standard.
I would argue the language is more a superset than "non-standard". Most everything for us just worked, and often I found areas of SQL that I could reduce significantly due to the "non-standard" extras they've added. For example: Did you know they have built-in aggregate functions for computing retention?!
> Knobs. Lots of knobs that are poorly documented. It's unclear which are mandatory. You have to restart for most.
Yes, there are a lot of knobs. ClickHouse works wonderfully out of the box with the default knobs, but you're free to tinker because that's how flexible the technology is.
You worked at Google for over a decade? You should know. Google's tech is notorious for having a TON of knobs for their internal technology (e.g. BigTable). Just because the knobs are there doesn't mean they must be tuned, it just means the engineers thought ahead. Also, the vast majority of configuration changes I've made never required a restart...I'm not even sure why you pointed this out.
(Disclaimer: I have been using ClickHouse successfully for several years)
I do NOT work for ClickHouse, but I've been running super stable distributed CH clusters for years.
We have recently added support for Hudi and Delta Lake; you can check here: https://clickhouse.com/docs/en/engines/table-engines/integra...
It is a read-only implementation: ClickHouse can read and process the external data in the Hudi or Delta Lake format.
Apache Iceberg is pending. There is no good C++ library for it. But at least the overall structure is simple, as it is not hard to implement it.
The overall principle - whatever data format it is, ClickHouse should support it in a fast, stable, and ALWAYS CORRECT manner.
If you have more experience to share, please do it.
Just tried it out. It seems to only support partitioned Delta tables, non-partitioned return CANNOT_EXTRACT_TABLE_STRUCTURE. Is that on purpose or is it a bug?
The difference between MPP and something like Databricks or Trino working with object store is that while MPP can likely get much better performance and especially latency from the same hardware, operating it is much harder.
You don't "backup" Databricks - the data is stored in object storage and that is it. You don't have to plan storage sizing quarters upfront, and you never get in trouble because there is unexpected data spike. Compute resizes are trivial, there is no rebalancing. Upgrades are easy, because you're just upgrading the compute and you can't break data that way. You can give each user group (like batch one and interactive one, or each team) their dedicated compute over common data and it works. That compute can spin up and down and autoscale to save some money. You don't have to think about how to replicate my table across a cluster or anything. And so on, and so forth.
Running a big data and analytics platform - place where tens of teams, tens of applications and hundreds or thousands of analysts come for data and where they build their solutions - is already enough of a challenge without all this operations work, and that is why Snowflake and Databricks are worth that crazy money.
If someone could solve the challenge of having MPP that is as easy to manage as Snowflake or a Lakehouse, that would be quite the differentiator. And maybe you people already did and I just didn't notice, I don't know :)
But one thing to be aware of if your data gets big is that (to paraphrase the film) “We’re gonna need a bigger data lakehouse”.
Object stores are a terribly inefficient way to access and store changes of data.
- someone who thinks it’s often a bad idea
you can't do OLTP type workload with lakehouses - meaning updating the same value in a row million times per second, because object storage is not supposed to be used in OLTP way. however you can easily do that with relational DB, just UPDATE the value and thats it. underlying RDBMS engine will keep memory buffers updated, will update pages automatically and keep WAL for durability.
for lakehouse the proper way is to setup streaming processing and use in-memory cache to do hot data processing, and once data cools down - write to lakehouse table (object storage) like once per batch (like once every few mins). and restart batch from the beginning in case of failure - for durability.
The root of the problem is using object storage improperly.
>The root of the problem is using object storage improperly.
I don't see anything improper for any Parquet-with-metadata format.
The initial concept was created by Databricks in the CIDR Paper (https://www.cidrdb.org/cidr2021/papers/cidr2021_paper17.pdf) in 2021.