Converting a Parquet table to a Delta table is an in-place, cheap computation. You can just add the Delta Lake metadata to an existing Parquet table and then take advantage of transactions and other features. I don't think it's a meaningless comparison.
Iceberg is cool too.
Yes, Parquet can be compressed with zip, but snappy is much more common because it's splittable.
Parquet tables can be registered in a Hive metastore. Delta metadata can be added to a Parquet table to make it a Delta table.
This is my point though? This is an apples to oranges comparison. A directory of Parquet files is not a table format. Comparing Delta to Hive or Iceberg is a more apt comparison. I have worked with all types of companies and I have yet to work with one that is just using a directory of Parquet files and calling it a day without using something like Hive with it.
I don't really see how Delta vs Hive comparison makes sense. A Delta table can be registered in the Hive metastore or can be unregistered. If you persist a Spark DataFrame in Delta with save it's not registered in the Hive metastore. If you persist it with saveAsTable it is registered. I've been meaning to write a blog post on this, so you're motivating me again.
I've seen a bunch of enterprises that are still working with Parquet tables that aren't registered in Hive. I worked at an org like this for many years and didn't even know Hive was a thing, haha.
You are right about Delta tables in the Hive metastore but if you are writing from the perspective of "there are companies that don't know what Hive is" then I feel the next step up is "there are companies that just stuff files in S3 and query them with Athena(which handles all the Hive stuff for you when you make tables). Explaining what Delta gives them over that I feel is something worth explaining.
Also, as someone who chooses between Parquet and Delta on a fairly regular basis, I can say from experience that there are as many situations where both are viable options than there are situations where such a comparison is invalid. So, it's hardly an apples to oranges comparison. At worst it's maybe a pomelos to oranges comparison.
Delta Lake is a storage framework that uses Parquet files. Parquet files are a thing, but Delta Lake files are not a thing. The Delta Lake framework uses Parquet files, plus some additional stuff (transaction log, checkpoint files) that enable capabilities that Parquet files alone do not.
That collection of files is called a "Delta table". If the title of this was, "Benefits of Delta Tables vs. Parquet Files Alone", and the article was revised to be more careful about not conflating Delta Lake and Parquet, I think that would benefit everyone.
I think that the people elsewhere in this thread who are trying to be pedantic about this are maybe more familiar with some of the individual open source technologies that are used in data lake applications than they are with conventions around how these technologies get assembled into a full-fledged data management system in a business setting?
In other words, it seems like people who are trying to talk about the forest are getting downvoted and picked on by a bunch of folks who seem to maintain that there is no forest, only trees.
Absolutely, and standardizing that has really cool benefits. You're clearly knowledgeable enough to know where the author means "Delta table" instead of "Delta Lake," "Parquet tables" instead of "Parquet", etc., but not everyone is.
I can understand if the author feels picked on, but I'm sure he knows the bar for technical correctness for developer content marketing is high (especially on HN). Honestly, if I were Mr. Powers I'd be happy for the "strict mode" feedback!
Almost as if people were trying to use HN's voting system as an ersatz referendum on their preferred big data packages rather than as a way to self-moderate the quality of the discussion.
We're supposed to be a crowd that favors mature discussions about technical topics. We can do better.