But with Snowflake, the data never comes out. Can't use Spark/Trino/Flink... on data in SF.
This is the problem. Both Snowflake and Databricks are spreading FUD and otherwise smart people are falling for it.
For all intents and purposes, large amounts of data are locked into Snowflake. Is it theoretically possible to export a petabyte out of SF? Sure.
Do I want to spend money on it? Not really. That is what I mean by the "data doesn't come out".
"Exporting" a petabyte out of Databricks is a no-op. I can already read Deltalake from other open source tools.
There are different ways to lock customers in and both Databricks and Snowflake are playing the game.
Every vendor, be it Snowflake, Databricks, EMR, Athena, BQ, … charges for use of the engine. The difference with a Lakehouse is that one doesn’t have to pay the vendor for the simple ability to use the data with another offering. That’s what you have to pay for with closed systems, whether it’s data on the way in or data on the way out.
This is just FUD.
Consider a scenario where data is coming in periodically, say daily, from some source, server logs, sensor data, whatever. And the user wants to train models daily on the data and they also want to do some SQL. Maybe they ingest the data directly into SF and copy it out for training, or they do it the other way round, land it in object store and the ingest into SF. This is unlikely to be a humongous amount of data, it's probably not a PB. However, this adds up, maybe for some use cases it becomes a PB in a month, maybe in a quarter, maybe it only adds up to a PB in a year.
Thing is, without a Lakehouse architecture, the user will pay to store and copy that data multiple times (at least twice) no. matter. what. They may not pay for a PB in one shot, but you can bet that eventually they'll pay multiple times to store and copy that PB.
Do you have to pay to export data out of Databricks? No, it's already sitting where you want it.
Which one is open? I wonder