Is there any advantage to having both a Data Lake setup as well as Snowflake. Why would one also want Snowflake after doing such an extensive data lake setup?
Many BI / analytics tools don't have great support for Data Lakes, so part of the reason could be supporting those tools (e.g. they still load some of their data to snowflake to power BI / dashboards)
We've solved that issue with Trino. Superset and a lot of other BI tools support connection to it and it's a very cost efficient engine (compared to DWH solutions). Another way to go even cheaper is using Athena, if you're on AWS.
Athena packages Trino - it’s in part a managed Trino service.
They are several versions behind, support for delta was added just recently. Also consider that with Trino you can build a cache layer on Alluxio, making it really fast (especially on NVMe disks).
Saving money 100% also lower latency on distributed access. Accessing file partitioned S3 doesn’t require to spin a warehouse and wait for your query to go on a queue, so if every job runs in like k8s you don’t have to manage resources and auto scale in snowflake is a “paid feature”
I believe just not having to handle a query queue system is already.
For one, Snowflake is expensive (you pay for the convenience and simplicity) and the data in there is usually stored in S3 buckets that Snowflake owns (and they dont pass along any discounts that they get from AWS for the cost of that storage).