Having a data lake - which I understand as a repository of raw data of diverse types, regardless of the tools - structured in a tool like S3 is very useful when you have multiple use cases over data of different kinds.
For example, you could store audio files from customer calls and have them processed automatically by Spark jobs (e.g. for transcript and stats generation), structure and store call stats on a database for analytics, and do further analysis via notebooks on data science initiatives (e.g. sentiment analysis). This is akin to having a staging area for complex and diverse data types, and S3 is useful for this because of its speed, scale and management features.
Teradata or Snowflake aren't a great fit for use cases like these, but they are great if the use case is to get answers to questions like "top 3 operators per team in volume of calls, by department and region, in last quarter" if the volume of calls is big.
If I understood correctly, I think your comment was more focused on why use new tools when the existing are mature, but I think big data tools have had to become more specialized and targeted for specific use cases. But if the question is "why build more than one data lake", the only reason I can see is organizational: teams or different areas of an organization either need their own data lake because they have specific needs (which is rare) or won't/can't collaborate with others to have a shared asset.
> they are great if the use case is to get answers to questions like "top 3 operators per team in volume of calls..."
You are straw-manning Snowflake/BQ. Just because they are SQL database systems doesn't mean you have to do 100% of your analysis in SQL. You can use other systems, like Spark, PyTorch, Tensorflow, to work with data that you manage inside a RDBMS. There are some issues with bandwidth getting data between systems, but these issues are getting solved (by Arrow!) and in the meantime unload-to/load-from S3 is a good workaround.
I've heard a lot of people make these same arguments and I've tentatively concluded it's mostly motivated reasoning. Engineers like to engineer things. They start by trying to make the obvious, boring system work, but when they run into an obstacle they immediately jump to "I need to build a new system using $TECHNOLOGY."
Databricks Delta Lake has some use cases, but there are some aspects that are rough around the edges. vacuuming is very slow, the design decision to store partitioned data on disk in folders has certain pros / cons, etc.
There are a lot of great products in the data lake space, but lots more innovation is needed going forward.
This is just flatly untrue---they are nearly the same cost/TB as object storage, and they store everything in compressed columnarized format, so they're about as efficient as you can get.
I have heard many people make the same claim. I can't figure it out. Is there something wrong with my calculator???