So amongst the cloud providers, AWS calls a combination of S3 + Glue + Athena (for example) a "data lake", where S3 is the object storage which can store data in various formats, and Glue and Athena are used to transform/process/query the data. See a more detailed article/guide here: https://aws.amazon.com/lake-formation/
If you didn't want to put anything into the cloud and keep all your services on-premise, a local Hadoop cluster could be a data lake, for example using HDFS + Zookeeper + YARN + Hive.
[This is a huge over-smiplification because it's late and I really should be going home :)]
When people just dump data in their storage they end having really hard time sharing them in their organization.
Trying to come up with some unified standard or common API for the extraction, transformation and implementation of useful data from a heterogenous collection of systems sounds like a problematic task at best.
The other, more likely, interpretation of 'data lake' is that it is the staging ground between your ability to do the above stated activity and other downstream systems interested in the data. If the idea is that you are creating the actual normalization layer, I feel like this is still more in the realm of SQL/ETL, as there really isn't any other direction to move that would reduce your entropy in a valuable way (relative to your time invested).
SQLite or Postgres is usually the right choice. This simple rule can help you avoid a lot of pain. Once you have convinced everyone that Postgres is to be used, the only other real barriers are your ability to get a data transport to each business system and the authoring of some SQL scripts. The workload of building a SQL representation of any particular business system is fairly predictable once you break it down to entity-relationship abstractions. Also, using a language with powerful class/object, serialization and database support such as C# or Java can cut your workload by orders of magnitude if you choose a SQL architecture. In C# for instance, you can just write POCOs and use Entity Framework to build out all of the SQL for you. This is not the most performant option by a long shot, but it can get you going incredibly quickly on a first iteration.
... the data has already disappeared.
Systems get shut down and replaced. Operational systems may discard history.
By the time you get a fully operational data warehouse set up, it may be too late to preserve the data.
The key line for me:
"The data lake stores raw data, in whatever form the data source provides."
The emphasis on"raw" was his, not mine.
A data lake is like you said a collection of data stores, and the industry as a whole hasn't defined it very well past that.
IMHO - A (useful) data lake is a platform that can support any type of data store in any format (be it relational, flat, graph, document, etc) and offer a way to consistently query it. A (useful) data lake does much in the way of managing metadata about those data stores that makes it easy to consume.
This means ingestion is faster (no transformation) and you don't throw away any data that you might want later. If multiple teams want to query the same data in different ways they have the ability to do so. And ideally it prevents data silos because everyone can stuff their raw data into a master data lake and each team has access to all the data but is responsible for doing the work to make it look like they want.
Reality of the above obviously doesn't always match the theory but schema-on-read/ELT are the easiest ways to handle the above. Typically this involves some kind of Hadoop-style technology, like Hive or SparkSQL for SQL-based querying, Spark for non-SQL, etc. But you've always got the raw data and can go back and re-ELT it from the data lake if your needs change.
Data lake = place to store unstructured raw data. Usually as files in an object store or Hadoop/HDFS cluster. Analyse with data processing/SQL frameworks. Schema may be part of the data (parquet, avro) or on-read (raw csvs or json modeled into tables).
Data warehouse = place to store processed (semi)structured data. Usually in a distributed columnstore database. Analyze with SQL. Schema is pre-defined by the tables. Usually for smaller, faster, real-time queries or as a cache in-front of a data lake.
Some cloud data warehouses like BigQuery and Snowflake can also query unstructured files and even run on top of the object store so the boundaries are getting blurred. Will probably converge at some point in the future.
A data lake is when you store your data in object storage (S3, GCS etc) as opposed to a filesystem (HDFS) or some indexed datastore (Redshift etc).
This potentially saves a lot of money because you can scale compute separately from storage, and object storage can be extremely cost effective compared to running a distributed filesystem.
Where the two overlap is when you store something like parquet in object storage, the file format is somewhat indexed already so you spend a bit more money preprocessing it but save a lot of money querying it.
I think whether its "raw" json or log files or preprocessed parquet doesnt really differentiate whether its a data lake or not
On another tangent, I wonder if it would be possible to make an un-warrantable cloud? Would it be feasible to create a distributed cloud within the borders of the US or another industrialized country which doesn't actually exist at a particular address?