Whenever I encounter a "data lake" it's just a bunch of random CSV and JSON files on S3, and a bunch of half baked Python scripts to query it. And usually somebody tells me it's "big data", and that's why it is has to be that way. So I check the size and it would easily fit on a laptop from 15 years ago.