Spark's initial path to success was "a faster way to process your data in HDFS". Cloudera was selling users Spark before DataBricks was even founded. The idea was that Hadoop was an ecosystem of tools for processing data built on commodity storage and compute hardware, for when your data was too big and expensive to transfer to the cloud.
Over time it became increasingly popular to use cloud storage instead of running HDFS. This really destroyed Cloudera's moat, because there was no operational overhead to putting your data in S3 or GCS. You just needed to run some stateless compute, and if you fucked up it didn't matter. Nowadays your "data lake" is a bunch of files in commodity storage someone else runs.