Azure Data Lake Store: a hyperscale distributed file service for analytics
blog.acolyer.org
blog.acolyer.org
There are a lot of similarities between Scope and USQL, but USQL is also a cleaner (and in some ways more rigid) implementation.
I see it described in many different ways but ultimately that's what it sounds like to me.
Obviously I'm sure there's massive effort involved under the hood in presenting it as "just a filesystem" to the end-user (or end-process), but am I missing anything else?
A distributed file system is just a convenient way of representing it. Let's say you're processing some incoming feed of data. You place new batches of data in this distributed file system. Your processing pipelines consume the data, store some intermediate output, which is then consumed by the next step, etc. — your pipelines will typically transform (clean up, extract, split, etc.) data, potentionally running asynchronously correlations with other data sources (e.g. your feed has street addresses but you want lat/lon coordinates), aggregating it, etc. The steps may store their outputs, and the end of the pipeline will also emit data back into the data lake, from which it can be fed into other, structured data management systems.
One benefit of modeling this as a classical (albeit distributed) file system is that it's familiar to users and can potentially be mounted directly (via something like FUSE) and used interactively, e.g. end-users can manually upload job data or take output data into programs such as R or Mathlab.
Amazon S3 can be considered a data lake, for example.