Big data applications tend to use other structured binary formats like parquet and avro, which any big data tool can typically parse.
Big data applications tend to use other structured binary formats like parquet and avro, which any big data tool can typically parse.
If you have any idea how to do it, I'd be more than glad to hear about it.
There are free alternatives, but many require programming.
SQLite on the other hand has a lightweight SQL REPL that can be invoked from the command line.
Spark can work with SQLite via JDBC, though obviously it isn't as native as Parquet. Between SQLite and Parquet, I might pick Parquet under most circumstances.
But it seems to me SQLite ought to at least be a better option than CSV for Spark jobs (less work needed to do type inference, predicate pushdowns are trivial, etc.)
SQlite is usually a single file representing a database though, I don't know how it would work with partitioning and stuff, and then how to handle the schema evolving across sqlite files.