Show HN: Querying JSON documents using SQL-like language in Scala
github.com
github.com
The first one, which is definitely close in spirit to this project is OctoSQL[1]. We too are trying to query all the things with SQL, including JSON, though we extended the idea to supporting more data sources and file formats, with the ability to mix and match in joins and subqueries. Curious what direction this project will take though, as we’re now heading for streaming sources (like Kafka) in pure SQL!
A totally different one is jql[2] where I’ve been exploring more of a lispy continuation passing style based approach to a JSON query language, as I didn’t really feel like the current ones are ergonomic.
If you like this, make sure to check them out too! (And all the others people will be posting, as I think they all are fascinating! This is a big area with lots of space for innovation left.)
We ourselves mainly work on the command line interface and concentrate on it.
Here's an example query: https://datasette-jq-demo.datasette.io/demo?sql=select%0D%0A...
EDIT: to add context to the story - Apache Drill is running on JVM and can be embedded, so it can be run from Scala code as well.
[1] https://stedolan.github.io/jq/ [2] https://github.com/doloopwhile/pyjq
However, I think this may be indicative of the task you are trying to complete rather than the tool itself. At a certain level of complexity your resulting code will just look bad regardless of which language you use. One thing jq does to make this easier is providing the pipe operator, so you can break a complex task into stages: A | B | (C .... | C.5) | D | E...
It is more verbose for simple queries, but with named variables and functions, larger queries do not become complicated
I still prefer jq for JSON processing though. I'm hooked to pipelines. EDIT: The great power of jq is in transforming documents, not querying.
In saying that, I guess its still better than having yet another custom query language a la mongo.
Come one read the second line of my reply before strawmanning.
Disclosure: works for Databricks but not on spark