Crux SQL
juxt.pro
juxt.pro
I just switched the forked driver repo to public if anyone wants to test it out [2]. The Metabase driver docs are pretty straightforward to get things running. There's definitely work to be done though and I didn't get very deep into it really - but I hope to pick it up again soon!
[0] https://github.com/metabase/metabase/issues/5562
[1] https://youtu.be/StXLmWvb5Xs?t=996
[2] https://github.com/crux-labs/crux-metabase-driver
(I work on Crux :)
EDIT: it's also worth a mention that the linked issue #6230 discusses a lot of problems deriving from Druid's lack of support for prepared statements, but Crux and Dremio don't have that limitation. Although reading the most recent comment it looks like Druid may have overcome that hurdle now, so with a bit of luck there might now be more traction to get mainline Metabase support for a generic Calcite driver!
https://docs.datomic.com/cloud/analytics/analytics-concepts....
I.e. group by, window aggregates, order by, limit
I'd be interested in seeing the equivalent datalog that Crux makes given some of those statements for learning what efficient datalog looks like for some of those problems
As for your specific examples, Crux already implements an order-by+limit+offset combination that automatically spills to disk if needed, but for efficient pagination you would probably want to maintain additional pre-sorted value ranges. For basic aggregation we have an alpha API decorator that composes Clojure transducers to great effect [1].
[0] https://docs.mongodb.com/manual/reference/operator/aggregati...
[1] https://github.com/crux-labs/crux-decorators/blob/master/tes...
I also enjoyed this talk: https://youtu.be/oo-7mN9WXTw though it’s not so much about syntax as the logic of it.
I see 4 main reasons why someone may want to be aware of Crux:
- if you have a bitemporal problem
- if you have a graph problem, i.e. something you might initially look to Neo4j to help with
- if you want to use Datalog because it can make writing an application simpler
- if you are thinking of building something similar (immutable event log + indexes) and want to save time
Crux is very different from Cassandra (strong consistency, fat nodes, arbitrary joins etc.), but you could definitely use Cassandra in your Crux architecture.
The closest "yet another" comparison would be Datomic. Crux and Datomic both strive to reimagine what a "general purpose" DBMS should look like, with the primary goal being developer productivity/sanity, whereas Cassandra's goal is simply to be a highly-scalable document store.
Hope that helps!
Temporal graph analysis of "evolving graphs" is an active research field with some strong motivating use-cases, for instance: profiling networks of fraudulent transactions across N bank accounts with data pulled from M source systems. This paper discusses the analysis of research citations over time, as another example: https://pdfs.semanticscholar.org/110b/0db484a1303eda30aa7e34...
That said, Crux's indexes aren't optimal for making all kinds of analytical time-range queries efficient just yet. Instead Crux is currently focussed on point-in-time queries, but the temporal R&D is still happening as it feels very ripe.
Looking for examples more generally, I think wherever you have a meaningful use-case for a graph database you probably, eventually, will want to capture and model history. If you then find yourself with two or more such databases that you want to integrate, then you will greatly benefit from a bitemporal graph DBMS.
As a fun example, I like to envisage integrating our two federated evolving knowledge graphs. Imagine a tool for "networked thought" like Roam Research that could allow us both to visualise the evolving connections between our independently recorded thoughts, before, during and after this conversation. Graphs of knowledge encoded in time.
2) performance of ad hoc as-of queries
3) ingestion throughput (RocksDB is _fast_)
4) eviction/excision throughput
5) a lazy query engine doesn't demand so much memory (because there is no need to hold entire intermediate result sets at the same time), and automatic join re-ordering makes the Datalog inherently more "declarative"
6) use of protocols for modularity allows you to create a massive range of possible topologies to support the non-functional requirements of your host environment
7) benefit from the RocksDB roadmap (or other embedded KV storage - see LMDB / rocksdb-cloud)
8) absence of a prescriptive data model
On the flip side:
1) absence of a prescriptive data model (though transaction functions can give you equivalent power)
2) API maturity
3) lazy caching of data at peers (vs Crux' fat nodes, though again, see rocksdb-cloud for one possible resolution)
4) query features: multiple data sources, lazy entity API, other niceties
There are definitely things still missing from both lists :)
That's not to say that using the SQL API is a bad idea, but if you're using Crux from Clojure, you'd be missing out on some stuff.
One can make one solve problems of the other but they're really different mental models, which impacts the representation of the data needed to enable their use.
For Datalog one needs relationships- a graph- not what SQL calls relations, which is just data. Datalog requires a richer set of opinions- at least conceptually- about the data.
Languages like Datalog will not become mainstream until graph modeling is mainstream.
> The REST API also provides an experimental endpoint for SPARQL 1.1 Protocol queries under /sparql/, rewriting the query into the Crux Datalog dialect. Only a small subset of SPARQL is supported and no other RDF features are available.
There is an open issue in regards more general RDF support with details on the kinds of things Crux would need to add: https://github.com/juxt/crux/issues/317
[0] https://dsg.uwaterloo.ca/watdiv/
[1] http://swat.cse.lehigh.edu/projects/lubm/
[2] https://en.wikipedia.org/wiki/Subgraph_isomorphism_problem
Off-topic, but I quite liked CRUX linux when I did use it. It introduced me to BSD-style init scripts