Eva – A distributed entity-attribute-value database in Clojure
github.com
github.com
While there are many excellent ideas embedded in Datomic and these projects, for me just being able to persist the same data structures you're using at a repl and query for them with data is a huge win vs having to start translating types and concepts and query strings to and from SQL is a huge win.
It's worth noting that time now has two levels of meaning in this context - transaction time and valid time.
Disclosure: working on Crux
i do like passing in a timestamp and getting the state of the world as Crux seems to implement, whereas Eva seems slightly less oriented on time based querying, more so for a sort of log setup (where they have EAV+T+added?)
`EAV+T+added?` is the canonical data representation used by Eva, whereas Crux relies on two discrete "document" and transaction logs as the canonical representation. There are many reasons for this (data eviction, "unbundled" scalability etc.) but bitemporality (i.e. the ability to transact into the past/future) is definitely the primary reason why Crux doesn't follow the `EAV+T+Added?` pattern. You can get a feel for the indexes Crux uses here: https://github.com/juxt/crux/blob/master/src/crux/codec.clj#...
Now if only they’d make it into open source projects in languages other than Clojure. I want my Datomic-for-Elixir, darn it! (And, if I wasn’t too busy, I’d be the first person to volunteer to build it!)
(Though I don’t think there’s been a new major infra component project started since Elixir became a viable contender, so maybe that could change.)
Mnesia has had features like disc copies bolted on after-the-fact, but you can tell from the way they’re implemented (e.g. manual node-crash recovery) that they’re not used by Ericsson in the form you find them in OTP, but rather that these are just framework hooks where Ericsson has built (and expects you to also build) specialized-to-your-use-case DBMS logic on top of.
And, likewise, Mnesia assumes OTP’s “distribution set” model (i.e. a static set of known operational relationships between nodes) rather than a clustering model; Mnesia has no clustering support per se—you need to take down the whole Mnesia application across the distribution set and start it back up if you want to change its node membership.
These things are fixable, but the result wouldn’t be really be “Mnesia” itself any more, but rather a freestanding DBMS system that uses Mnesia as its storage- and transaction-linearization engine, maybe Lasp for clustering, CRDT-annotated field-types in table schemas for resolving state after crashes, etc. Still, I’m surprised nobody has bothered to build such a thing and open-source it.
That said, I generally agree with you. Mnesia is perfectly serviceable, but it is, as you say, really only a good fit for manual intervention in the event of service interruption. That manual intervention can be after the fact, but it can't automatically resolve a partition in the event of a netsplit. It's not too bad to extend to write code to automatically resolve those partitions (I've POCed it; we ended up with a manual button to execute the code to do it just because of our own concerns, but it worked quite well), but you have to start making some decisions, and write some code, on how to do that.
Which is why people tend to use other things; there are other solutions that have "what do I do in the event of a netsplit" already built in. And while they may be faulty under certain circumstances (as Jepsen tests have shown is usually the case), most people will take mostly correct self healing over no/DIY self healing.
A few other things; Erlang's default distribution (which Mnesia leverages) was not built for remote distribution (i.e., nodes located in different data centers), so no clue what happens there. It also was built with a fully connected topology, which limits how many nodes it's reasonable to connect to the cluster.
Datomic seems to be one of the best ones in terms of execution of the ideas behind it and usage in real world projects, to say the least.
I'm confused -- does this mean that Workiva themselves are not using Eva? Or are they still using it, but not officially developing it any more? If they were really invested in it, why would they only allow employees to work on it in their 10% time?
Source: I work there, although I have literally nothing to do with this project.
I have created almost identical implementation of this as an embeddable library in a year 2000 along with proprietary SQL language. The reason was that my client had inventory of products with some crazy amount of attributes and each product can have its own set and the client kept changing, creating, deleting those. It was in memory but with persistence and atomic transactions. No history though. It was blindingly fast on complex queries. And the schema of the database was kept as a set of entities with some predefined names, values and range of id's .
For a while I was contemplating releasing it as a standalone product but as I had enough tasks on my plate decided not to do it. Kinda feel sorry now ;(
So all in all very close by idea.
In the context of Magento, it is a real _nightmare_ and is one of the major contributor of the slowness of the Magento platform (at least for magento < 2.0).
The reason for this slowness is that in a relational database, the EAV model makes it so that e-v-e-r-y s-i-n-g-l-e SQL query is one gigantic query made of tons of JOINs.
To give you an example, querying a product in Magento may need to join no less than 11 tables!
catalog_product_entity, catalog_product_entity_datetime, catalog_product_entity_decimal, catalog_product_entity_int, catalog_product_entity_gallery, catalog_product_entity_group_price, catalog_product_entity_media_gallery, catalog_product_entity_text, etc, etc.
In order to fix this issue, the Magento team created what they call "flat tables" which are tables that are created by querying the database with an EAV query (i.e. the query with a million joins) and putting the results in a table with as many columns as there is attributes being returned by the original query.
In theory choosing to use EAV was an amazing idea. In practice, this idea did not scale for large Magento stores and it has made Magento hugely complex, slow and hard to use.
We use Magento at betabrand.com and I can confidently say 90% of the slowness of our website is due to Magento's EAV tables and we have spent a humongous number of engineering hours optimizing this.
My main complaint with Magento 2 is with the feature that is it's biggest selling point: its flexibility. The fact that any public method on any class is wrappable/replaceable, and that any class in the system can be replaced wholesale, and any javascript or any template in the system can be wrapped or replaced, from anywhere, by any module, at a distance, just makes the whole thing a huge cluster to deal with once you get any number of 3rd party modules.
And, yes, you want your tools and application to abstract away some of the messiness involved in something as simple as “load/save this object from the database,” otherwise you’ll end up with a much larger amount of code for such a simple operation. I’ve seen it done in a Django app, and it’s not bad to work with once those abstractions are in place.
EAV performance at large scale is really dreadful. So many queries.
Additionally, the ability to pin a database at a specific point in time is useful for APIs. For one, on pagination so that you don't see pages change out underneath as records are updated. Another is in APIs that join configuration values with time-series data or aggregations. If our data decisions drive revenue, for example, looking at 2018-08-31 revenues with today's business data has a lot less insightful value than being able to see them with other data pinned at that same date. Pointing the queried DB to a fixed timestamp in Datomic is free.
One of the motivations for Crux was that `as-of` queries still weren't powerful enough for what the business really needed, which was something like `as-of` that could cope with retroactive corrections for out-of-order and delayed ingestion (of upstream transaction data).
[0] https://gist.github.com/grafikchaos/1b305a4e0b86c0a356de
I used EAV in an application and it was fine (I used Datascript on the client-side, with server-side data being stored in a JSON document database, RethinkDB). I did run into some annoying issues with the lack of nil value handling. I'd say writing queries was difficult, but not significantly more difficult than in any other language, if you care about performance and want to know what the query actually does.
Overall, I felt there was a good mapping between my domain model and EAV.
I eventually dropped this solution in favor of Clojure data structures: if you have an in-memory (in-browser) database anyway, why keep the data in Datascript, if you can simply keep it as ClojureScript data structures?
TopLink was the ORM that made me hate ORMs.
Versant was building systems that were updated on the fly without shutdown, which was quite an achievement back then (and probably still is: Imagine updating a running instance of PostgresQL)
The Object-Oriented Database (OODBMS) space is basically just Versant. Actian however seems not to be developing it actively anymore and it's generally not that much advertises (they rebranded it as "Actian NoSQL" apparently).
I wonder if anybody else still uses it (besides us).
* Generic/"bootstrappy" looking site, mostly marketing speak.
* Not even one code example of what it looks like to solve a problem with this DB.
* Gigantic "Get a Trial" button that links to a lengthy form that I'm sure has 99% abandonment rate.
There are a lot of other software products that are marketed like this, so I'm genuinely curious how this works.
1: https://www.actian.com/data-management/nosql-object-database...
congrats on speaking about a product you haven't even tried.
The transaction log design is a fundamental design tradeoff that Eva/Datomic/Crux share, which means system throughput is limited by the throughput of a single process. The argument in favour of such a design is that most businesses & business applications don't actually experience transactional data volumes over 10K tx/sec.
What are the failures of EAV except when applying it to a database optimized for something different?