Thinking in Datomic
pelle.github.com
pelle.github.com
Since I'm writing a history-aware application at the moment, I recently looked into different patterns for this and trying a mixed strategy at the moment (SQL DB used as document store and a single event log, that accumulates changes - a lean approach though, a few hundred lines python for the data access layer; what always gets ugly is the validation, which your application must take care of).
I wish, there was more hands-on material on the subject (some resources dive depth into bi-temporal modeling, but I feel your schemas can get complex (= expensive) very fast).
[1][2] 'persistent' means different things here.
If only there was an open source implementation of this concept that I could run on my own hardware. Does anyone know of such a beast?
And, no, I don't want to roll my own versioned/timestamped row schema in an RDBMS - been there, suffered that.
I'm working on a project to bring persistent data-structures to disk ( https://github.com/DanielWaterworth/Siege ). It's currently undergoing a rethink, hence the lack of recent activity, but it's still in development.
(that is, concatenate the entity, attribute name and timestamp in a lexically sorting string)
At larger scale, hyperdex http://hyperdex.org/ might do the job but I don't know much about it
[1] http://en.wikipedia.org/wiki/Persistence_(computer_science)
It might be of interest to you given your current project.
Unfortunately, it is also ridiculously expensive.
A more extensive description/case study of bi-temporal database design: http://www.cs.arizona.edu/people/rts/tdbbook.pdf
This case study tried to be somewhat rigorous, but from what I recall the SQL did get complex and expensive pretty quickly. There are a few attempts to create products to mask the complexity as well:
http://www.timeconsult.com/Software/Software.html
For living in time and dealing with it regularly in a common sense sort of way, it is really rather challenging to manage effectively as data.
That is, I've seen this pitched as the solution to all our data woes numerous times now, why is this time different?
Their Entity class makes it feel a bit like an oodbms though in that it kind of acts like a hash map.
My previous explanation already papered over a ton of fundamental differences in order to make a somewhat understandable generalization, but I can't make the jump to triple-store from my understanding of Datomic.
I basically see it as "just a single-writer many-reader distributed fully persistent data-structure."
If it's because you store objects as many triples with the same subject and different adjectives... err, well, yes, that's how you store objects in a triple store. Not remotely new, and still subject to my initial question.
The main difference is that in Datomic the timestamps don't identify a single fact but the entire object graph at that point in time.
To make an analogy; saying that Datomic is just a triple store is about the same as saying that git is just a store for lists of diffs.
With everything else, the standard RDBMS table could be considered as having a 'snapshot' of the Datomic values.
I'm still not sure what benefit this has over a traditional DB. Perhaps I'll just have to wait for the next post.
Both of those sound like bad deals if there's a packaged solution like Datomic that's built specifically for the use case.
Also as it does it on a datom level it is a lot more efficient than versioned rows.
Interestingly one of the early selling points of Postgresql was this time travel functionality. But it was yanked out in 6.2
I'm sure that Datomic has it's place, but the examples you gave in this post aren't that convincing.
I will be getting into more detailed examples later. My real point with this post was more to talk about how to model the data using datoms and not specifically the temporal aspects of it.
I made many mistakes in my original data models in datomic based on many years of rdbms thinking.
Playing to the Well I'm obviously above average gut feeling. Cute.