How we test at Nubank [video]
youtu.be
youtu.be
Complecting the two into, e.g. an Account object that has both metadata related to an account and e.g. rules related to transactions that can be part of an account quickly turns into an expensive maintenance nightmare.
And for finance it particular it's a very natural fit because there's just a lot of transformation of data and business logic.
I.e., you can ask to look at the entire database as it was yesterday, and run arbitrary queries against it.
You can also do speculative updates to it, in the sense of "show me the entire database as it would be if I were to have pizza for lunch".
It models this as a strictly linear succession of assertions and retractions of facts. Yesterday, `A` was true, today `A` is no longer true. While this new fact is recorded, it doesn't change the fact that yesterday, `A` was true.
What we see in reality is that append-only database is unusable without making additional "projections" or whatever you call them, databases that are ready to be queried/updated, with maybe specific denormalizations, indexes and so on.
And oh, btw, those later databases are not "imutable".
I highly recommend the talk “Domain modeling and with Datalog”[1]. It gives an explanation of how all this works, including indexing.
I don't even know how they have manage to scale datomic to that level, the support contract we had for datomic was only really used to report bugs[0] but they have more than 2000 datomic transactors? ouch.
[0] Yes, too much bugs and slow, but databases are hard so I guess this was expected for a closed-source niche DB with little users.
I was at the talk, they don't use spec they use prismatic schema iirc while waiting for spec2 to stabilize.
There is no "Datomic scale" problem. Datomic transactors are just another singleton service to deploy with your microservice pod. The underlying storage can obviously be consolidated and scaled independently. I don't recall what they're using, would guess it is pg.
You don't have to scale your postgres daemons with your microservices- each of which have their own transactor- which would be painful and out of the ordinary for an ops team.
Scaling datomic on top of postgres is no different from scaling any other microservice.
That all said, the architectural point that did sound painful in the presentation- and which is a common pain point in microservice architectures, not unique to what NuBank is doing with Datomic- was having to maintain an ETL for analytics purposes, to pull together all of the distinct microservice-specific Datomic-hosted data sets into a single uniformly SQL-queryable data set. The details of that implementation, and whether it made use of the new SQL interface supported by Datomic, were not discussed. But it smelled brittle and fragile.
yes
,each talking to a named "database" hosted by that one daemon.
No. That would imply distributed writes and break the single write serializability property of datomic. Think about it, transactors don't sync with each other(only one for HA but that's orthogonal).
To put it simply, multiple transactors can share the same storage but only one can write to a single database at a time.
I don't know of a reliable multi write system for mysql that doesn't make significant trade offs
At work we use multiple mysql servers to handle "scaling" so it's not a surprise to me that Nubank are using multiple Datomic servers
Unless you start getting into eventual consistency territory I don't think distributed writes are a trivial problem and targeting Datomic for this is a bit odd