Seven Commandments for Event-Driven Architectures
rjzaworski.com
rjzaworski.com
---
I've not really used Kafka, so couldn't comment on that. I did some work for a customer that involved multi-DC microservices with isolated databases (ie one DB per DC). We used event sourcing with vector clocks to do manual reconciliation of the databases including during partition. Reconciliation involved custom logic depending on the event type, so not sure how a transport mechanism like Kafka would handle that.
A talk about the approach is here: https://youtu.be/hYVh8PbbeJw
I get to vector clocks at the end.
Edit: to add on about Kafka, it guarantees a serializable ordering on the incoming messages and is driven internally by a vector clock. This may not be suitable for applications, but the throughput is high and allows multiple subscribers to get a consistent ordering.
Vector clocks I typically want to stay away from as far as possible.
What do you do during a network partition? Accept writes that you’ll throw away eventually?
Imagine a customer has £100 in their account. System partitions. Customer withdraws £70, hitting one DC. Customer then hits the second DC, this time withdrawing £50. Each DC thinks the transaction is valid, and so serves it.
Later when the partition is restored, events are played back, and divergent history is detected via the vector clocks - the two withdrawals are not causally related. Remediative action can then be taken.
Transactions prevent bad things happening, but require CP semantics. Eventual consistency allows AP, allows bad things to happen, meaning you have to be able to detect them and clean them up later.
> Remediative action can then be taken
Sounds expensive and error-prone; taking a “read only” outage makes more sense in many use cases
Wouldn't it be useful to send both, since if the expected balance does not match up when the event gets processed, that must mean something was processed out of order?
[1]: https://rjzaworski.com/2019/03/7-commandments-for-event-driv...
This way you keep two separate, simple mechanisms to deal with two problems, rather than getting one more intricated problem.
Even if you "have a policy" people will ignore it. Better to make it explicit.
For instance, if you're a payment processor in the EU, and a withdrawal was done in the US, you may not have received it yet, so you have an incorrect balance. If a withdrawal is made in the EU, and you add the balance in the event, you will have mixed a perfectly legit observation with something, which is not an observation but an inference, an artifact processed from your incomplete set of observations.
Trust is something that changes over time. When your EU platform was implemented, maybe your US platform didn't exist. But the difference between observations and inference will remain true.
The reason a lot of games have frame-perfect bugs / exploits / glitches is because they use an update-driven architecture. So its possible that for one frame after certain things happen the game is in an inconsistent state.
The author talks about general cases, and while respecting his rules in your particular situation is not going to give you much, it isn't going to cause trouble either. On the other hand, if you try to disrespect these advices in a different problem, you may run into huge problems, because the simplifications you made don't hold.
Lets say you have a character doing some animation clip and at the end you want to run some callback.
That character could be killed by the player or the player could quit to the main menu or any number of things.
Now your callback is stale. You have to deal with cleaning it up, ensuring it does actually fire if you need it to, and handling any issue with the now deceased character.
Most of the points in the article are about using an event driven architecture in the world of distributed systems, where the messages triggering events are delayed by a certain amount, can be lost, malformed, or different to what the receiver expects.
None of that applies to a single program running on one machine. Thats why using an event driven architecture for a game is so great. I never have to worry about most of the stuff in the article. Also (as opposed to the more common Update driven model) if I do a good job its possible for my games logic to be flawless, it can be impossible for it to be in an illegal or weird state. So an event driven architecture in a game is fantastic for that reason, and the article is quite useless when it comes to event driven architecture in games. I'm sure its fine for event driven architecture in distributed systems, but the title is certainly too broad.
Like with many of the articles posted here, CRUD / web devs forget there are other fields of programming.