What do you mean by “Event-Driven”?
martinfowler.com
martinfowler.com
The idea looks deceptively simple. If you need something to happen, you just broadcast the right event, someone else acts on it and everything is fine. You code it, it works in the demo, and you deploy to production.
And then you run across network failures. Missed events. Doubled events. Events that arrived while a database was down. Events that were missed due to a minor software bug. And so on. You layer on fixes, and the system develops more complex edge cases, and more subtle bugs.
Then you rewrite from a push based system to a pull based system, and all of your complex edge cases disappear. :-)
We planned for duplication of messages. Operations are idempotent. If the message already ran, we could rerun the calculation safely. With Rabbit the only duplicate messages were the ones humans sent to rerun a model that broke.
And why are they rebooting all the time anyway??
This is important. There is no "exactly once" system; you can have at-most-once where things get occasionally lost, or at-least-once where you resend things that you can't be sure have been recieved. The latter loses less data but needs methods for detecting or otherwise ignoring duplicates.
NFS goes for "idempotent operations", SMTP goes for message-IDs that can be used for deduplication.
I think doing "event driven" without an event log for critical system functionality is probably the anti-pattern you are describing here. with my event log, worst case scenario is i need to reprocess all the messages in my system to recover from all of the above.
That said, most of the things u mentioned can be mitigated by a decent event bus with a competing consumer.
ie. isolate your critical components (write model) into small succinct peices that use highly reliable message delivery techniques. once committed, broadcast to other queues where delivery isnt that critical.
Or you can just level up and go https://martinfowler.com/articles/lmax.html
All the options you mentioned are doing unwarranted smart work without context.
Alternately you can have some sort of confirmation protocol. It is easy to go too far, but the kind of confirmation/resend logic that turns UDP into TCP has very much demonstrated its value in practice.
Sure...if all the ilities weren't accounted for in the original design.
Async or out-of-process event systems have increased points of failure. If those aren't accounted for, then yes -- problems occur.
Doesn't make an anti-pattern, though.
Virtually nobody is able to account for everything in the original design. Convincing yourself that you got it right is easy. Actually getting it right is HARD. It is possible. For example zookeeper seems to have. But your odds of success are very, very low.
See https://aphyr.com/tags/jepsen for many, many examples of competent people who thought they had it right, being proven wrong. Over and over again.
> You layer on fixes, and the system develops more complex edge cases, and more subtle bugs.
"Accounting for" != implementing to address.
Failure is expected in distributed systems. If your architecture allows no room to address future failure, that's the problem.
Unfortunately, as aphyr proves over and over again, is that the guarantees we are given for how distributed software is supposed to work don't hold. Over and over again, across virtually every piece of well-known distributed piece of software. And my experience is that in house software is at least an order of magnitude worse.
However you think your software will work, it doesn't.
An example is inter-process communication.
See Erlang for life lessons on how to get it right.
Of course, if the events become irrelevant after the fact you don't need this, but only very few systems are like that.
This same thing applies to time-based events: your system should not assume that the process is always running, so if something needs to happen at exactly 9:00am (and it's not okay to just skip it if missed), it should be able to run anytime later with the same outcome as if it ran at 9:00am.
When I was first getting into this, it helped me to understand that events must follow one of the semantics: at-most-once, at-least-once and exactly-once -- with the trick that exactly-once is not strictly possible [1].
There are only two hard problems in distributed systems:
2. Exactly-once delivery
1. Guaranteed order of messages
2. Exactly-once delivery
-Mathias Verraes [2]
[1] http://bravenewgeek.com/you-cannot-have-exactly-once-deliver...[2] https://twitter.com/mathiasverraes/status/632260618599403520
Is there any receipe for this in general? Once you start with a model of a push architecture... don't you kind of have it as the only sane model?
You'll end up implementing some kind of pooling loop, and sooner or later voila, you've reimplemented an event loop, and now you have a badly ad-hoc implemented event driven system anyway.
The only way to handle a "naturally push based system" is to accept that this is the natural way for it to be, a find a declarative way to express it as a rules based system instead of tangles of imperative event handlers. But this is really hard! So you settle for the "push based system" or "event driven" system instead quite often...
Yes, there is a polling loop. Is this bad? Nope. You just have to understand the system you're building. Will you get data in a mostly consistent timing? Polling works. Is is sporadic? Try event driven. One size does not fit all. Event driven should not be used for everything and it's not in any way more natural than polling.
There's trade-offs either way. No way is "best"; it all depends on what you're doing. I communicate with equipment over TCP, that I have to poll. There's _no_ mechanism for push. You just work with what you have and do the very best you can.
I worked with a team that were making quite a complex multi-part network manager some years ago. Development was going to be superfast because the lead developer had come up with this new framework (red flag! red flag!) based on publishing events and then subscribers picking them up and acting on them.
The problems began when the subsystem running user interaction started to need replies to specific things so it could tell what was going on with specific user interactions. Then message timeouts where needed so that pieces could react if their response didn't manifest in time.
What this framework unintentionally did, I cynically worked out after a while, was implement UDP multicast over TCP.
What I was watching happen, each time messaging problems come up, was the reimplementation of TCP on top of that. That's about when I bailed!
I need to think about it some more :)
The backend implementation ideas involve REST, caching, replication, eventual consistency, CAP, and especially PACELC.
The continuous improvement ideas involve upgrading existing software projects from imperative styles to functional styles, phasing in publish-subscribe modules, growing codebases to be more event-driven, and planning for evolutionary architecture.
In practice, I see these areas need management help, such as funding and time for training, and help for teams creating infrastructure as code (IaC) to deploy many event-oriented modules as well as event monitoring and management.
Related links:
https://en.wikipedia.org/wiki/Command-query_separation
https://www.wikipedia.org/wiki/REST
https://www.wikipedia.org/wiki/PACELC_theorem
[1] http://www.paulgriffiths.net/program/c/srcs/hellosrc.html
[2] http://www.paulgriffiths.net/program/c/srcs/winhellosrc.html
UI events are typically of two kinds: user performed an action and you need to update the model / execute a command, or the view is asking for information and you need to translate it from the model. Both are fairly encapsulated; the larger operation doesn't extend beyond the scope of the callback, and if it does (e.g. showing a modal dialog), you can use a nested event loop to take care of event dispatching while showing the modal.
The events in a UI are also naturally scoped to the screen being shown. When the screen goes away, you can rely on those events not being dispatched any more. Events get to party on a shared state (the model, perhaps attributes of the view) but scoped within the view being displayed. It's all pretty flat. You also don't need to worry much about concurrency, unless doing async.
IMO this comes up a lot in HN/Reddit discussions about either term. Event-sourcing implies CQRS, but the reverse is not true.
> I'd love to write some definitive treatise [...] Sadly I don't have the time to do it.
The "Fermat's Last Theorem" approach to computer architecture :P
void Stuff::AddThing(Thing *thing) {
this->AddEvent(new AddThingEvent(thing));
}
void Stuff::AddEvent(Event *event) {
m_events.push_back(event);
event->Execute(this);
}
void AddThingEvent::Execute(Stuff *stuff) {
stuff->m_things.push_back(m_thing);
}
and maybe AddThingEvent is a friend of Stuff, or it's an inner class, or whatever you like.Maybe this is what he meant though.
If so, that means that instead of a CRUD interface like:
if(customer.getApprovalLevel() < 5){
customer.setApprovalLevel(5);
customer.setApprover(currentUser);
}
saveChangesSomehow(customer);
So you'll end up with an API that's a bit more CQS (no 'R' yet) with: if(customer.getApprovalLevel() < 5){
customer.promoteApprovalLevel(5, currentUser);
}
saveChangesSomehow(customer);
That way you capture each change-event at the right granularity with the necessary data. Finally, CQRS comes into play because you want to do something with that rich event data, such as eagerly populating a table: recentApprovals = getRecentApprovals();
if(recentApprovals.length() > 0){
sendApprovalAuditReport(recentApprovals);
}I don't know of any other generalized event-sourced OSes. But almost every principle Fowler describes in special-purpose event-sourced designs also seems to apply in the general-purpose case.
For example, the connection between CQRS and event-sourcing seems very deep and natural. The connection to patches in a revision-control system is also completely apropos.
When Fowler starts pointing at advanced problems like "we have to figure out how to deal with changes in the schema of events over time," you start to see a clear path from special-purpose event-sourced apps to general-purpose event-sourced system software.
Broadly speaking, one category of answer is: abstract the semantics of your app, until it becomes a general-purpose interpreter whose source code is delivered and updated in the event stream. Then the new schema is just a source patch. And your app is not just an app, but actually sort of an OS.
I've looked through its documentation, and while I couldn't make heads or tails of it, Urbit seems to be everything but an operating system. Espeically that it needs unix system to run.
Urbit is an OS in that it defines the complete lifecycle of a general-purpose computer. It's not an OS in that it's the lowest layer above bare metal.
Sorry you had a bad experience with the docs. Urbit is actually much stupider than it looks.
Unfortunately, premature documentation is the second root of all evil, so some of the vagueness you sense is due to immature code. It's always a mistake to document a system into existence -- that is the path of vaporware. Better to run and have weak docs, than have good docs but not run.
I'm not criticizing, I am exposing that I am part of the problem and lost, and I'd like to know how people handle this on their lives. I want to document more but there's always this rush to implement implement implement and not once, during some catastrophic breakdown of something, management wants everything solved quickly but they expect you to remember instantly of everything you did 2+ years ago :-)
PS: ... and they'll get mad if you say "you didn't let me document this damn thing"
- Alan Kay, from a discussion he had with Rich Hickey on HN (Rich Hickey disagrees with the idea that you need to send an interpeter)
If event-handlers are genuinely decoupled from one another, and there isn't a higher-level state machine that's being driven by events, then it can be an excellent way to structure logic.
Events-as-deltas that are durably stored and can be reliably replayed is another fine way to go for a system with checkpoints, auditing, history, rollback, time-series reporting and similar requirements. But like the article says, code will be much simpler if it deals with an eternal present and the calculation and application of the deltas is centralized in code that doesn't change often.
There's an isomorphism between building a bigger program out of event handlers, and distributed programming. Distributed programming is known to be difficult, and doing it at the message send level is IMO too low a level - one should build more expressive and composable primitives and work at a higher level. That's if it's warranted at all.
For both, whenever I read about em or I hear someone discussing em, it makes so much sense in my head! But later when I try to execute on the ideas, I can't make it "come together". It's like some puzzle piece refuses to click into place.
In the cases when I've tried hacking something together, I've always ended up with something that seemed more brittle than if I'd just gone with a safe relational database. With that said, I'm also completely open to the possibility that I've just been trying to apply the pattern in cases where it's not a good fit.
I'd seriously love to poke around in a real-world application showcasing CQRS. I don't expect something perfect, but getting a chance to look at what tradeoffs and concessions a bunch of engineers made would surely be incredibly insightful.
There is a 95% chance that event sourcing and CQRS were entirely the wrong patterns for your use case. If you're not dealing with a very complex domain (like, say, automatically orchestrating a global logistics or drug discovery pipeline with audit logs for regulatory compliance), that chance becomes 99.9%. If you're working on something that only you or a few people will ever work on, it's 110%. Just the tooling for a reliable event sourcing system would be a herculean task for a single developer and that's without even writing a single line of business logic!
I don't know of any open source real world examples of CQRS off the top of my head and the only successful ES/CQRS systems I've seen in the wild were projects that involved hundreds of domain experts and programmers, dealing with extremely complicated fields like biotechnology and electrical engineering where nonprogrammers spent about as much time looking at code as the programmers. Any good ES/CQRS project will contain a lot of the business's secret sauce so it's not something Google or Facebook would open source.
I don't know how real-world it is, since I'm still pretty much the only one who is using it, but you may enjoy checking it out for a slightly different view on ES.
It doesn't strictly fit the definition of CQRS at the moment. This is because I see CQRS as an implementation detail orthogonal to ES, but which dovetails nicely as they both work well with distributed systems. ES has immutable events (monotonically increasing), and CQRS implies a distributed system and allows for more effecient retrieval, leveraging ES's immutability.
But ibGib does provide what I see as an augmentation, or maybe an evolution, to ES. It has shifted over the years away from ontological "events" and now is more biological in approach. All of the data is in terms of ibGib, which is a flat database structure comprising `ib`, `data`, `rel8ns` and a `gib` which is a hash of the other three fields. The ES "event" analog specifically about what is replayed to hydrate domain aggregates can be seen in a specific `rel8n` called `dna`. This contains the code necessary to "rebuild" that ibGib immutable frame. This came largely from my dislike of the foundation of the entire river of the source of truth being necessary. With the DNA structure, it actually keeps a dependency graph, which ends up being a projection of the "entire" source of truth. This allows for nodes to "replicate" ("reproduce" would be a better term) as projections of entire graphs. These kinds of deep aspects are really neat!
But anyway, I digress. The server side is written in Elixir, which runs on top of Erlang's vm (the BEAM(!)). And you can visually interact with the structure, as the web app runs on d3. I have some demos on YouTube on the website if you just want to look at it, though I don't go into the DNA/ES aspect of it.
Anyway, I had to say something, since ES is a topic near and dear and so closely related to ibGib's current architecture. My design decisions could very well help you & others understand ES more deeply. Plus I've been enjoying getting to a point where ibGib is actually usable (and useful).
NOTE: I'm somewhat arbitrarily responding to you in this thread, as I used to play halo with an AceOfHearts.
Andre Staltz' lessons have been a great help for me. Most of them are for-pay though, on egghead.io.
With a simple shopping cart, you'd write 10x more code just for the tooling to get ES/CQRS up and running reliably.
I wasn't aware of the term 'event sourcing' at the time though.
In commercial systems, transactions as actions and events, recorded for business or legal reasons, are central. That includes significant "events" like purchases, sales, reservations etc. In that sense, they all are event driven.
Event-Driven as an architecture that is based on recording changes to a baseline state, should be applied only where really suited.
^may not be true for spooky actions
A couple of my favorites:
https://www.martinfowler.com/articles/injection.html
https://www.martinfowler.com/articles/refactoring-dependenci...