CQRS: Command Query Responsibility Segregation (2011)
martinfowler.com
martinfowler.com
I fail to see how CQRS is a factor in increased complexity. After all, CQRS boils down to separating the interfaces for queries/reads/immutable operations and commands/writes/mutable operations. Unless you're bolting on unrelated concepts, like event sourcing, then CQRS is not a significant source of increased complexity in non-trivial applications.
Can you shine some light on which aspects led to higher complexity in your use of CQRS?
I mean, look at GraphQL, which essentially makes it trivial to implement "plain" CQRS because there are always explicitly separate operations for mutations and queries, and it's easy to have separate object definitions for these operations (this is in direct contrast to many REST architectures, where your verbs - GET, PUT, POST, DELETE, etc. - are fundamentally operating on the same objects). However, underlying in the actual DB there is just a single representation.
But I think when people talk about CQRS these days they are really talking about the underlying data being different (e.g. a log of write operations vs a repository of queryable objects), and that usually implies something like event sourcing, and with that model you have a ton of complexity involved in keeping the different "sides" in sync, and since many update operations also involve queries (i.e. only do this update if something else is true about another object) you get into transactional difficulties as well.
FWIW, these days I usually do go with GraphQL over REST in my APIs, partially due to some of the advantages you get from separating the queries and mutations. But like you said, I prefer to ultimately store the info in a DB in one consistent format unless it's absolutely unavoidable. Sometimes it is unavoidable, but like the article mentions you can frequently get away with having a "Reporting Database" separate to your "Primary data" database.
Once you start spreading your data across multiple databases and storing it in different structures managed by different apps with different APIs, suddenly getting to even eventual consistency can quickly become a non-trivial problem, which itself is usually solved or mitigated by adding more complexity.
Sometimes this complexity really is required because of incredibly large scale, large/numerous engineering teams working on one project, or unique business requirements - but my original point was that for most cases, particularly new products/projects that aren't going to have 10M+ daily users from launch, you should keep it simple until you know for a fact that it needs the complexity of a full CQRS + event sourcing system.
Not really. Event sourcing is typically implemented with CQRS, mainly due to how ES relies on streams of command messages and as CAP-related requirements already require mutable and immutable commands to be processed differently. However, CQRS has zero requirements other than segregating interfaces.
Hell, in languages like C++ you already get CQRS out-of the-box by following const-correctness.
What do you need to add?
1) Internal design guide/rule/law system so everyone understands and abides by. i.e. "C" command services should NOT talk to other command services, only to query services. But query services can only talk to other query services. You need service mesh and discovery. A bit complex.
2) Sink or aggregator. It's not efficient to have a query service listen to both events from message bus to store in DB and serve requests. It's best to have dedicated sink service(s) where data is streamed to database(s) with certain level of guarantees.
3) Rules engines or workflow system. Since command services cannot talk to one another, you'd need to orchestrate actions among them, that's what workflow systems are designed for. They connect to both query and command services and make sure long or complex tasks get done without the need to do hacky stuff!
4) Sanity! The moment you go with CQRS + Sink + Workflows, it'll become easy to feel overwhelmed and loose control. Start with a small set of MACRO based micro-services. Jam pack all commands into one service, same for query and sink and slowly dissect into multiple smaller services over time as workload grows and you need scalability (chaos monkey). This way you have only ~4 microservices to manage, well 5 if you need an API gateway.
5) With all the services and probably technologies, you need a contractual way to communicate. You need a system where CODE is the DOC! You need Protobuf or similar to design your schemas and api. You need GRPC because of protobuf...
Could you elaborate on this? We’ve deployed a view aggregate services that we expose through a HTTP interface, but you’re suggesting to stream this state to a database instead? Is that so that you can avoid the whole snapshotting trouble?
Sink/aggregators listen to events and, as per instructions, store payloads into db for efficient storage or optimized querying. The query service is only connecting to that db to execute queries. It's possible to have multiple dbs that store the same exact data but for different purposes. i.e. sql for normalized data or acid transactions and nosql for fast map reduce ...
A properly designed system, treats dbs as disposable. Meaning, delete any db at anytime and just play back events from the beginning or from the last snapshot and rebuild a new db.
Unfortunately making message buses the single source of truth is not that easy, and scary. Kafka is an expensive family relative, you better make sure you have the $$$ & the initial configuration is near perfect but still streaming requires some experts to help you get it right.
Nats Steaming is good alternative to Kafka, I use and recommend it personally, but not perfect at high scale, like billions of events/day. Jetstream, successor to Nats Streaming, is coming but it's barely in preview stage and would need another year or so to be qualified for prod usage.
We have implemented our event store on top of PostgreSQL. I’m not a big fan of using Kafka as an event store.
You are totally right - the CQRS pattern alone is not enough. In my experience, most shops that go down the CQRS path, also tend to make use of other event driven patterns like event sourcing.
As with most things, there's pro's and con's and the most obvious con here is increased complexity in exchange for higher reliability and scale.
Your advice is great, especially deciding on a structured message format such as protobuf. I would 100% avoid JSON schema as it's possible folks will forget to fill out fields. Protobuf has excellent support by now and there is a ton of support tooling for it. I'd be hard pressed to choose Avro at this point (unless you're a Java or Kafka shop that has it already in their ecosystem).
In addition, if you are already on CQRS, I would advise looking into fully embracing event driven - by utilizing an event bus of some sort (rabbit, kafka, eventbridge), you could make message passing completely async (and avoid having to use gRPC, which is another layer of complexity, as you mentioned).
A good friend of mine once said that in order to successfully do event driven arch, you must be OK with eventual consistency and I think that lies at the heart of this. If you are OK with eventual consistency, you understand the burden of complexity you're bringing on.
MQTT is lightweight and fits nicely into IoT sources, backend streaming and even scalable frontend distribution to browser or mobile clients. I recently did a quick'n'dirty PoC with paho.js and D3.js for live charting in the browser.
In a pub-sub event-driven world, CQRS just means having distinct events for: Commands to do something; Query matches (CEP-style); and other Notifications that something has happened. Queries are compiled and unfolded into streaming operator graphs, delivering match events to the client. Queries and their operator graphs can be factored to share common sub-expressions (like the RETE rules algorithm).
A good start-up idea would be MQTT 'appstreamer' framework, with model-driven message formats, validation and APIs, together with pluggable engines for CEP queries, rules, workflows, constraint solvers and other business logic.
I started by sketching out a plan on how I would rigorously apply DDD & CQRS to a problem, but in the full-knowledge that it would likely be overkill. I then dropped out components which I felt were unnecessary for my particular situation.
The result has been something fairly lightweight, without too much boilerplate, and very reliable. The clients concerned have certainly found it to be a rock solid part of the core business process.
The need for this did not come from any performance requirement, but because I felt there had a to be sensible way to architect SME business software. Especially in a startup environment where business needs and processes were often changing. I felt that DDD allowed me to actually model businesses processes in software, rather than requiring business needs to adapt to the software I was developing.
This may well be old news to enterprise-scale developers, but as a freelancer coming from the smaller-company end of the spectrum this was a big deal.
As part of this process I developed a Python library [1] to facilitate writing evented/RPCed systems (event sourced or otherwise), and wrote up some very rough architecture tips [2] as part of the documentation. They are pretty high level, but may be an interesting (somewhat opinionated) starting point.
[1]: https://lightbus.org
[2]: https://lightbus.org/dev/explanation/architecture-tips/
Edit: Link update
The answer straight from the horse's mouth (Greg Young): https://web.archive.org/web/20190211113420/http://codebetter...
CustomerService
void MakeCustomerPreferred(CustomerId)
Customer GetCustomer(CustomerId)
CustomerSet GetCustomersWithName(Name)
CustomerSet GetPreferredCustomers()
void ChangeCustomerLocale(CustomerId, NewLocale)
void CreateCustomer(Customer)
void EditCustomerDetails(CustomerDetails)
goes to
---------
CustomerWriteService
void MakeCustomerPreferred(CustomerId)
void ChangeCustomerLocale(CustomerId, NewLocale)
void CreateCustomer(Customer)
void EditCustomerDetails(CustomerDetails)
---------
CustomerReadService
Customer GetCustomer(CustomerId)
CustomerSet GetCustomersWithName(Name)
CustomerSet GetPreferredCustomers()
---------
and that's it. No Task/mediator architectures, no event sourcing, nothing.
With CQRS, this is pretty trivially implemented in GraphQL as GraphQL explicitly separates queries from mutations (this is my favorite blog on the topic, https://www.apollographql.com/blog/designing-graphql-mutatio... ), but for some reason nobody talks about it as CQRS when they implement it like that. In fact, in the past 5ish years I've only heard about CQRS in an architecture that also used event sourcing.
see also https://cqrs.files.wordpress.com/2010/11/cqrs_documents.pdf
1. CQRS+ES works great, if you implement it carefully, which requires discipline. We ended up creating our own framework (ugh) to make it easy to go down the right path.
2. There are trade-offs, of course.
3. You are only one smart agile thinker-in-a-pinch away from stepping away from the discipline and creating really tricky complexity and far more negative trade-offs.
4. There is very little market demand for this approach among the typical enterprise business solutions we work on; though as I understand there are some sharp peaks in certain areas of financial services.
(The odd thing is... we have financial service customers, just not in the particular domains where there is demand for CQRS.)
The idea is that for banking, it is not enough to just get the _current_ state - the more important thing is how someone _reached_ that state.
Adding history to transactions is not new - so rather than bolting on a history/audit mechanism, you knock out both - a higher resilience, distributed system + built-in audit/history mechanism.
I love this sentence. That seems to be a common pitfall, regardless of framework or pattern, which ironically puts a hero aura around the person who did it.
We also went down this path (CQRS + ES) a few years ago at my old company. Implementing it was fun, but it did add quite some complexity. Granted, it was very scalable.
Interestingly enough we used it for an enterprise messaging app, although I left for other things before it was released. I wonder how that codes looks today.
> Despite these benefits, you should be very cautious about using CQRS. Many information systems fit well with the notion of an information base that is updated in the same way that it's read, adding CQRS to such a system can add significant complexity. I've certainly seen cases where it's made a significant drag on productivity, adding an unwarranted amount of risk to the project, even in the hands of a capable team. So while CQRS is a pattern that's good to have in the toolbox, beware that it is difficult to use well and you can easily chop off important bits if you mishandle it.
It's not a fit, imo, as a system architecture. It's more a subsystem design where you apply it to something which has many writes and a eventual consistent read-store, which can of course be built from the write-store. That doesn't really matter.
This can aslo easily be done without adding event sourcing, and the two does not per se go hand in hand.
The key point being that you sometimes cannot build a view "fast enough", based on data which is constantly entering the system. So the system with in the CQRS design, which handles a query is very important to design based on this.
I see it as a pattern that really helps you to perform on the read/query load due to the almost impossibility to make a query perform over a massive amount of commands/writes
My sense is the crux is that you make explicit with CQRS is that you can offload reads to read replicas as long as you're OK with the read replicas being a bit more behind in the writes. Which gives you greater scalability like you say.
In our case, queries state aka projection were asynchronous which made it very flexible and fast to work with. Projection could live anywhere, within the web server, redis or just database.
In practical terms where I had to use it we ended up with the query portion being fulfilled by a materialized view that organized & aggregated the data in a way to make the queries feasible in milliseconds instead of minutes or tens of minutes.
This pattern can be very important with respect to enormous data sources in the cloud but can even be useful in smaller data sets in "on prem" applications.
Like anything though if you try to apply it everywhere that's just bad architecture.
We use Lagom framework to implement CQRS+ES over Akka Actors and related technologies. There's enough guardrails/guides within the Lagom framework and documentations that it's hard to screw up too badly. And IMO I think the different pieces fit together very nicely!
Overall the tradeoffs in system complexity vs scalability isn't too shabby.
CQRS usually means that Commands (modifying the system) are handled differently than Queries (checking the current state of the system).
Of course, you can absolutely create a REST API where there is one set of resources that you can modify but not read (Command resources), and one set of resources you can read but not modify (Query resources). But that would definitely not fit the general idea of what people expect from a REST API.
Also, such an API would be relatively hard to make HATEOAS-compliant, for those that care about "ultimate"/"true" REST.
I don't know if it would be any more idiomatic for GraphQL than it would be for REST.
In the end it's about whether the term tells you something about the architecture, and in neither of those cases would I gain useful knowledge of the system from someone saying "it's CQRS."
A better litmus test is probably something to do with enumerating possible commands. So a GraphQL API with a specific set of mutations could be a CQRS design. But if your mutations are just CRUD spelled differently, it's probably not.
As a sibling comment said: if you implement GraphQL with DTOs that look exactly the same for reading and writing, sure, you won't have CQRS, but it seems like the structure encourages you not to do that since you explicitly CANNOT use the output types as your input types in a GraphQL schema. Besides, real world usages seem to mostly treat it in a CQRS-style way instead of just CRUD operations.
Not really. REST just models everything as resources that are targetted by requests. REST states nothing about how rewuests should be segregated by commands and queries.
If you go fully event sourced, use gateways: https://martinfowler.com/eaaDev/EventSourcing.html
https://medium.com/fiverr-engineering/fiverrs-microservices-...
https://medium.com/fiverr-engineering/fiverrs-microservices-...
I have never used Google Spanner. But my understanding of Google Spanner database is it was designed from scratch so that reads never will slow down writes so it scales out of the box.
"Using Event Sourcing and CQRS with Incident - Part 1"
https://pedroassumpcao.ghost.io/event-sourcing-and-cqrs-usin...
If you design your system around idempotency you should be able to replay events and get the same result.
Not sure it was what you wnted to hear?
As the previous poster mentioned, if you are are able to handle events idempotently, you should be able to ingest events, regardless of the destination - monolithic or distributed.
With that said, in a monolith - there is still a chance of missed events - deploy, bugs, etc. and at that point, you'd still want to employ replay to get your system back into a good state.
Your initial question was about missed events - replay is still the answer and if you design your system with idempotency - you will be able to reingest events (even if they're dupes).