Data-Oriented Architecture
blog.eyas.sh
blog.eyas.sh
If you can't keep track of what services call each other, what makes you think you'll be able to keep track of who is writing out what data?
Reading how they did SOA (spaghetti of services with highly coupled interactions) I don't think putting that at the data layer is useful in the long run. They even suggest many services operating on the same data as a way to decorate the data as it haphazardly makes it way through this mess of an app. Their original problem was a poorly planned separation of concerns leading to cross service coupling. In DOA that still have it. Its just coupled directly through the datastore with no hope of cleaning it up without a massive migration.
And to be honest, the services are loosely coupled between each other that is a great benefit. It’s not a “spaghetti of services” because a good design will carefully consider them.
Plus it could give some good benefits for auditing trails too and consistency in data storage/archiving.
I don't want patterns where I have to be good. Anything is possible if you're good but how does the pattern help and protect the developer?
That said, as a developer, how do I know some data should be owned by one service or another? Should I notify other services when I change data? Where is the separation of concerns? Is the DAL a monolithic codebase that knows who to notify or what event to write when things are updated?
I like event queues but instead of using something like ZeroMQ they seem to suggest using a custom rolled table. If I want to add a new event listener such that I have more than one, do I have to know that at event write time and write n messages? Do I store message completion separate from the message itself? Does the DAL handle that by knowing who needs what? Aren't we back to thinking about all those service interactions but in the DAL? So much work is glossed over that I can only assume the author hasn't worked with queues or an ESB very much.
The ultimate question though is, why are these services even separate if they're mutating the same data? Just merge the services (or at least the operation in question) instead of rewriting the entire app. Way easier than orchestrating all of your systems through a monolithic DAL.
Because if you did that you'd have a monolith, and microservices are really fashionable right now.
Being fashionable is the worst possible reason for making technology choice, but sadly one which many developers mindlessly choose.
Every rdbms left on the stage was built with the assumption that lots of disparate systems would be coordinating within them.
I’ve yet to see a SOA system with half the tooling to support this. That’s before you get into the performance advantages.
If you’ve done a bad job coupling your data tier it’s because you didn’t know/follow best practices we’ve had for at least 30 years. Don’t blame the architecture for that.
Also assuming you are sane and give every service/user/etc their own user account into the DB, your DB logging will easily show user X did action Y.
Databases these days scale very well. PostgreSQL, which is a great OSS database can scale very well out of the box on any single system, and the size of X86 boxes are getting pretty giant these days(not to mention other platforms). Plus there are loads of 3rd party, but well supported, options for scaling past a single instance.
There are other systems like FoundationDB, etc that scale well out of the box, but I'd argue most people don't actually need to scale that large, 99% of us will never get to Google size and one can get very far on a single DB instance. Especially for new projects, scaling should be near the bottom of your todo list, until it starts to hurt, and then the general answer is, just throw money at the problem, since if you are having scaling problems, you better not also be having money problems or you likely have larger problems than scaling your DB.
Monoliths certainly can reach their breaking point for other reasons. But they can take you very, very far too. Depending upon your exact use case, a stateless monolith, Postgresql, and something like Redis can handle millions of daily users.
It's been a while since I've had to do this, about 10 years. The breaking point for our use case was around 6 million daily users. I would imagine you can take it further today, but haven't had the need to so I don't really know.
Multi-core CPUs are cheaper than ever, and terabytes of RAM are not uncommon.
The only problem is that to have an HA solution with these monster machines, you need at least 2 of them, and maybe a similarly sized test system.
Well that's exactly my point. The author is saying this architecture will save them and it wasn't the architecture at all.
- A key reason for splitting up a monolith into services is because collaboration becomes too costly with 100s to 1000s of developers working on the same code. Your build & test system can't handle the number of commits. Your code takes too long to build. Would the data-access layer not them become the development bottleneck in such scenarios, meaning that it doesn't scale as well as SOA?
- What are the advantages and disadvantages of centralizing in the data-oriented layer over centralizing through a network-related tool like Envoy, Kubernetes, or Istio?
- The database layer itself often becomes a performance bottleneck, which requires us to run a sharded database. Some companies go the extra distance by running in-memory databases, document databases, and time-series databases. In such cases, wouldn't the data access layer need to support federation, which is a hard problem according to database research?
- Is O(N^2) really that big of a problem? It seems like the problem can be reduced to something simpler: developers cannot easily understand which services communicate with one another. If that is the case, would a visualization tool be sufficient?
while they do not use the term "data-oriented architecture", many of the largest web companies use what is effectively this approach and have teams dedicated to building and maintaining a shared data layer. some examples:
google - spanner
youtube - vitess, migrated to spanner
facebook - tao
uber - schemaless
dropbox - edgestore
twitter - manhattan
linkedin - espresso
notably absent is amazon. amazon has taken the full blown microservices approach where anyone can do whatever they want. worth noting is that amazon is in a very sad place when it comes to data warehousing and analyzing data across teams/products/etc. while the shared database approach is strictly intended for OLTP use cases and explicitly not meant for OLAP use cases, having a common interface to all data and something approaching a data model makes it extremely easy to replicate all your data out into a data warehouse or data lake or whatever you want to call your system for your OLAP workloads. with the 'every service has its own database' model, each team has to be responsible for replicating their data to analytics systems, and that is usually not super high on their priority list relative to product features. this problem is magnified when people from a different team want to consume data from that team's product/service but the team producing the data has no incentive to make it available. in large organizations (including amazon) this is a huge issue for teams who mostly do analysis, reporting, marketing, and other activities where they primarily consume data produced by others.
You give examples of all the Big Tech having such shared DBs but that seems like more of a reason to not use that pattern. Good DBAs are hard to find and not many people choose to become DBAs anymore. Big Tech can hire the experienced ones since they can compensate them pretty well; most companies can’t. The shared DB therefore becomes a critical bottleneck to the business.
fortunately this type of environment is available today as a managed service in a few different offerings. gcp has spanner and vitess is available as a managed service on multiple cloud providers from planetscale.
In terms of development, no, the data layer code grows sublinearly with the with the size/breadth of the schema/data. The data access layer is not much more than a database (plus usually, to enable event-driven programming, some semblance of subscriptions/notifications when data in your query changes). But it's fairly generalizable, and doesn't depend on the size of the team or schema using it.
> Is O(N^2) really that big of a problem? It seems like the problem can be reduced to something simpler: developers cannot easily understand which services communicate with one another. If that is the case, would a visualization tool be sufficient?
It really depends on how complex the system is. At some point, a visualization stops being helpful. There's obviously room to simplify a SOA dependency graph to look reasonable, and many do this successfully. But DOA is another interesting option in the toolkit: turn the problem on its head and say: maybe there's no graph at all.
You can't actually remove complexity like this, just push it to the dev ops layer. And also it makes setting up a development environment a lot more difficult.
There are many strong points for this the general challenges are...
Databases don't/didn't have too much in the way of integration primitives...
Databases can become a performance bottleneck that can only be scaled vertically...
SQL language is not that great as an application programming language
EDIT: forgot the biggest one which is, it is completely up to discipline to produce any kind of separation between implementation details and public api since everything lives in the database. That's really the biggest challenge.
Hum... They have the best and most diverse set of integration primitives available. Services architectures (micro, SOA, and whatever) did never actually reach parity to them.
Your other points are good (DBMSes do scale horizontally, but it's not easy nor nice), agreed on everything. But they still do not beat the capacity DBMSes have for integrating stuff on most applications, so this is still a good paradigm.
My experience has been the opposite here are some examples:
Number of times I had to implement or maintain hand rolled queues in the database.
Number of times I had to implement a web service whos only purpose was to expose database data to the world.
Number of ETL processes I wrote just to handle some data daily data input from a third party.
Number of times I had to use comparatively complex SQL techniques to iterate over a list of rows because SQL is intended for set based operations not iterative processing ..
Number of times I had to do tedious text templating to generate HTML or XML or JSON or any kind of heirarchical data format that is easily consumable by a non database system.
That's what I am talking about
Which isn't to say this is a bad architecture. Just that database as integration platform has its challenges, a lot of them.
Many of the problems described are problems from SOA systems which are using a number of common anti-patterns for SOA systems.
Also the way this person describes DOA is prone to end up with a monolith in _data_ and a just the logic is not monolithic but might end up accidentally being quite tight coupled. (Through it does have some benefits).
Also even with DOA you can end up with internal state coupling between services if you do it wrong. It's harder then in some bad designed SOA systems but IMHO roughly as likely as a SOA system communicating with events.
---
- So use events if you do SOA (for inter service communication)! - Never rely on the internal state of another Service. - Make sure you don't send events to a specific service, instead "just" send events, then all services interested _in that event_ will receive it. (Make sure to subscribe for events independent of their source not services; E.g. use a appropriate event broker or service mech; Sometimes just broadcasting is fine; Oh and naturally storing the event and making that trigger other services works, too. At which point we are back at DOA )
----
DOA is not bad just IMHO misguiding. If you want to use it look at common problems with event based systems for vectors of potential problems wrt. accidental internal state coupling as many of this will apply to DOA, too (if your system becomes complex enough).
Note that I don't mean all problems caused by combining eventual consistency + high horizontal scaling with event systems. Sadly this is just very often mixed up.
Give me a minute I will try to find at least some of the sources, but don't get your hopes up.
- https://www.youtube.com/watch?v=STKCRSUsyP0
I wasn't able to find any other talk I watched back then or any of the stuff I did read, but I then last time I looked into talks and reading material about this was ~2.5 Years ago and while a lot of new software and tooling was done since then the principles didn't change. Try some of the other GOTO; talks about it if you like listening to talks, they tend to be quite good.
This was interesting as far as I remember but not what I was looking for:
While this seems closer to SOA, the key difference here is that single Type or Table can still have multiple produces (of non-overlapping records). In a trading system, you'd have producers of RFQs from makretplace A, B, C, etc. but for each single row in that table, the same service "owns" it. So you still get the benefits of not caring about the DAG/callgraph or knowing about the individual service that calls you.
Locking might be hard if you're doing some transactional change. Those become harder to do. But the half good news is that shifting to an "Event-based" programming mindset might mean you run into less of these.
But yeah, there's a whole new set of drawbacks.
Won't having several independent (and maybe physically remote) tables, one per service, solve the problem better? You can still `union all` them for analytical purposes.
But sure, you can probably implement that with views etc. also.
This could get very messy at scale.
IMO one of the reasons CRUD is so hard now is that the schema is different at every layer of the product
Slight differences between layers are necessary for permissions / privacy, but there are probably better ways to get that done than to reimplement the schema at every layer.
- I think a major problem is in the difference between the way you can layout thinks in the storage layer and your application.
- Another is where to evaluate correctness (e.g. in service checks + DB-system constraints, etc.).
- Another major think is that different actions/taks take different slices of the same data. One area where dynamic typed languages can have a clear benefit.
- Different actions/tasks in the same system work better with different representations.
---
What I currently think is helpful is to:
- learn from Entity-Component Systems (from Games) for slicing data of the same entity. (but can lead to problems with consistency across slices, transactions can help if doable).
- I would love to have a DB which can somehow do algebraic data types/sum types/tagged union/rust enum (all different words for roughly the same concept).
- Be very strict about preventing coupling of internal state, mixup of service responsibilities in logic/endpoint _and_ mixup of this responsibilities in data. Which e.g. means that e.g. you avoid any FK between schemas owned by different services, even through it often seems usefull at the beginning.
- Have a _system wide_ schema for all entities split up into small slices which if combined together from the entity and only use that schemas in the system no service specific schemas.
- Specify this schema somewhere, preferably generate data types and similar from it, preferably have some opt-in strict schema validation (enabled during part of testing).
- Use events for communication, use the schemas from above here, too.
- Have a well defined way to represent "patch"/"update" queries, I have run to often into the problem with JSON of "reseting/deleting" a values vs. just not changing it (null value vs. field not given) and/or nested optionality.
I'm aware of Esper[1] which tries to do this. And maybe/arguably Firebase (?)
One of the critical improvements to SOA is DDD (domain-Driven Design) where context matters and boundaries should include the service and its data.
Data oriented architecture a bad idea. Period.
Data storage should reflect a bound context and its domains. It could be relational, graph, document, or key/value.
Putting all data in one place just because it’s convenient is ignoring your business capabilities.
DDD proves that aligning your technology to our business, reducing complexity, having real boundaries, is the way forward.
https://www.amazon.com/Data-oriented-design-engineering-reso...
Also, don't just try to watch conference talks and read blog posts to understand it. It leaves too many gaps and a fuzzy understanding.
DOA is often implemented taking advantage of event-driven programming, pubsub, and message passing, which as generalized practices were not as prominent when SOA began, imo.
[The Pure Function Pipeline Data Flow v3.0 with Warehouse / Workshop Model](https://github.com/linpengcheng/PurefunctionPipelineDataflow)
1. Perfectly defeat other messy and complex software engineering methodologies in a simple and unified way.
2. Realize the unification of software and hardware on the logical model.
3. Achieve a leap in software production theory from the era of manual workshops to the era of standardized production in large industries.
4. The basics and the only way to `Software Design Automation (SDA)`, just like `Electronic Design Automation (EDA)`.