The Clean Architecture (2012)
8thlight.com
8thlight.com
Here's where the story breaks:
- Most app's domain objects and rules are small. You could describe them on a few sheets of notebook paper. The technology implementation (e.g. database or http logic) is a much larger portion of the codebase.
- In most teams, especially startups, the domain rapidly evolves. The technology stacks (databases, SQL, etc) are quite stable and have already been heavily abstracted for reuse over the years. Using these proven technologies and excellent existing interfaces is how you go fast.
- Technology choice isn't just about code interface. Each comes with its own assumptions of system behavior and theory of operation. Plugging a different implementation into your storage adapter interface is the least piece of work you need to think about unless it's nearly a 1:1 swap like postgres -> mysql.
So following a strict clean architecture approach will probably pour concrete around something changing all the time and create friction and extra work in using excellent available technologies.
Instead I'd rather follow pragmatic design guidelines:
- The concept of each component should be clear
- Sensible responsibilities for each component
- Abstractions should serve the composability of your existing design; introduce adapter layers if you truly have more than one implementation
- Keep it simple; optimize for reading the code end-to-end and easier refactoring
These aren't inconsistent with Clean Architecture but probably more productive than adding religious rules.
That could be the right choice in some situations -- it certainly worked out well for Salesforce and allowing extensibility in their model. However if you want to build an app with consumer usability you'll be putting a lot of unique development effort into your view layer.
That not true. Poorly understood business logic is where most of your code is. Once you create and depend on business object it is really difficult to change. Because you cannot change you end with writing more code...
Your code will largely reflect organisation and process. Only way to have "Clean Architecture" is when business owners are fully committed to project and willing to adapt/change organisation. But because it is easier to change code than people, we end up with multi million LOC projects in COBOL.
What this blog post is for enterprise managers that believe in bullshit graphs and layers. Technology is important because it will guide you towards solution. People selling ideas about "You can swap out Oracle or SQL Server, for Mongo" to my corporate masters should be hanged. Next year I will have a project to change database: CP to AP thanks to people like author.
> 4. Independent of Database. You can swap out Oracle or SQL Server, for Mongo, BigTable, CouchDB, or something else. Your business rules are not bound to the database.
MongoDB vs Oracle.. Ok architecture astronaut, you're gonna have a real bad time when you swap those out for each other.
Anyone who wants "database neutrality" (least useful common denominator), vs leveraging the awesomeness of pgsql, mssql, oracle, should just use flat files. Or mysql.
I write mine in raw sql normally.
I currently agree.
There was a time when the database engine was the "app server". Before client/server & ODBC.
I miss those days. That strategy is overdue for a comeback.
A database is a perfectly good application core, if you application primary about creating/updating/deleting records.
If your application is business workflows, and complicated business logic. Then this is something to be considered.
Almost foolproof.
If you tie your model to your database, you might be out of business sooner that you think. Just ask any of the oracle customers who paid $$$ to migrate away from that license fee sinkhole.
> Anyone who wants "database neutrality" (least useful common denominator), vs leveraging the awesomeness of pgsql, mssql, oracle, should just use flat files. Or mysql.
Alternatively, simply invest a little bit more time into proper design and implementation. In one of my products I am "leveraging the awesomeness of pgsql" but switching to another storage engine is still a matter of one or two hours. After all, it's only about 200 lines of code: The connection itself and the adapters for the query engine (and no, it's not loosely coupled - it's strongly typed and my compiler inflects on the model and the constraints imposed on it by the storage backend).
Lots of developers think database first, then build on top of that. That's good for simple crud apps. Creating/updating/deleting stuff is what databases are good at.
For complicated apps you should think domain or business problem first. Model your domain, make it good for reasoning about your business issues, and then afterwords think about persisting that to a database.
You domain is your application core, database persistence is a technical issue outside of that core. You write an adapter to persist that domain to whatever database you want. It just so happens this means you can write many adapters for any database to persist your domain to it. Utilizing whatever advanced features your database has to speed things up in that database adapter.
Yes, you should design your work first, and figure out what type of data you need to store, and how it should be stored.
Just about any application that uses Oracle cannot be "adapted" to use mongodb. It's an entirely different scope, different persistence model, different features... If we were comparing mysql vs postgres, there would be novel-sized comments about how you can't just switch between them.
mysql : postgres :: bicycle : recumbent bicycle
mongodb : oracle :: tricycle : underwater nuclear submarine
They're just soluctions to entirely different business problem domains. This kind of "if the architecture is clean enough it can run on your microwave" attitude is a disease.
However there are plenty of applications why there is no technical reason why they can't work on both though.
Writing application for Oracle I know that I have ACID and transactions, Writing my data access layer I will take advantage of this. I will also consider locking issues and add some kind of cache like Redis to boost read performance.
Writing for MongoDB I have schemaless data and different consistency guarantees. Because my data objects do not have consistent schema my access layer need to be flexible enough to deal with object of different shape. I would probably use dynamic languages like JS/Ruby.
Database choice will guide design and implementation data access layer. It is non trivial to change this. Even MySQL to PostgreSQL can take a lot effort when you have a lot live data.
You can implement it however you want.
How you will create an adaptor that support transaction for both Mongo and Oracle? Answer: You don't because Mongo do not support transaction.
This is like fridge and bookshelf. Both provide storage but cannot be used interchangeably. What would be adaptor for bookshelf to store meat?
Most applications copy a string from here and paste it over there. Input, processing, output. What we used to call data processing.
Add some defensive programming for sanity. Validation rules, "schemas", type systems.
Favor composition over inheritance, a useful programming language over "dynamic typing" (aka type hostile).
Extra credit for logging, monitoring, auditing, alerts, rolling deployments.
Life time victory achievement bonus award for setting breakpoints (debuggable) and easy reproduction steps.
---
Instafail if you use mappings (eg ORM), observer/listener, factories, singletons.
Wait, why?
As far as I can tell Factories have been replaced with DI frameworks or just hand rolled DI.
Singletons aren't so bad if you're working with an object oriented language and the singleton is just an instance of a service which holds no state and has methods that operate on a limited number of data types. At that point it's essentially just a namespace for a functional library.
Observer/Listeners ... Not sure here unless the commenter is advocating message queues or eventing.
In my experience I've never seen an enterprise company change databases over night. Uber perhaps being the exception to the rule. If you're choosing the target platform that your software runs on you should exploit that platform for all it's worth to get the best benefit from it. Yet I've seen plenty of software shops that insist on writing/running heaps of code to abstract away the database server on some imagined future point where someone decides they're going to run it on Mongo now instead of MySQL.
Long term, the result was a 45-minute test suite which spent most of its time setting up records in the DB, and which could not be parallelized except by adding more instances of the real DB (otherwise the tests would interfere with one-another's assertions about the DB state).
To provide a counterpoint though, an advantage touted by advocates of complete decoupling from the DB is that your UT suite can run in O(one minute), rather than O(ten minutes). E.g. see https://www.youtube.com/watch?v=tg5RFeSfBM4.
I'm not 100% sold on this approach yet (separating your domain objects entirely from the ORM wrapper is an uphill struggle), but it's interesting.
Also since we strive to make our databases upgradable, it's important that the actual schema update scripts themselves are tested and used directly.
In fact i just write a domain model, solve the problem, then write database adapter to persist the domain model. I always think about modeling the domain first, then persistence is an after thought. The persistence adapter can be done however you want, raw sql orm, nosql. So it's not really "abstracted away", it's just no the focus of the application.
Everything just plugs into the domain model.
I've seen too many programmer bugs to trust putting business logic outside of the DB. Separation of concerns here too just on different lines of concern.
That's a rather sweeping statement, that many in the industry do not agree with.
https://en.wikipedia.org/wiki/Domain-driven_design https://martinfowler.com/
> If there's a business rule for how that data is handled then it's most consistent to define it in the [db]
To be clear, you're proposing that all business rule validation should be implemented as stored procedures in the DB? One of my domain model aggregates consists of thousands of lines of code just enforcing business rules and constraints. Am I to put that all in the DB?
And some who do
https://dataorientedprogramming.wordpress.com/tag/mike-acton...
I'm not going to appeal to authority here. I'm speaking from experience. We all know how source code gets over time with hundreds of programmers working on it. A clear, consistent specification of your data model, rules, invariants, and transformations is far more valuable than the abstract-soup of trying to model your business domain in source code. I think we can all agree that the less code there is to understand then the easier it is to verify it is correct.
Verifying the requirement that "when record A is written to the database then B is appended with the delta change if such and such is True" is guaranteed at the database level along with all of the other constraints on those relations. If it's nested in one of these rings behind an abstract factory somewhere it's harder to verify.
> Am I to put that all in the DB?
That's where I would start. But I'm not you and I don't understand the problem you're trying to solve.
My original point was that abstracting out the platform if you're not really concerned about switching platforms is a form of premature pessimization and a source of errors. If you control the platform, target the platform and don't bother with the abstractions.
I find most SQL languages are not particularly great for general purpose programming. Especially tooling for testing.
That's why most mature RDBMS servers ship with at least one procedural language. Though it'd be nice if there was an option to use OCaml or Haskell in PostgreSQL.
I'm not suggesting to throw out all your code and build your entire application in SQL. I'm just saying that if you control the database, use it, exploit it and don't abstract it out unless you absolutely have to (because you need to ship your application on-premises to clients who may run MySQL servers and others who run Oracle.
That is a big if in the enterprise. It is getting better, but for most enterprise apps I have worked on, there is a db team that owns the db and you have to go through them for all changes.
Perhaps db abstraction can be thought of as an instance of Conway's Law.
Bugs are just as likely to occur in the database programming language as they are in application logic.
In fact i think tooling around unit testing is far more mature in general purpose languages.
I've seen databases like that and similar teams that didn't take care with their change management.
Either way it's never pleasant to work with such systems.
When I build web applications on top of this they have less to do. They literally parse HTTP and shuttle data. No big MVC framework needed. When the data model changes we change the data model. In the database. And we use the abstraction facilities in our server to keep the public schema clean.
Premature abstraction is just as dangerous as premature optimization. Maybe even worse in my experience.
Lots of developers think database first, then build on top of that. That's good for simple crud apps.
For complicated apps you should think domain or business problem first. Model your domain, make it good for reasoning about your business issues, and then afterwords think about persisting that to a database. You domain is your application core, database persistence is a technical issue outside of that core.
You first step in implementing in a feature is understanding the business domain with the help of people who know it well, then modeling it in the core of the application.
I'm currently working in the retail and warehouse distribution domain.
Essentially all big ERP packages follow the model of bunch of relational tables directly exposed to user with bussiness rules and processes as an afterthough. One can say that this stems from historical reasons, but unforeseen usecases that have to be somehow handled right now are also significant reason.
You end up creating a really beefy single database, with tons of ram, infinityio etc
We have a bunch of different DBs for specific purposes. Some have to be super fast for reading, some just store streams of events, others need to maintain consistency.
I see our customers' data as our core and any code we write as a liability. We're constantly looking for ways to reduce our liability and ensure our customers' data will always be consistent. We push all of our business logic down to the database layer where the RDBMS server is responsible for ensuring its consistency and integrity.
While we're not perfect we do have some business logic in our web processes. However for the most part the job of that component is to parse HTTP queries into queries on our public schema. If we wrote business rules into our software it would be very difficult to verify them, search for them, and keep track of how they've changed over time. It's also too easy for a programmer to make an error which could cause our customers' data to enter into an inconsistent state or worse. We avoid that as much as possible.
Too deep a stack will lead to code "scavenger hunts" when you want to figure out what something does.
I prefer Gary's explanation of it because he jumps right into the meat of the problem and shows code that models this architecture. The video doesn't require you to know Ruby to understand what he is saying, just as long as you know some basic testing phrases; Mocking, Stubbing, Etc, you should be able to follow along.
https://www.destroyallsoftware.com/talks/boundaries
Note that the Functional Core architecture makes an additional restriction on top of the Clean/Hexagonal architecture, namely that the core should be functional; the OP doesn't make such prescriptions on how you implement the Entities and business logic (though it doesn't discourage you from doing so either of course).
"2. Testable. The business rules can be tested without the UI, Database, Web Server, or any other external element."
If that doesn't scream functional, then I don't know what does. :)
http://degoes.net/articles/modern-fp http://degoes.net/articles/modern-fp-part-2
Not everything is black or white (no architecture vs clean architecture). You have to think about what you are doing (having fucked up on previous projects help) and don't follow anything you've read blindly.
Silver bullets, yadda yadda
https://www.amazon.com/Clean-Architecture-Craftsmans-Softwar...
I'm sure the book will have many real world code examples, as is fairly typical of his previous works.
Disappointing.
Domain driven design. You model you business problem using plain objects and methods. The domain should do nothing else other than modeling the domain, and solving the business problem you app is designed to solve. It should be persistence ignorant.
Then everything else simply has adapters for interfacing with the domain. Including a persistence adapter.
This is basically ports & adapters, and a lot of these architectures are basically variations on that.
You can get pretty far just worrying about this part.
I tend not to find these acronym names and the diagrams very helpful in terms of actually creating code. Separating concerns, as we all already know, is crucial, as is being careful about where knowledge is located. But I've found that if you buy in to these architecture patterns, you quickly become confused about which part of the pattern a given class or module or area of responsibility falls under. Like, is my `BazFrobber` a Gateway, or a Presenter, or a Controller... ?
The thing with these abstract architecture principles is they are not always practical if you try to be puristic about enforcing them.
I get where these ideals come from. I've seen novice programmers take a stab at writing mildly complicated apps, and the code is a nightmare to read because everything is jumbled up together. This is bad and we can all agree on that.
But I've also seen projects where everything is split into a tiny little function or object that don't seem to be doing anything meaningful. Presumably these projects are following the principles of "single responsibility" and "loose coupling", but it's so loose that it's hard to put the pieces together in your mind and nearly impossible to follow the flow of the program.
I consider generic rules of the form: "<X> related objects should not be doing <Y> related stuff" to be harmful. (An example of such "bad rules" can be found in the article under the heading "Use Cases")
Instead I prefer practical rules that make the code easy to read, understand, and updated, without making it any more complicated that it needs to be.
* If a piece of logic can be contained in a function, let it be contained in one function without splitting it across 10 different objects and factories and coordinators. Even if that function is slightly long, there's no need to split it apart just because it's over the arbitrary threshold of say, 25 lines.
* Create abstractions around a set of vocabulary that you can use to describe the problem domain and the process of doing things within the application. Make sure to document your vocabulary well and try to keep it as small as possible (but not smaller)
A good example of this is git's vocabulary for commits, trees, and blobs.
Every operation doesn't need to be a class (or worse, a series of classes and factories). Operations can be functions that operate on the structure's (objects) you've defined.
* Keep related files and functions together.
If you have a server side module implementing an json API, and an html page designed to display the API, and a javascript module designed to drive the UI on the html file, and a sql file describing the data you're displaying, then let's put all those files together in the same directory, instead of spreading them thin across separate folders
app/controllers/api/X/X.py
app/templates/X/X.html
app/js/views/X/X.js
app/css/X/X.css
Why not put them all under: app/X/X.py
app/X/X.html
app/X/X.js
app/X/X.css
Related material:Object-Oriented Programming is Garbage: 3800 SLOC example
I say this because we follow it pretty religiously at my work and it never feels impractical nor does it really feel like it adds too much extra work to anything we do, nor does it feel like related code is far apart nor are our functions too small.
In practice, following the clean architecture means, for example, that my business logic can't explicitly depend on some database related code. If we have business logic that needs do stuff in the database, it defines an interface of some methods that it expect the database repository to implement that it can use. It then has an implementation that satisfies that interface passed to it using dependency inversion (as pointed out in the article).
So it really means that if some business logic related to, say, tweets needs to retrieve some from the database, it calls something like "tweetRepository.findAllTweets()", which returns a bunch of "Tweet" domain objects. It never sees database rows, it knows nothing about which database we use, etc. Our business logic is focused solely on dealing with the use case, and nothing to do with the nitty gritty of how we interact with the database or how we eventually return that data to a user that needs it.
Which is great, because it means if we ever change how our database layer works, we only need to change our database repository code. We don't need to worry about changing business code if we can satisfy the same interface as before.