Monoliths Are the Future
changelog.com
changelog.com
Most people think a micro-service architecture is a panacea because "look at how simple X is," but it's not that simple. It's now a distributed system, and very likely, it's a the worst-of-the-worst a distributed monolith. Distributed system are hard, I know, I do it.
Three signs you have a distributed monolith:
1. You're duplicating the tables (information), without transforming the data into something new (adding information), in another database (e.g. worst cache ever, enjoy the split-brain). [1]
2. Service X does not work without Y or Z, and/or you have no strategy for how to deal with one of them going down.
2.5 Bonus, there is likely no way to meaningfully decouple the services. Service X can be "tolerant" of service Y's failure, but it cannot ever function without service Y.
3. You push all your data over an event-bus to keep your services "in-sync" with each-other taking a hot shit on the idea of a "transaction." The event-bus over time pushes your data further out of sync, making you think you need an even better event bus... You need transactions and (clicks over to the Jepsen series and laughs) good luck rolling that on your own...
I'm not saying service oriented architectures are bad, I'm not saying services are bad, they're absolutely not. They're a tool for a job, and one that comes with a lot of foot guns and pitfalls. Many of which people are not prepared for when they ship that first micro service.
I didn't even touch on the additional infrastructure and testing burden that a fleet of micro-services bring about.
[1] Simple tip: Don't duplicate data without adding value to it. Just don't.
This line of argument fails to take into consideration any of the reasons why in general microservices are the right tool for the right job.
Yes, it's challenging, and yes it's a distributes system. Yet, with microservices you actually are able to reuse specialized code, software packages, and even third-party services. That cuts down on a lot of dev time and cost, and makes the implementation of a lot of of POCs or even MVPs a trivial task.
Take for example Celery. With Celery all you need to do to implement a queuable background task system that's trivially scalable is to write the background tasks, get a message broker up and running, launch worker instances, and that's it. What would you have to do to achieve the same goal with a monolith? Implement your own producer/consumer that runs on the same instance that serves requests? And aren't you actually developing a distributes system anyway?
There are many benefits to having microservices that people seem to forget because they think that everyone interested in microservices is interested in splitting their personal blog into 4 different services.
They take coordination, good CICD, and a lot of forethought to ensure each service is cooperating in the ecosystem properly, but once established, it can do wonders to dev productivity.
I think if the database gets too overloaded I'll partition certain tree nodes across multiple masters (this is feasible because the framework doesn't rely on a single timestream).
With the level of shared code (the framework) and the single database, it's somewhat monolithic but the actors themselves are quite well-behaved and independent on top of it.
that's a little bit of a straw man because that's not the "microservice" architecture this post is talking about. I personally wouldn't call that a "microservice" architecture, I'd call it, "a background queue", although strictly speaking it can be described as such.
what this post is talking about are multiple synchronous pieces of a business case being broken up over the http / process line for no other reason than "we're afraid of overarchitecting our model". This means, your app has some features like auth, billing, personalization, reporting. You start from day one writing all of these as separate HTTPD services rather than just a single application with a variety of endpoints. Even though these areas of functionality may be highly interrelated, you're terrified of breaking out the GOF, using inheritance, or ORMs, because you had some bad experience with that stuff. So instead you spend all your time writing services, spinning up containers, defining complex endpoints that you wouldn't otherwise need...all becuase you really want to live in a "flat" world. I'm not allowed to leak the details of auth into billing because there's a whole process boundary! whew, I'm safe.
Never mind that you can have an architecture that is all separate process with HTTP requests in between, and convert it directly to a monolithic one with the same strict separation of concerns, you just get to lose the network latency and complex marshalling between the components.
Independent services that create a coherent whole enforce isolation barriers. I don't believe in microservices. These things don't need to be micro. They can be just normal-sized services. I don't even particularly care internal to those barriers how bad things get. There are programmers that write shit I think is hot garbage and will cause all kinds of bugs as time goes on. But if they are confined to their space space, and they're solving their problem and are happy and making management happy then w/e. That piece of the whole will collapse and die eventually, but it won't take everything with it. It's just a piece, and it provided value for a while, so probably worth it from a business sense.
But when you have a monolith? And Developer Dave and Programmer Pete tell Big Boss Bob that they could sure get that feature out quick if only it weren't for those pesky rules preventing them from putting in a few mutexes in that one module so they can just read some data directly from it, and boy wouldn't it that be swell? Well Big Boss Bob says we need to get this feature SHIPPED BOYES so buckle up and put in those mutexes, and now the fucking thread goes into a deadlock state every so often but it's really intermittent so you spend hours debugging the damn things and late nights because yeah features gotta ship but shit gotta work, and you trace it back to Developer Dave and Programmer Pete and their change but what do you do? Big Boss Bob said do it that way and what? You gonna whine to Director Dan about it? Is that gonna get you back your late night spend figuring out what was going on? Nah.
Make systems that are just small that when people break them completely it doesn't mean the end of the world to throw it out and start over.
What is there about microservices that makes them better at this? Most of the stuff I have seen has been highly coupled.
Now that said, with the new jigsaw module system in the jvm, and the multi-language support that is constantly getting better, a disciplined enough senior management team could enforce module boundaries within the process. It means any change to jvm command line flags would require the approval of the most senior tech lead, because that's how module boundary enforcement can be disabled, but if you have that and it works you would get significant performance and simplicity benefits.
Also, like I said, not a fan of the "micro" terminology, it just confuses the issue.
And keep in mind that coupling is not the opposite of encapsulation. Encapsulation just means hiding internal state. So unless you're going completely out of your way to break things, you'll have services that communicate through some message passing protocol, which means services are inherently encapsulated. It's not that you can't break that, it's that generally it's hard.
Here's an example of breaking it, though:
You have an API that talks to the database. You need to get users, groups, and preferences. Instead of writing different endpoints to access all of these things individually, clever you decides to simply make an endpoint that is "postQuery" and you give it a sql string and it will return the results. Great. You have now made the API effectively pointless.
Another example: you have an API that needs to perform standard accounting analysis on an internal dataset. Instead of adding endpoints for each calculation (or an endpoint with different inputs), you create an endpoint "eval" that will internal eval a string of the language of your choice. Congrats, you can now use that API to execute arbitrary code! No need to define any more pesky endpoints.
So yeah, absolutely people can make shit decisions and build garbage. It's entirely possible. But hey, at least it's pretty obvious this way and if you see your team do this you can look for another job.
So you have encapsulated services... boss comes and says "we need feature X right away". What if feature X spans all of your microservices? The bad programmer will hack together a monstrosity between multiple services. It's the micro-lith problem: instead of a monolith, you have a monolith disguised as micro services. Now their poor choices are spread across a lot of services and distributed, it's not really confined to just one service.
Wait, what is a good programmer supposed to do in this scenario?
After my team had taken that project and split it into a bunch of smaller packages, it's not like everyone magically became better programmers. The same people who were introducing fucking spaghetti code that did who-knows-what in overly complex ways were still around, and they were given effectively their own sandboxes. In fact quite a few more teams within the company began using those packages, which just became modules they shipped with. We no longer had to deal with people screwing with core packages because they no longer had ownership, they could no longer make merges into that area of the code base, so we could keep things stable.
So I'm a hard sell on monoliths. Like, I'm not actually pro micro-services, per se. I'm mostly just anti-monolith. Giant code bases with multiple teams simultaneously contributing are doomed to fucking disaster.
It is not a strawman; it's a concrete example of the technical, practical, operational, and economical advantages of microservices architecture, more specifically service reuse, specially managed services provided by third parties.
While you're groking how a multihreadig library is expected to be used, I already fired a message broker that distributes tasks across a pool of worker instances. Why? Because I've opted not to go with the monolith and went with the microservices/distributed architecture approach.
I'm so confused by that statement... because I can't for the life of me figure out how you got there.
You absolutely can have a monolith which is multi-threaded or asynchronous and "resource/task pools." The JVM for instance has threads, I use BEAM (Elixir) personally, and it's even pre-emptively scheduling my tasks in parallel and asynchronously... but, I still don't get what multi-threading has to do with microservices.
Microservices and monoliths are boundaries for your application they aren't implementation details (i.e. all microservices must be asynchronous is strictly not true) in and of themselves, they're design details. That design can influence the implementation but they are separate.
Ex. there are plenty of people who use Sidekiq and redis like you're using Celery but don't call it a microservice. It's just a piece of their monolith since it's largely the same depdencies.
I mean consider a gen server mounted in your supervisor tree. its 5 min of work; tops. Doing the same with kubernetes would require coordinating a message broker, picking a client library, creatig a restart strategy and networking. all of which would add considerably to your development time
You would take Celery, get a message broker up and running and launch some worker instances.
1. Microservices do not have some inherent property of having to duplicate data. You can have data in single source and deliver that data to anyone who needs it through an API. There are infinitely many caching solutions if you are worried about this becoming a bottleneck
2 and 2.5. There are tools for microservice architectures that counter these problems to a degree, mainstream example being containers and container orchestration (e.g. Docker and Kubernetes). One can even make an argument that microservices force you to build your systems so they are more robust than your monolith would be. If the argument for monoliths is that it's easier to maintain reliability when every egg is in one basket then you are putting all your bets into that basket and it becomes a black hole for developer and operations resources, as well as making operations evolution very slow
3. There are again tools for handling data and "syncing" (although I don't like the idea of having to "sync") the services, for example message queues / streaming processing platforms (e.g. Kafka). If some data or information might be of interest to multiple services, you should then push such data to a message queue and consume it from services that need it. The "syncing" problem sounds like something that arises when you start duplicating your data across services which shouldn't happen (see my argument on 1.)
Again not to say microservices are somehow universally better. Just coming into defense of their core concepts when they get unfairly accused
Once you add in this consolidated data service, every other service is dependent on the "data service" team. This means almost an change you make, requires submitting a request. Are their priorities your priorities? I would hate to be reliant on another team to do every development task.
Theoretically you could remove this issue by allowing any team to modify the data service, but then at that point you've just taken an application and added a bunch of http calls between method calls.
This same problem ops up with resiliency. If you have a consolidated data service, what happens if your data service goes down? How useful are your other services if they can't access any data?
Re. teams: For any project above a certain size, you'll have teams. If that's a network boundary, a process boundary, or a library boundary doesn't change that you'll have multiple teams for a large project.
I'm not sure I get the resiliency point. I worked on a project where the dependent data service was offline painfully frequently. We used async tasks and caching to keep things running and were able to let the users do many tasks. For us our tool was still fairly useful when dependencies went down. If we used monolith then everything would be down, right? That doesn't sound better.
For sure, and one of the big selling points for micro-services is you can split those teams by micro service, with each team having an independent service they are responsible for. But when a big chunk of everyone's development is done on one giant service everyone shares you don't get the same benefits you would if the services were independent. Or put another way, splitting micro-services vertically can yield a bunch of benefits, but splitting them horizontally introduces a lot of pain with few benefits.
> I'm not sure I get the resiliency point. I worked on a project where the dependent data service was offline painfully frequently. We used async tasks and caching to keep things running and were able to let the users do many tasks. For us our tool was still fairly useful when dependencies went down. If we used monolith then everything would be down, right? That doesn't sound better.
I'm not saying to never spin off services. If you have a piece of functionality that you just can't get stable for the life of you, splitting it off into it's own service, and coding everything up to be resilient to it's failure makes a lot of sense. (I am very curious what the cause of the data service crashing was that you couldn't fix.)
But micro-services aren't a free lunch for resiliency. You're increasing the numbers of systems, servers, configurations, and connections which by default will decrease up-time until you do a ton of work. Not to mention tracking and debugging cross service failures is much more difficult than a single server.
Now, there is a pattern called Event Sourcing (ES) which proposes that the source of truth should be the event bus itself and microservice databases are mere projections of this data. This is all good and well except it's very hard to implement in practice. If a microservice needs to replay all business events from months or years in the past it may take hours or days to do this. What about the business in the meantime? If it's a service that significantly reduces the usability of you application you effectively have a long downtime anyway.
Transactional activity becomes incredibly hard in the microservices world with either 2 phased commit (only good for very infrequent transactions due to performance) or with so called Sagas (which are very complex to get right and maintain).
Any company that isn't delivering its service to billions of users daily will likely suffer far more from microservices than they will benefit.
In my experience it's not hard to implement, but of course it depends on the problem domain (and probably also on not splitting things up willy-nilly because of the microservices fad). I think the key to event sourcing and immutability in general is to not overdo it. For example you will likely need to redact certain data (e.g. for legal compliance), so zero information loss is out. Systems like Kafka are a poor choice for long term data storage, the default retention is 1 week for a reason.
But the things that are wonderful about event sourcing (the ability to inspect, replay and fix because you haven't lost information) mostly materialize over a 1 week timeframe.
If you need to recover a lot of state from the event log you will need store aggregated event data at regular intervals to play back from to have acceptable performance. But in practice, in many cases the data granularity you need goes down as the data ages anyway, and you do some lossy aggregation as a natural part of your business process (as opposed to to deal with even sourcing performance problems). I.e. for the short timeframe kafka is the source of truth, but for the stuff you care long term it's some database, and this happens kinda naturally. So often you don't need to implement checkpointing.
We learned and is continuing learning that.
We're currently translating a 20 year old ~50MLOC codebase into a distributed monolith (using a variety of approaches that all approximate strangler). I have far less motivation to go to work if I know that I will be buried in the old monorepo. I can change, build and get a service changed in less than an hour. Touching the monorepo is easily 1.5 days for a single change.
We seem to be gaining far more in terms of developer productivity than we are losing to operational overhead.
I also should have mentioned that it's definitely more pleasant for those in purely development roles. Troubleshooting, resiliency and system effects don't impact everyone (and I actually like those types of hard problems). I'd also suggest that integrating tracing, metrics, and logging in a consistent way is imperative. If you're on Kubernetes, using a proxy like Istio (Envoy) or LinkerD is a great way to get retries, backoff, etc established without changing coded.
Finally, implementing a healthcheck end-point on every service and having the impact of any failures properly degrade dependent services is really helpful both in troubleshooting and ultimately in creating a UI with graceful degradation (toasts with messages related to what's not currently available are great). I have great hopes for the healthcheck RFC that's being developed at https://github.com/inadarei/rfc-healthcheck.
^ this
I think the naming decision of the concept has been detrimental to its interpretation. In reality, most of the time what we really want is a"one-or-more reasonably-sized systems with well-enough-defined responsibility boundaries".
Perhaps "Service Right-Sizing" would steer people to better decisions. Alas, that "Microservices" objectively sounds sexier.
Whereas if you write a single monolithic program its connections are described in code, preferably type-checked by a compiler. I think that gives you at least theoretically a better chance of understanding what are the things that connect, and how they connect.
So if there was a good programming language for describing micro-services then that would probably resolve many management difficulties, and then the question would be simply do we get performance benefits from running on multiple processors.
The runtime system has built-in support for concurrency, distribution and fault tolerance. Because of the design goals for Erlang and it's runtime, you get services that can all run on one system or be distributed across a network, but the code that you actually write is relatively simple; the entire distributed system acts as a fault-tolerant VM that has functional guarantees.
If your startup node fails, then other nodes are elected. If a node crashes while in the middle of a method, another node will execute the method instead.
The runtime itself has some analogies to functional coding styles. It runs on a register machine rather than a stack machine. The call and return sequence is replaced by direct jumps to the implementation of the next instruction.
Nit: if service X could function without service Y, then it seems to follow service Y should not exist in the first place. And equivalently, the functionality of service Y before some microservice migration.
Recommendations are not necessary, and not showing them will significantly affect the bottom line, so you don't want to skip them if possible. But not showing the product page (in a timely manner), because the recommendation engine has a hiccup, is even worse for your bottom line.
What about running multiple versions of a microservice in parallel -- don't each need their own yet separate databases that attempt to mirror each other as best they can?
The short answer is "no," as succinctly stated by the, I assume from the name, majestic SideburnsOfDoom. The versions shouldn't _EVER_ be incompatible with each other.
E.g. you need to rename a column.
Do not: rename the column, e.g. `ALTER TABLE RENAME COLUMN...`. Because, your systems are going to break with the new schema.
Do: Add a new column, with the new name, and migrate data to the new column, once it's good upgrade the rest of your instances, then drop the old column. Because, you can use both versions at the same time now without breaking anything. Yes, it can be a little tricky to get the data sync'd into the new column, but that's a lot less tricky than doing it for _every_ table and column.
Not much is said about S3's design principles in public, but that one was one of them.
Disclaimer: Recalling from memory.
1) "A" Component
2) "B" Component
3) "A" Server
4) "A" Client
So for example if you started off with a single "project" in your favorite IDE, you now have 4, give or take. You might be able to code-gen your server and client code out of a single IDL file or something, but generally speaking you're going to be writing code like "B -> A client -> A server -> A" no matter what instead of simply "B -> A".Now you have to worry about network reliability, back-pressure, queuing, retries, security, bandwidth, latency, round-trips, serialization, load-balancing, affinity, transactions, secrets storage, threading, and on and on...
A simple function call translates to a rats nest of dependency injection, configuration reads, callbacks, and retry loops.
Then if you grow to 4 or more components you have to start worrying about the topology of the interconnections. Suddenly you may need to add a service bus or orchestrator to reduce the number of point-to-point connections. This is not avoidable, because if you have less than 4 components, then why bother to break things out into micro services in the first place!?
Now when things go wrong in all sorts of creative ways, some of which are likely still the subject of research papers, heaven help you with the troubleshooting. First, it'll be the brownouts that the load balancer doesn't correctly flag as a failure, then the priority inversions, then the queue filling up, and then it'll get worse from there as the load ramps up.
Meanwhile nothing stops you having a monolithic project with folders called "A", "B", "C", etc... with simple function calls or OO interfaces across the boundaries. For 99.9% of projects out there this is the right way to go. And then, if your business takes off into the stratosphere, nothing stops you converting those function call interfaces into a network interface and splitting up your servers. However, doing this when it's needed means that you know where the split makes sense, and you won't waste time introducing components that don't need individual scaling.
For God's sake, I saw a government department roll out an Azure Service Fabric application with dozens of components for an application with a few hundred users total. Not concurrent. Total.
Yeah...
Analysts expect to be able to connect to one system, see their data, and write queries for it. They were never brought into the microservices strategy, and now they're stumped as to how they're supposed to quickly get data out to answer business questions or show customers stuff on a dashboard.
The only answers I've seen so far are either to build really complex/expensive reporting systems that pull data from every source in real time, or do extract/transform/load (ETL) processes like data warehouses do (in which the reporting data lags behind the source systems and doesn't have all the tables), or try to build real time replication to a central database - at which point, you're right back to a monolith.
Reporting on a bunch of different databases is a hard nut to crack.
It's not necessarily a bad idea though :-/
Maybe I'm wrong.
You need to use something like Spark, Presto or Drill if you want to run queries across different data sources.
Right, that's the data warehouse method that I described. "Put them into their system" is a lot harder work than just typing in "put Kafka on top of your database."
Which all gets back to the point of the OP.
The article talks about micro services being split up due to fad as opposed to deliberate, researched reasons. Putting Kafka over the database also makes the data distributed when in most cases, it’s not necessary!
Disclaimer: working on Debezium
In a non-append-only scenario, Debezium tracks each source operation (insert, update, delete) from the replication log (oplog, binlog, etc) as an individual operation that's emitted into a kafka topic. How does one efficiently replicate this to a Data Warehouse in an efficient manner?
I have not been able to use Debezium as way to replicate to a Data Warehouse for this very reason. At least not without having to resort to very complicated data warehouse merge strategies.
Note, there exist Data Warehouses that allow tables to be created in either OLAP or OLTP flavor. I understand that Debezium could easily replicate to an OLTP staging table. But are there any solutions if this isn't an option?
Your casual tone is at odds with what I've seen when teams run Kafka clusters in production. Not a decision I would take so lightly.
That's exactly how that problem has been solved successfully for the past 20 years.
It is extremely expensive and slow to move all the data to a data warehouse. Ignoring the cost element, the latency from when data shows up in the operational environment to when it is reflected in the output of the data warehouse is often unacceptably high for many business use cases. A large percentage of that latency is attributable solely to what is essentially complex data motion between systems.
For it to work, online OLAP write throughput needs to scale like operational databases. This is not the case in practice, so operational databases can scale to a point where the OLAP system can't keep up. The technical solution is to scale the operational system to absorb the extra workload created by the data warehousing applications, but current database architectures are not really designed for it so it isn't trivial to do.
Simple enough. Surely you wouldn't run analytics directly on your prod serving database, and risk a bad query taking down your whole system?
Uhh, yep, that's exactly how a lot of businesses work.
There are defenses at the database layer. For example, in Microsoft SQL Server, we've got Resource Governor which lets you cap how much CPU/memory/etc that any one department or user can get.
Also, you complicate locking down access to the database. A reporting database can typically contain less sensitive info, reporting would not have password (hashes) for user accounts for example.
Reading other comments Brent’s made here, I’m not so sure.
No no, it's not a good idea, but it's just the reality that I usually have to deal with. I wish I could lay down rules like "no analyst ever gets the rights to query the database directly," but all it takes is one analyst to be buddy-buddy with the company owner, and produce a few reports that have really high business value, and then next thing you know, that analyst has sysadmin rights to every database, and anybody who tries to be a barrier to that analyst's success is "slowing the business down."
I work on a monolith that does this, but its usually not even necessary a single db server on modern hardware with proper resource governing can handle quite a bit.
Just because a lot of businesses do it, doesn’t mean it’s a good idea. A lot of businesses don’t do source control of any kind, so should we all do that too?
This is not a big deal if your system is designed for it, sophisticated database engines have good control mechanisms for ensuring that heavy reporting or analytical jobs minimally impact concurrent operational workload.
Right, that's the data warehouse method I described, keeping a central database in a reporting system. But now you just have to keep that database schema stable, because folks are going to write reports against it. It's a monolith for reporting, and its schema is going to be affected by changes in upstream systems. It's not like source systems can just totally refactor tables without considering how the data downstream is going to be affected. When Ms. CEO's report breaks, bad things happen.
No, because as soon as you change your schema, you have to plan ahead with the reporting team for starters. The reports still have to be able to work the same way, which means the data needs to be in the same format, or else the reports need to be rewritten to handle the changes in tables/structures.
IMO the devops folks should define some standard containers that include facilities for reporting on low-level metrics. Most of the monitoring above that should be managed by the microservice owner. The messages that are consumed for BI and external reporting should not have breaking changes any more than the APIs you provide your clients should.
This is a great point. The way to make backward-compatible changes to an API is by adding additional (JSON) keys, not changing / removing keys. The same approach works for a DB -- adding a new column doesn't break existing reporting queries.
Yes, but data lakes don't fill themselves. Each team has to be responsible for exporting every transaction to the lake, either in real time or delayed, and then the reporting systems have to be able to combine the different sources/formats. If each microservice team expects to be able to change their formats in the data lake willy-nilly, bam, there breaks the report again.
Schemas evolve as the needs of the product change, and that evolution will always outpace the way the business looks at the data.
The best way I've seen to deal with this is to handle this at report query-time (e.g. pick a platform that can effectively handle the necessary transformations at query-time, rather than at load-time).
Now databases themselves are different stories, they are the persistence/data layer that microservices themselves use . But it's actually doable and I'd even say much easier to use microservices/serverless for ETL because it's easier to develop CI/CD and testing/deployment with non-stateful services. Of course, it does take certain level of engineering maturity and skillsets but I think the end results justify it.
That all depends on the API.
It's easier to join tables in databases that live on a single server, in a single database platform, than it is to connect to lots of different data sources that live in different servers, possibly even in different locations (like cloud vs on-premises, and hybrids.)
1) what if your data set is larger than you can practically handle in a single DB instance?
2) nothing about a monolith implies you have a single data platform, let alone a single DB instance
I have clients at 10TB-100TB in a single SQL Server. It ain't a fun place to be, but it's doable.
ETL is really database focused and batch focused , Extract, Transform, Load.
Data pipeline, is a combination of streams and batch. For example, you can implement a Data Capture using something like https://debezium.io/
Here's how Netflix solves it https://netflixtechblog.com/dblog-a-generic-change-data-capt...
Overview
"Change-Data-Capture (CDC) allows capturing committed changes from a database in real-time and propagating those changes to downstream consumers [1][2]. CDC is becoming increasingly popular for use cases that require keeping multiple heterogeneous datastores in sync (like MySQL and ElasticSearch) and addresses challenges that exist with traditional techniques like dual-writes and distributed transactions [3][4]."
Then it doesn't seem to be solved. Seems like teams operating at a lean scale would have an issue with this, especially teams with lopsided business:engineering ratios
And for data scientists working on production models used within production software, most inference is packaged as containers in something like ECS or Fargate which are then scaled up and down automatically. Eg, they are basically running a microservice for the software teams to consume.
Real time reporting, in my opinion, is not the domain of analysts; it's the domain of the software team. For one, it's rarely useful outside of something like a NOC (or similar control room areas) and should be considered a software feature of that control room. If real-time has to be on the analysts (been there), then the software team should dual publish their transactions to kinesis firehouse and the analytics team can take it from there.
Of course, all of this relies heavily on buy-in to the world of cloud computing. Come on in, we all float down here.
The complexity you and the OP seem to be describing are more in the management and prioritization of analytics projects than in the actual "this is a hard technical problem" domain. It's just a lot of it is tedious especially compared to "everyone just put all your data in the Oracle RACs and bug the DBA until they give you permission" model of the past.
What ended up happening is each application uses its own database, nobody offered applications that could be configured to an existing data base, and all of our data is in silos.
I think one reason to avoid this approach is because SQL and other DB languages are pretty terrible (compared to popular languages like C#, Python, etc...) But why has no one written a great DB language yet?
Where did you get that from?
In any case, designing a single schema that encompassed all the needs of the organization and could grow and change as the organization did was nearly always too much to ask.
This was in the days of enterprise data modeling where people believed there really was just one data model or object model that could represent the whole org, independent of the needs of any given application. I don’t think anybody believes that any more.
For write-heavy workloads, best of luck to you :)
* Testing is god-awful. To test a simple thing you had to know how the whole application worked, because there's validation in triggers, which triggers other triggers, which require things to be in a certain state. This made refactoring really hard/risky so it rarely got done.
* There's a performance ceiling, and when you hit that, you're done. We did hit a ceiling, did months of performance tuning, then upgraded to the biggest available box at the time, 96 cores, 2TB ram, which helped, but next time the upgrade won't be big enough. You're limited in what one box can do (and due to stored procedures being tied to the transaction there's limits to what you can do concurrently as well)
Debugging PL/SQL without the PL/SQL debugger is nearly impossible. Unfortunately a lot of shops cheap out on developer tools after they buy the server licenses. I never liked the idea of Java on the database. The good thing about PL/SQL is that nobody wants to write it so it has a tendency to not be overused by most developers.
The performance ceiling probably wouldn't be too low if there weren't too much extra activity with each update. As with all databases, sharding and replication are your friend.
SQL is actually hard to compare to programming languages, because in SQL you say what you want, while in iterative language you say how. I only know one language that was competing with SQL (and lost) it is QUEL (originally it was used by Ingress and Postgres).
BTW for triggers and stored procedures you actually can use traditional language, I know that PostgreSQL supports Python, you just need to load a proper extension to enable it.
Perhaps refactoring, had it been better understood around 1970, could have gone a long way toward harmonizing diverse schemas, allowing experimentation with eventual refactoring into the common database.
Our current environment makes this impossible. There's no way that Salesforce is going to ship a version that works with your company's database schema. You're going to have to supply that replication yourself. Same for Quickbooks. To get that kind of customization you need to be spending hundreds of thousands for enterprise software.
Edit - I just saw that you addressed this in your original post: "nobody offered applications that could be configured to an existing data base"
One of the key concepts in microservice architecture is data sovereignity. It doesn't matter how/where the data is stored. The only thing that cares about the details of the data storage is the service itself. If you need some data the service operates on for reporting purposes, make an API that gets you this data and make it part of the service. You can architect layers around it, maybe write a separate service that aggregates data from multiple other services into a central analytics database and then reporting can be done from there or keep requests in real time, but introduce a caching layer or whatever. But you do not simply go and poke your reporting fingers into individual service databases. In a good microservice architecture you should not even be able to do that.
This is why I distrust all of the monolith folks. Yes, it's easier to get your data, but in the long run you create unmaintainable spaghetti that can't ever change without breaking things you can't easily surface.
Monoliths are undisciplined and encourage unhealthy and unsustainable engineering. Microservices enforce separation of concerns and data ownership. It can be done wrong, but when executed correctly results in something you can easily make sense of.
There's going to be a relationship between data in your services, but it shouldn't be directly referential.
But to speak directly to your concern, you have to think about service boundaries and granularity correctly. Nobody is saying make a microservice out of every conceivable table. Think about the bigger picture, at a systems level. Wherever you can draw boxes you might have a service boundary.
Why would you need to join payment data to session and login data?
Do you need to compare employee roles and ACLs against product shipping data?
These things belong in different systems. If you keep them in the same monolith, there's the danger that people will write code that intertwines the model in ways it shouldn't. Deploying and ownership become hard problems.
The goal is to keep things that are highly functionally related together in a microservice and expose an API where the different microservices in your ecosystem are required to interact. (Eg, your employees will login.)
When the data analytics folks want to do advanced reporting on the joins of these systems (typically offline behavior), you can expose a feed that exports your data. But don't expose an internal view of it to them or they'll find ways of turning you into a monolith.
Also then what also happens is microservices are created using different languages which in turn adds so much complexity to understand what is going on on the whole big picture level.
And code gets repeated a lot more. If there is change in a microservices or update everyone will need to figure out what services depend on and how they will have to adapt. With monolith you can just use your IDE to see what will break if you make a change. So much repeated business logic. Creating a new feature involves having to have many meetings to figure out what services in which way have to be updated.
It is crazy mess in my opinion.
I have been with a company that had monolith application which they split up to more than 15 services (some python, some js, Scala, Java, etc...). Monolith still is used for some parts that are not migrated. I was working on single service having no idea how the whole system worked together. Then I had to do something in the old parts and I very quickly got an understanding how everything works together.
With microservices, without a good documentation how it connects, it's going to leave a very bad impression.
It is still nowhere close to ability to jumping around with IDE.
It might be in a different language, different design patterns and to get to the details you have to check out that project anyway because you can't document absolutely everything out of code base. And if you do you will end up with multiple sources of truth.
It is so much more likely that for every little issue which you otherwise might be able to find an answer to yourself very easily you will have to contact the team owning that microservices.
It is not only mentally exhausting. It is time consuming, it requires so much back and forth. It creates so much dependence on other people because figuring out how things are related is so much more difficult.
Sometimes I have 8 or more different IDE windows open to understand what is going on.
This is what people mean when they say "distributed monolith" vs. microservices.
And holy fuck is debugging that stuff difficult. HUUUUUGE waste of time, but management looooooooooves their blasted microservices...
This is the hardest part.. I'd argue that this is almost impossible to do correctly without significant domain modeling experience.. also microservices by nature make this hard to refactor these boundaries (compared to monoliths where you'd get compile time feedback)
I prefer to make a structured monolith first (basically multiple services with isolated data that are squished together into a single deployable) and pull them out only if I really need to... Also helps with keeping ms sprawl under control
That's what SOA and microservices is supposed to solve.
At that scale you do reporting from a purpose-built service.
Allegedly.
Quite the self-fulfilling prophecy there.
> Yes, it's easier to get your data, but in the long run [...]
Systems can and should be evolved and adapted over time. E.g. deploying components of the monolith as separate services. You can't easily predict what the requirements for your software going to be in say 10 years.
And depending on the stage a company is, easy access to data for business decisions outweighs engineering idealism.
Microservices require your organization to have an engineering culture. I would be afraid of introducing them at, say, Home Depot where (I've heard) your average programmer doesn't even write tests.
If you have engineering talent within a small multiplicative factor of Google (say 0.5), then you can pull off Microservices at your org.
Edit: I'm being downvoted, but I don't think it's a dangerous assumption or point to make that it takes a certain amount of discipline and experience to implement microservices correctly. When you have that technical capacity and the project calls for it, the benefit is tremendous.
I've seen good and bad in each approach. It's certainly possible to enforce good SOCs and proper boundaries in monorepos, and also possible to plough a system into the ground with microservices.
They're all just tools in your toolbox and both have a part to play in modern development.
I think there are different levels of sophistication of "engineering idealism". GP talks about "data ownership", and I get the desire to keep the data a microservice is responsible for locked in tightly with it. But let's be precise why it's good: because isolating responsibility reduces complexity. Not because code has some innate right to privacy.
In my own engineering idealism, there's no internal data privacy in the system. Things should be instrumentable, observable in principle. If an analyst wants to take your carefully designed internal NoSQL document structure and plug it into an OLAP cube for some reason, there must be a path to doing that; if that's an expected part of the business, the service needs to have it on the feature list, that this should be doable without degrading the service.
Software needs to be in boxes because otherwise we can't handle it mentally, but the boxes really shouldn't be that black.
YMMV, but the tradeoff is less complexity at the SWE/prod department, and more at the analytics team.
The thing is, it just shifts around complexity. Once you have microservices, you have to deal with a bunch of new failure modes, plus a bunch of extra code whose only purpose is to provide an interface to other services. And in terms of separating data, the worst part is that you've prevent access this data with some other data within the same transaction.
As the article says though, you can't fix a people problem (bad engineering practices and discipline) by going from one technology to another (monolith to microservices).
It's about code quality, microservices are easy replaceable. Modules are too.
With both systems, the core part ( eg. mesh, Infrastructure, ... ) Is crucial.
I think experienced developers can see this, the ones that actually delivered products and had big code changes. The ones that handled their "legacy" code.
Microservices are just a way to enforce it, there are others. None are perfect or bad, both have their use-case.
In a microservice architecture it's harder to pretend you're doing it right.
The trick with microservices is that the ecosystem is maturing and there are still lots of ways to screw up other things that are harder to screw up with monoliths. In time 95% of those will go away (my specific prediction is that one day we will write programs that express concurrency and the compiler/toolchain will work out distributing the program across Cloud services--although "Cloud" will be an antiquated term by then--including stitching together the relevant logs, etc and possibly even a coherent distributed debugger experience).
After seeing a few of them, I'd say: "it's less embarrassingly obvious that you're doing it wrong."
But dig into the code for a few endpoints and it usually don't take long to find the crazy spaghetti and the poorly-carved-out separation of responsibilities breaches.
How does creating a tangle of microservices (effectively distributed objects) really solve the problem?
My understanding of microservices is a bunch of loosely connected services that can be changed with minimal impact to the others
Problem with the ideal is in reality this never works as complexity grows the spaghetti code moves to spaghetti infrastructure ( Done a network map of a large k8s / istio deployment lately ? )
So do modules/classes/interfaces etc. You don't need a layer of HTTP in between components to have abstraction.
In addition, it feels like microservices solve a problem that very few people really have. I've never run into a case where I though "boy, I'd sure like to have a different database for this one chunk of code". If that did happen, then sure, split it out, but I can hardly believe that splitting your entire code base into microservices has a net benefit. The real problem in nearly every project I've worked on is complexity of business logic. A monolith is much easier to refactor, and you can change the entire architecture if you need to without having to coordinate releases of many different applications.
You can enforce separation of concerns and data ownership in a monolith just as much as you can not enforce these two characteristics in a micro service architecture. Microservices and monoliths are a discussion about deployment artifacts, full stop.
The same folks aren't going to magically learn how to do distributed computing properly, rather they will implement unmaintainable spaghetti network calls with all the distributed computing issues on top.
Then your monolith is just all the modules glued together.
Depending on where you work, it can be a problem, because the separation is not always appropriate, and can for political reasons be much harder to revert when visible at the service level (for example because the architect doesn't understand consistency, or because your manager tells you that the distributed architecture documentation has been sent to the client so it cannot be modified).
In case of undue separation, reworking the internals of the enclosing monolith should have less chance to cause frictions.
I agree. In a monolith architecture, though, you CAN do that (and many shops do.) That's where their pains come from when they migrate from monolith to microservice: development is easier, but reports are way, way harder.
Not even that -- that idea is still highly debatable.
I would argue that it absolutely isn't easier, and the stepping-back-in-time of developer experience is one of the biggest problems with microservices.
Microservices in general, are way, way harder.
Side point: This is a needlessly hostile and unprofessional way to refer to a colleague. Remember that you and the reporting/analytics people at your company are working towards the same goals (the company's business goals). You are collaborators, not combatants.
You can express your same point by saying something like "The habit of directly accessing database resources and building out reporting code on this is likely to lead to some very serious problems when the schemas change. This is tantamount to relying upon a private API." etc.
We can all achieve much more when we endeavor to treat one another with respect and assume good intentions.
Anyway, what’s the reason not to treat people on hackernews with the same respect you’d treat a coworker with?
We have a whole industry around Analytics and Data and the tools and processes to build this reporting layer is well established and proven.
Nothing will give you as many nightmares as letting your analysts loose on your micro service persistence layers!
Our databases are open to way too many people. What's worse, they are multi tenant making refactoring really hard.
We used to have a few of those, especially on exadata clusters. Finally carted them out of the local dc after moving to RDS Aurora databases with strict policies. Might have caused 3 or 4 people to quit, but totally worth it for the 500+ people that stayed who now can own their data, schema and development (and be held responsible for it! -- another issue of multi-db-access, it's always someone else's fault). Went from deploying once a day with a 'heads up' message to no-message deploying multiple times per hour.
I cannot imagine people not doing that and having need to have stats in real time. For most shopping/banking stuff you can get away with once in 24 hours dumps and then analytics can be done on that.
Most APIs are glorified wrappers around individual record-level operations like- get me this user- or constrained searches that return a portion of the data, maybe paginated. Reporting needs to see all the data. This is a completely different query and service delivery pattern.
What happens to your API service written in a memory managed/garbage-collected language when you ask it to pull all the data from its bespoke database, pass it through its memory space, then send it back down the caller? It goes into GC hell, is what.
What happens when your API service when it issues queries for a consistent view of all the data and winds up forcing the database to lock tables? It stops working for users, is what.
There are so many ways to fail when your microservice starts pretending it is a database. It is not. Databases are dedicated services, not libraries, for a reason.
It is also true that analysts should not be given access to service databases, because the schema and semantics are likely to change out from under them.
The least bad solution? The engineering team is responsible for delivering either semantic events or doing the batch transformation themselves into a model that the data team can consume. It's a data delivery format, not an API.
Can you exapand on this a little? Or a paper that I can read?
Since a micro service deals with only its own data and reporting is then across services, we’d need to query across services to get data and make sense of it. If we’d ever need to query all records, then such records would become domain objects in the micro services first before being passed along. A large number of domain objects would require a large amount of memory. Processing and releasing domain objects will result in GC on the released objects.
You could say but oh, why not just return the underlying data without making objects? Well now you are exposing the underlying data format, which is what we’re trying to avoid by giving this job to the service.
Now I try to think about problems as "I have input data of shape X, I need shape Y" and fractally break it down into smaller shape-changes. I am kinda starting to get what those functional programmers are yammering on about.
It’s not like they come up with every report they think they might need while the micro service is being architected. They come up with a new report long after engineers have moved on. If it’s a SQL database, no problem. If it’s some silly resumeware data store, then what?
If you can't ask questions you didn't think of in advance, you didn't collect data.
Its not perfect but what we do is create a bunch of table views that represent each of the core data types in the system. We can then do all of the complex joins to collect the data analysts want in to an easy to query table as well as trying to keep the views consistent even as the db changes.
Why pull it into memory like that? Why not just pump it through a stream?
This really makes sense to me. I love the idea that part of a microservice team's responsibility is ensuring that a sensible subset of the data is copied over to the reporting systems in such a way that it can be used for analysis without risk of other teams writing queries that depend on undocumented internal details.
At what point in time?
In the end everything involves tradeoffs. If you need to partition your data to scale, or for some other reason need to break up the data, then reporting potentially becomes a secondary concern. In this case maybe delayed reporting or a more complex reporting workflow is worth the trade off.
DataWarehousing has drastically improved recently with the separation of Storage & Compute. A single analyst's query impacting the entire Data Warehouse is a problem that will in the next few years be something of the past.
(SRE here, but I work on databases as well all day)
It can access all those different databases.
You can also make your own connectors that make your services appear as tables, which you can query with SQL in the normal way.
So if the new accounts micro-service doesn't have a database, or the team won't let your analysts access the database behind it, you can always go in through the front-door e.g. the rest/graphql/grpc/thrift/buzzword api it exposes, and treat it as just another table!
Presto is great even for monoliths ;) Rah rah presto.
Presto and SparkSQL are SQL interfaces to many different datasources, including Hive and Impala, but also any SQL database such as Postgres/Redis/etc, and many other types of databases, such as Cassandra and Redis; the SQL tools can query all these different types of databases with a unified SQL interface, and even do joins across them.
The difference between Presto and SparkSQL is that Presto is run on a multi-tenant cluster with automatic resource allocation. SparkSQL jobs tend to have to be allocated with a specific resource allocation ahead of time. This makes Presto is (in my experience) a little more user-friendly. On the other hand, SparkSQL has better support for writing data to different datasources, whereas Presto pretty much only supports collecting results from a client or writing data into Hive.
I know Hive can definitely query other datasources like traditional SQL databases, redis, cassandra, hbase, elasticsearch, etc, etc. I thought Impala had some bit of support for this as well, though I'm less familiar with it.
And SparkSQL can be run on a multi-tenant cluster with automatic resource allocation - Mesos, YARN, or Kubernetes.
1. Pay a ton of money to Microsoft for Azure Data Lake, Power BI, etc.
2. Spend 12 months building ETLs from all your microservices to feed a torrent of raw data to your lake.
3. Start to think about what KPIs you want to measure.
4. Sign up for a free Google Analytics account and use that instead.
Okay, sounds reasonable enough for a complex enterprise.
> to feed a torrent of raw data to your lake
Well, there's the problem. Why is it taking a year to export data in its raw, natural state? The entire point of a data lake is that there is no transformation of the data. There's no need to verify the data is accurate. There's no need to make sure it's performant. It's just data exported from one system to another. If the file sizes, or record counts match, you're in good shape.
If it's taking a year to simply copy raw data from one system to another, the enterprise has deeper problems than architecture.
Yes, it makes life harder for the data engineers in the future, but it might turn out that analysts only ever need 5% of the data in the lake, and dealing with these schema changes for 5% of the data is easier than carefully planning a public schema for 100% of it.
It can be helpful to include some small amount of metadata in the export though, with things like the source system name, date & time, # of records, and a schema version. Schema version could easily be the latest migration revision, or something like that.
But if I haven't spent the effort to extract it, do I really own it?
If you collected it, you are responsible for it.Then you just don't keep red data for longer than 30 days.
On the other hand, the notion that “microservices == completely independent of everything else” is an unrealistic one to hold.
Consider:
Source Data -> Data Lake -> ETL Process -> Reporting DataWarehouse(s)/DataMart(s) -> User Queries
vs
Source Data -> Data Lake -> User Queries
vs
MonolithDB -> User queries
vs
MonolithDB -> ETL Process -> Reporting DataWarehouse(s)/DataMart(s) -> User Queries
A schema change in the source data should be easily updated in the ETL process in example 1. Most changes are minimal (adding, removing, renaming columns). And for a complete schema redesign in the source data, a new entry in the data lake should be created and the owners of the ETL process should decide if the new schema should be mangled to fit their existing reporting tables or to build new ones. Across the four models I outlined above, the first is by far the easiest to update and maintain, IMO.
Another strategy is that the service has an explicit API or report specification; that way, the team that owns the services also owns the problem of continuing to support that while changing their internal implementation.
Of course, whether the benefits are worth the cost is probably organization specific, just like microservices in general.
The main benefit of a dumb copy is that the production service is not impacted by reporting, only a copy is. This relates to performance (large queries) but also implementation time.
To me this sounds great. And honestly you should do the same thing with a monolith. Nothing worse than "oh you can't make that schema change because a customer with a read-only view will have broken reports".
https://en.wikipedia.org/wiki/Data_lake
The reasons for it are hard to explain succinctly in HN comment but if you look up data lake there will be a lot of explanation. But it basically comes down to "is it better for the data integrators or the data consumers to massage the data?" And data lakes was the insight it's really great when the consumers decide how to massage the data.
Use off the shelf stuff but be prepared to have to move in a (relative) hurry.
“Data lake” may not be the right answer, but GA certainly isn’t.
But yeah, it's funny how these projects get complicated in larger organizations. Personally I would have rolled something even simpler on gnu/posix tools and scripts, in rather less than a month.
I'd argue that (given a large enough business) "reporting" ought to be its own own software unit (code, database, etc.) which is responsible for taking in data-flows from other services and storing them into whatever form happens to be best for the needs of report-runners. Rather than wandering auditor-sysadmins, they're mostly customers of another system.
When it comes to a complicated ecosystem of many idiosyncratic services, this article may be handy: "The Log: What every software engineer should know about real-time data's unifying abstraction"
[0] https://engineering.linkedin.com/distributed-systems/log-wha...
Recently the Netflix engineering blog mentioned a tool called DBLog too, but I don’t believe they’ve released it yet.
Maybe, but your business analyst already needs to connect to N other databases/data-sources anyway (marketing data, web analytics, salesforce, etc, etc), so you already need the infrastructure to connect to N data sources. N+1 isn't much worse.
It will lag behind by some extent, roughly equal to the processing delay + double network delay, but can include arbitrary things that are part of your event model.
Though, it's not a silver bullet (distributed constraints are pain in the ass yet), and if system wasn't designed as DDD/CQRS system from the ground up, it would be hard to migrate it, especially because you can't make small steps toward it.
A quick fix might be to split different customers onto different databases, which doesn't require too many changes to the app. But now you're stuck building tools to pull from different databases to generate reports, even though you have a monolithic code base.
Analytics has very different workloads and use cases than production transactions. Data is WORM, latency and uptime SLAs are looser, throughput and durability SLAs are tighter, access is columnar, consistency requirements are different, demand is lumpy, and security policies are different. Running analytics against the same database used for customer facing transactions just doesn't make sense. Do you really want to spike your client response times every time BI runs their daily report?
The biggest downside to keeping analytics data separate from transactions is the need to duplicate the data. But storage costs are dirt cheap. Without forethought you can also run into thorny questions when the sources diverge. But as long as you plan a clear policy about the canonical source of truth, this won't become an issue.
With that architecture, analysts don't have to feel constrained about decisions that engineering is making without their input. They're free to store their version of the data in whatever way best suits their work flow. The only time they need to interface with engineering is to ingest the data either from a delta stream in the transaction layer and/or duplexing the incoming data upstream. Keeping interfaces small is a core principle of best engineering practices.
It came with many of its own challenges, too! A great deal of infrastructure had to be built to get from O(N) to O(1) infrastructure engineering effort per service. But we did build it, and now it works great.
There is a reason monoliths were traditionally coupled with quarterly or even annual releases gated by extensive QA.
The problem you raise was in fact fixed years ago in our org simply by properly decoupling our "monolithic" app into modules. The top-level build was simply a collection of all pre-built modules. After each dependency update, an automated regression would run, and if pass rate was less than X%, the change was not pushed upstream.
Really, microservices give you 2 things better than that: 1, the ability to combine more technologies; 2, much simpler (horizontal) scalability, since properly done microservices naturally scale well with multiple copies running on multiple machines. Of course, the costs are there, in terms of more debugging difficulties, more complex logging needs and usually a higher minimum overhead.
If you have a process that takes up a lot of a particular resource that others don't (e.g. disk throughput), maybe spin that into its own service. But in my experience, lots of things are just fine being lumped together and don't really exhibit resource use profiles that are all that different from eachother.
But I can see how on a live service-type product, microservices would help a lot more on this front.
And definitely don’t start migrating stuff that matters until you have mature implementations and operations around all those things, and probably many others too.
It worked great and I was honestly shocked when I worked for other companies that plan their releases like they're shipping shrink-wrapped CDs in 1999.
But I have seen people try to build a solution using microservices "because this is the way to go". But this often trades the problems of commits and unit verification for problems of architecture. It can take a lot of work and skill to design a robust architecture using microservices and architectural bugs and failures can be a lot harder to debug, understand and resolve. Pick your poison.
In this scenario, you deploy ~30 times a day. If there's a bad commit, you revert it, and then do another deploy. So there's no rollbacks, you don't lose as much velocity, and since deploys are safe and a reverted commit is just a new commit, the revert is safe too.
You can have individual dev teams, with their own repo ,backlogs, own stakeholders, etc all working at their own paces. They build modules (jars, nuget packages, npm modules) and deploy semver versioned artifacts to a repo like Nexus or JFrog. Any frontend/consumer applications can build towards versions of those modules and upgrade on their own schedule. Only the consumers need to worry about deployment.
This gives you the organizational flexibility but not the infrastructure overhead.
The discriminating factor that makes microservices necessary if these individual services have divergent hardware needs.
Yeah. Not my favorite things either. Luckily there are alternatives in the Java world.
Outside of perhaps stodgy banks, technical folks are not choosing to run their JVM projects today on Websphere. Gradle is pretty darn nice.
Say you version each module and pull in specified versions. It'll work fine, right up until two modules both try to pull different versions of a third module. In practice, you have to update multiple modules at once to avoid conflicts, which, in turn, can require updating other team's code.
It also turns out some tools like Maven don't prevent conflicts by default. You can end up exploring pom.xml files in Eclipse, trying to add exclusions, or figure out which repo is dragging them in.
For example, if master is v1.5.0, and I'm an app that uses v1.1.0, then if the library owner bumps to 1.5.1 for a critical update, I need to go from 1.1.0 -> 1.5.1, which might involve changes I'm not ready for yet. I better have phenomenal integration testing to make sure the update is safe to do.
That could be solved by backporting security fixes to version 1.1.
Of course at a certain point you should deprecate older versions of your package. At which point it's the client's responsibility to upgrade (like for any third party library they would be using).
Occasionally people versioned their module's APIs, which seemed like a cleaner way to handle the module update, as you don't have to update everything at once. They only went to that effort when they realized they'd have to update thousands of callers.
I've switched to Gradle, and one of the first things I do is usually flip on dependency locking, and then even go so far to reject and/or flag anything that doesn't have a clear semantic version scheme.
Gradle's documentation can be a little overwhelming, but there's a lot more to these topics that developers usually overlook:
https://docs.gradle.org/current/userguide/dependency_locking... https://docs.gradle.org/current/userguide/resolution_rules.h...
Many companies would benefit from real dependency locking, and making sure they have reproducible builds. It's tricky, but, it can be a lot easier than "containerization", which I've often heard touted as a solution. (Containers are useful, but you should fix your CI separately.)
That makes an obvious kind of sense, when your number of contributors have grown too large to operate like a monolith... why not operate using models well-established for very large inter-entity communities?
I wonder why more very large companies don't do this, if they don't.
The system only works because of very good integration testing infrastructure and a culture of being able to roll back almost any change that broke you.
If for some reason the more traditional versioned release dependency approach made more sense for other reasons (as the GP suggested they were using it) -- it shouldn't be too hard to write automated tooling to go tell everyone to upgrade for a security release, or even make PR's dependabot-style; if an org already has "very good integration testing infrastructure", adding that tooling for security updates of dependencies is perhaps within the capacities.
Software tends to reflect the structure of the organization that creates it, so this makes sense to me. If you have multiple teams contributing to a stack, eventually it is easier to have the teams work on their own (micro-)service(s).
I recommend to most people that stacks should start out as monoliths though, and move to microservice architectures only when they encounter enough pain. I think starting out with microservices from the get-go just reduces initial velocity with little payoff until you hit a certain scale.
1. monolith 2. libraries 3. services
If you skip step 2, there is a high probability that the services you end up with are going to be just as disorganized as the monolith which is causing grief.
I've been a software engineer for over 30 years and have dealt with companies always trying to jump on the next bandwagon. One company I worked with tried to move our entire monolith application, which was well architected and worked fine, over to a microservices-based architecture and the result was an unstable, complex mess.
Sometimes, if it's not broke, don't try to "fix" it.
I can say the same regarding a lot of what is going on in the JavaScript ecosystem, where people are trying to replicate stuff that works fine in other languages in JavaScript. Mostly because they are only familiar with JavaScript and don't realize this stuff already exists and doesn't need to be in JavaScript.
As an example slip some DOM methods into your code and watch people go into convulsions like an angry zombie on cocaine.
It is several orders of magnitude faster. It’s also how I prefer to code because managing state isn’t challenging and I am not hopelessly paralyzed by imperative code.
DOM manipulation is perfectly fine for interacting with styled documents, the web's forté. If your website amounts to a set of configuration forms and a blog then you probably don't need React.
That said, I’m am quite free to choose not to install npm packages that use Babel or Webpack, and I that’s what I choose when I have a choice.
Certainly don’t have that choice at my work.
https://en.wikipedia.org/wiki/Structured_systems_analysis_an...
https://en.wikipedia.org/wiki/Booch_method
Any language that gets into enterprise architect hands, with projects spread around multiple development sites with several consulting agencies, gets their FactoryFactories and such.
C and Go also don't do operator overloading.
Code generation was a thing in C and C++ during the 90's. Borland C++ 1.0 came with a macro library (BIDS), which was later replaced by a new template based version in Borland C++ 2.0.
And the Go culture of //go:generate goes beyond anything that Java has had on its almost 30 years of existence.
AOT compilation exists since the early 2000's. The only distaste is that most developers don't want to pay for their tools, so they rather used the free beer JDK from Sun instead third party vendors. So only big corporations got to buy the JDKs from IBM, Oracle, ExcelsiorJET, Aonix,....
However now AOT free beer exists on OpenJDK, OpenJ9, GraalVM and although not strictly Java, Android.
The culture of code generation before the culture of XML configuration.
They are there for a reason.
I'm waiting for more fun with JSONSLT the day someone has the idea to make a first limited and poorly documented implementation.
Enterprise ORMs were created in Smalltalk, C++ and Objective-C, years before Java was created.
Poet and Enterprise Object Framework were two well known ones.
In fact J2EE was born of the ashes of a failed Objective-C project at Sun, Neo, after its collaboration with NeXT on OpenSTEP failed apart.
https://en.m.wikipedia.org/wiki/Distributed_Objects_Everywhe...
Things are not created in vacuum, it helps to actually know computing history.
Do you have a concrete example to illustate this, and what issues it causes?
On the surface I'm not sure I agree, if what you're saying is people wanting a certain feature should switch languages to get it rather than build it into the language they already use.
There were a plethora of other languages/runtimes available for writing such programs, but Node made this available to people who already knew & liked using JS without having to switch to and/or learn a new language.
It isn't really, though. This is Some Guy's Opinion™. There are many Some Guy's, and there are countless anti-monolith articles being penned at this moment (probably).
People have different experiences with different groups and different tech stacks and different needs. Results may vary.
Just to give my own Some Guy opinion, people fail with so-called microservices when it's not really microservices but instead is a monolith with artificial walls (in the same way that firms do waterfall but pretend that they're agile by having incredibly frequent "scrums" that are nothing but status meetings). When you actually divided into lots of different projects and teams and they each get to construct their own internal world so long as they provide the appropriate robust and documented external API, it can be absolutely liberating. For some projects.
I'm going to interpret this as "if your monolith is broke, breaking it up into microservices won't fix it"
Sometimes a microservice architecture is the best way to fix your problems. Sometimes it’s the worst. But you’ll only be able to tell after your monolith is, indeed, truly good and busted.
Just never start with it. That’s crazy.
I'm glad we did and today GitLab has a big monolith but also a ton of services working together https://docs.gitlab.com/ee/development/architecture.html#com...
I did an interview about this yesterday https://www.youtube.com/watch?v=WDqGaPGBZ9Y
The best analog I can come up with is monoliths in larger organizations are like a manifestation of Amdahl's law. The overhead of communication and synchronization reduces your development throughput. Each additional person does not add one persons worth of throughput when you cross a critical individual count threshold (mythical man month and all that).
I'm not describing this clearly so I should probably actually commit to writing out my thoughts on this in a post describing my experience with this.
If you’re moving to microservices because the number of people working on a project is growing too large to manage and you need independent teams, great. If you’re refactoring to microservices because “we’re going to do everything right this time,” this is just big-rewrite-in-disguise.
Whatever engineering quality improvements you’re trying to make—tech stack modernization, test automation, extracting common components, improved reliability, better encapsulation—you’re probably a lot better off picking one problem at a time and tackling it directly, measuring progress and adjusting course, rather than expecting a microservices rewrite to magically solve a bunch of these problems all at once.
If the modules of your system are already relatively independent with well-defined interfaces, microservices would be fine and yes would make changes like upgrading the language runtime version easier.
But when I think of messy, tangled, poorly-tested code that prompts people to start talking about needing to refactor to microservices, I’m thinking about different sorts of problems. The messiness I usually see has to do with lots of missing abstractions, lots of low-level code reading and writing directly to files and message buses and databases and datastores instead of going through some clean API. This makes it really hard to change things, because instead of updating some API backend, you have to find and update all the low-level accesses.
Now the problem is, typically when going to microservices, people aren’t looking at the question of, “What common stuff can we pull out to make all our messy code simpler?” They’re taking the existing, messy modules, with lots of cross-cutting shared abstractions dying to get out, calling the existing module a service, and putting a bigger barrier around it.
There are many ways to approach the problem of moving to cleaner, simpler abstractions, and microservices can help. But you can easily go to microservices without addressing all the needless complexity, instead crystallizing that complexity in the process, and many organizations end up doing exactly that.
I don't think this is really the intuitive outcome. Think of cabinets, dressers, shelves etc. They're all basically little messes but they're much easier to deal with than one large mess.
1. Organizational streamlining. If the team working on the monolith becomes to large, then coordinating and pushing out changes quickly can become incredibly difficult. One rule of thumb I've heard is the two pizzas rule. If two pizzas can't feed the team working on a system, it's time to break up the system.
2. Horizontal scaling. If some components of your workflow require much more computing power than others, then it makes sense to break up your system to move computationally intensive tasks to their own services.
While there are lots of other decent reasons to break up a system, if you can't invoke at least one of the two above reasons, you may be shooting yourself in the foot. I think he's dead on when he points out that if you don't have engineering discipline in the monolith, then you won't have it in the microservices.
You build a monolithic application. Everyone works on the same code base. Things are broken up into modules/classes/packages. From the programmers point of view it's just like working on a standard Java project or something similar.
The magic happens at the method and module boundaries. When the application is first started everything works normally. Methods call other methods using addresses. As the application runs some parts of it become hotter than other parts. At some trigger point an included process starts that spins up 1+ cloud instances. Only the hot code is deployed to the instances. If necessary the instance is load balanced on multiple nodes. You configure the triggers and whatnot as part of the applications config. The framework/language would either come with support for popular cloud services or allow you to create whatever system you need to create the instances.
My hypothetical language/framework would proxy all method calls and remap object instances to the new instance(s). If the extracted code cools down enough it is integrated back into the main monolith. At that point proxying is turned off and the methods use address again.
Using this approach you get the all the advantages of a monolith (interface compatibility checked by compiler, not needed EVERY service writing their own http code, etc). Of course you can't optimize latency as easily and merging is harder with monoliths. There's undoubtedly a hundred other reasons why this is a terrible idea.
This isn't really true other than network errors are more likely than a machine getting shut down but you should really be writing your code as if something could go wrong at any moment.
You can just structure your code in this way -- divide hard boundaries in your monolith. Segment things apart the same as they would be in microservices. Have a collection of methods for accessing each segment of code, and don't allow calling anything but that collection (API) from other parts of the codebase.
Set up monitoring/logging. If a segment of your code is using a huge amount of resources, it'll now be trivial to pull that segment out into its own microservice because you already have a defined API and hard boundaries.
I don't see the value in separating a monolith to allow independent scaling, unless there are wildly different performance demands across its components which makes reliably autoscaling difficult.
So with a hypothetical framework like this, yeah it would make microservices easier (in theory, though there are lots of technical problems too) in terms of "look ma, I made microservices", but it wouldn't actually address any of the problem microservice architectures actually try to solve. So, tail wagging the dog.
Some of the technical problems stem from memory working different from API calls: they can fail, the overhead is much higher (shouldn't call in a big loop), can't pass pointers, global state may differ on remote machine. So an application model that tries to abstract away those differences is bound to have problems. Also deployment: remotes will be temporarily broken if the API changes, until the deployment has fully propagated; a team running a microservice takes a different mindset with respect to API versioning than does a compilation unit. And that's really the crux of the whole thing: microservicing requires an entirely different mindset than monolithing, and approaching one with the mindset of another will cause problems.
For example, let's say in an eCommerce application that the shipping calculator is getting hit a lot. You'd like to be able to scale this independently as a service, so you can handle all the requests without also having to replicate all of the other resources, such as the cart persistence, user sessions, etc. that are a lot more memory intensive.
Assuming you allocate different resources for it. If you're using the same instance type for all your microservices, you aren't benefiting from this. In fact, you're paying for resources you aren't using.
You might even be paying more - allocating high memory instances to high memory services, and regular instances to low memory services that don't use that memory. You might be able to get by with only regular instances if you distributed your services in a monolithic fashion.
In my experience most small teams I've seen aren't that specific with their resources. Unless something is obviously super high memory, like a cache, they tend to just use default instances.
This was a pretty high volume API at Azure, and the savings after all was said and done was sadly only around 15K/mo, so it'll take a while to pay for itself. So, I'd say for the average website this should not even factor into consideration.
To begin with, anything with side effects is mostly a non-starter. You would need some way to annotate the boundaries where side effects can not happen, and this already requires considerable refactoring effort.
Even if you ignore this and suppose you are working with a purely functional language (this is probably what the "You just invented Erlang and OTP " meant), the overhead of the "smartness" can be pretty big. If you see a "map" over a list, you know you can parallelize it, but which one of the 100 "map"s should you really parallelize? If you try to be too smart, you will burn a lot of effort and may not get a payoff. If you try to be just a little big smart and you choose the wrong one you shuffle a lot of data for nothing.
In some way it seems to me a lot of today's big data systems like (surprise) MapReduce and Spark do exactly this, they offer ways to explicitly mark those boundaries where you want your program to be parallel (and require that you obey their rules about side-effects there) and contain some of this "smartness" for how to distribute your data.
Even in OpenMP you also have something like this, you can easily say "this piece of code is a task and the runtime can decide to run it in parallel", but you need to also tell it all the inputs, outputs, etc. (since figuring this out automatically is hard if not impossible) so it doesn't fit very well in a "big ball of mud" project.
In the (near?) future as computing power is more abundant and cheaper but the gap between sequential and parallel computing continues growing ever wider I can see those general "smart" approaches paying off, even if they are much worse than hand-optimized code. It only needs to be better than the average programmer.
At this point, many teams institute a merge queue. Which works only until more devs are added to the teams, which makes the merge queue very long and it can take several days in the ideal case to get things merged.
A monolith is hard to tune and often ends up being a money pit.
Monolith => Microservices => Monolith
I wouldn't say the journey was completely pointless, because the fact that we had to deploy 10+ services to make a single environment whole required us to build extremely powerful CI/CD management tools that we happen to be able to re-use in the (new) monolith case today. This journey was also a really good growth and learning opportunity for the team. Everyone who has touched this project and has seen both ends of the distributed<=>monolith spectrum is now radicalized towards preferring the monolith approach.
On the trip back into a monolith, we didn't just stop with the binary outputs of our codebase. We also made the entire codebase a monorepo. We have a single solution (VS2019) within that monorepo which tracks all of our projects. Prior, we had upwards of 15 different repositories to keep track of. Being able to right-click on a type, select "View all References" and legitimately get every possible reference to that type across the entire enterprise is the most powerful thing I have yet to see in my career.
I'm very happy with this approach.
> Splitting up your application in many services requires a lot of thinking and designing.
In my experience, you spend a bit of time asking "What does my project do, and how do I decompose it?" and that's about it. The same thing you'd do with a single binary that you decompose into modules.
> Challenges with syncing, communication etc are not always easy to deal with.
Depends on what you're doing certainly. For me, I'm basically doing an ETL and analytics pipeline, so I have no issues there.
I have found it much easier to reason about code and boundaries. I did a ton of experimentation with code and services (I didn't know how to build it when I started, so there was tons of iterationt) and the ability to just rewrite any service fairly trivially helped a ton.
I haven't had to do any "splitting". Merging services is trivial, splitting is not. I'd much rather say "Ah, the synchronization here is too hard, I'll shove it into one process" than "How the hell am I going to scale these two modules separately with all of this shared memory between them?".
I've gone through a "We have to start splitting things for reliability/performance/conway's law" and it's years of very difficult, dangerous work.
It's really a "to each their own" but I don't think I overengineered things at all. Microservices just make sense for my use case.
I ended up restricting database access, but wouldn't have thought that was necessary. To me it was obvious all services should only communicate via their public API, but I guess that's not so much of a no-brainer as I thought.
I had to make an authorization service and the idea to use a single authorizer to handle token generation and authorization was shot down by management due to worry about "lambda startup times". I complained that the startup time is less than a second (for nodejs) and honestly would not be an issue. They gave the task to someone else to have the token generated in their service. The developer did it by copying the code I'd written into their own service verbatim.
This is why I don't like microservices as we do them. Management would rather we wrote small programs with a lot of duplicated functionality in many repos instead of writing a large program where we can enforce some discipline. This is also better for them because they meet with us individually to ask for functionality rather than have design or architecture meetings where we can push back on implementation details.
On the other hand they are often sold as a way to increase developer velocity. And I do sometimes wonder if that is the case (based on personal experience).
People are selling it as easy. You can deploy them independently, each service is small with low complexity.
That it is not that easy in reality and that the total complexity increase is so big is not told as often. There is a team at my work with 6 developers developing something that processes a couple of gigs of data each day and has a front end serving maybe a hundred users total, I would guess 20 would be logged in at the same time.
Kubernetes, 10 microservices, 2 different types of databases.
That's undeniably true; but, it's too easy to slip from that to "good engineering doesn't require good architecture, good technology, etc." I think it has to be stressed that adapting the big picture features of a solution is as important as adapting the finer features.
It seems in many cases teams have decided to port their highly coupled monoliths over to highly coupled distributed monoliths and now they have the worst of both worlds.
Let's call them "tiers". I think 3 is a reasonable number..
In any case, there's already such a term. It's called SOA: https://en.wikipedia.org/wiki/Service-oriented_architecture. Microservices arguably evolved out of this.
Services is more like feature folders? Where you define everything you need related to a VERTICALLY sliced part of your application, ie. Products, or whatever.
I feel like if you end up talking about eventual consistency or needing ids to be passed around you've surely built it wrong and you've split what should be a single service into a mess just to feel cool having more services and needing the hyped new tools to manage them.
The touted benefit that you can scale the bottleneck separately also suggest the need for 1 monolith at most 1 or 2 services. If you have more than that many bottlenecks the code is doomed anyway presumably?
According to most of the intent of the definition of "microservice", we do have many of these within our monolith. There is a nice big folder in our core project called "Services" and within this lurks such things as UserService, SettingService, TracingService, etc. Each with their own isolated implementation stack+tests, but common models and development policies. All of these services are simply injected into the DI of the hosting application. We are injecting approximately 120 dependencies into our core project and it is working flawlessly for us in terms of scalability and ease of implementation. Microsoft DI in AspNetCore is awesome. For those trickier cases we will usually just pass IServiceCollection to whatever needs to arbitrarily reference the bucket of injected services (e.g. rules engines).
I think you can have the best of both microservices and monoliths at the same time if you are clever with your architecture and code contracts.
The big thing that people often underestimate is the complexity of facilitating communication between micro services. Not only you need to plan the API well, but also figure out the routing and you also get overhead.
What I'm not seeing is any attempt to go in the opposite direction. A compiler should be able to look at ordinary code and slice it up into microservices automagically, converting the header interfaces to API specifications like OpenAPI/Swagger. We should literally be able to write a monolithic program in any functional or C-style imperative language and get a conversion to a bunch of lambda functions. If that doesn't work, then something is seriously wrong (probably having to do with determinism, like inadequate exception handling for timeouts, etc).
So frankly, the first day I saw lambdas, I was skeptical. I don't understand the point of writing all of the glue code by hand. Incidentally, I reached this same conclusion after manually building a large REST API around the JSON API standard just before GraphQL went mainstream and made a mockery of my efforts.
I think that the HTTP spec and things like separation of concerns serve a purpose for human readability. But we're well past the point where the gains made by the early internet are providing dividends in today's highly-interoperating stuff like Rust, Go and Node.js. Basically 90% of the work done today would be considered a waste of time (bike shedding and cargo culting) in the 1980s and 1990s. Just my two cents.
Isn't there room for a middle ground with modularity that can live in between a full blown monolith or a full blown microservices pattern, particularly for operations that are more medium scale?
I think people sometimes forget that a healthy level of pragmatism is what keeps shipping. Just because someone said "microservices is the new all" you don't need to do it. Just because some said "monoliths are the future" it does not have to be true for you.
Take all these as case studies and solutions that apparently worked under specific conditions.
I really don't like the dogmatic view of architectures. Leaves no room for craftsmanship, and it's only really useful for creating code monkeys that have to follow a spec and need to be interchangeable cogs in the machine.
That is not what he is commenting on here though. It is that microservice architecture is starting to become the default pattern on how to build applications for a lot of people. Instead it should be an exception when you reach those very specific problems that most people don't have.
Thanks for the downvote, though.
It would be more constructive to give reasons why one argument is better than another though, rather than resorting to status, and you did not reference any of the commenter’s points.
> Thanks for the downvote, though.
I don’t have the karma required to give a downvote. I’m not sure who you should thank.
I am suggesting that the OP has little business writing off a legitimate expert's opinion in a domain where they are highly qualified to comment. This isn't controversial.
If you had the karma to downvote, would you give a hard time to the OP who started with "The author does not seem to understand when to correctly apply microservices."?
Once you've iterated on a monolith enough to see which parts are relatively independent and would actually benefit from decoupling, then you can move them into separate services.
One example that comes to mind: I wrote a recommendation service that also handled user feedback events. This was the easiest way to start. After about a year I saw that we were iterating faster on the event processing than on the actual rec delivery. We were also deploying this monolith across more machines mostly to scale up event handling capacity. So we broke the high volume event handling out into a separate service that was smaller and optimized exclusively for event processing.
I've seen X be a dozen things: UML, databases, User Stories, Functional Programming, Testing... It's too much to list.
Yes. If you do it that way it will hurt, and you should stop. I don't know this author, but I suspect that many people who jump into microservices are not getting the foundations they need to successful. The idea that microservices are just broken-up monoliths is a big clue. They're spot on about marketing and spend, though. In this community we're quick to hype and sell things to one another whether it's a good idea or not.
I've seen some great criticisms of microservices, some of which made me pause. Now, however, I think there's a reasonable way through the obstacles. It doesn't have to be a mess. Nothing is a magic bullet, but about anything will work if your game is good enough. You don't buy a bright and shiny to make your game better. Doesn't work like that.
I'll delete the comment if I was unnecessarily cruel or missed the sarcasm. It was not intentional. But it is important to understand that you want to think of persistence and deployment coupling as independently of your microservices strategy as possible. The vast majority of problems we see with people implementing microservices is people carrying baggage over from some previous project or pet technology. K8S's great. It's just not relevant here.
It is very relevant though as it has become so tightly connected with microservices and if you are one of the most well known people in the Kubernetes world you will see a lot of applications that should not be microservices.
When microservices started to catch on, it was just a name given to a really good solution to a specific problem. And there are plenty of problems for which creating an independent service is a great way to manage both technological and organizational issues. But it doesn't just magically solve those issues -- you can't apply that model to everything just because -- it has to fit the problem space.
I had a conversation on Reddit with a developer whose application had over 3,000 independent micro-services. He was very proud of this solution. But I can't imagine that could be anything but a monolith with function calls replaced with network I/O.
I consulted with a team last year that was moving to microserivces. They bought BigToolX and had already created a disaster ... and they weren't even through their design. Most all of what they were doing was just best practices in some other paradigm. It was painful.
Whenever I fall in love too much with a technology, I get paranoid. There's usually something I'm missing.
The text of the article reads in that same vein, but given that the title is “Monoliths are the Future” I think the author’s original intent was to describe the advantages of monoliths and point out that microservices have advantages only in relatively rare cases. Too bad they made the article about microservices instead of monoliths.
Sounds to me like a good opportunity for the author to write a followup post.
This 100x. If you aren’t able to maintain a monolith you will most likely mess up microservices too. Every approach has its own set of trade offs and problems. if you know what you are doing you can make things work.
It's an excerpt from a podcast.
For example some time ago, I talked with devs that were about to change their monolith to micro-services. I pointed out that having decentralized the data is going to be tricky deal with. It got immediately dismissed as not a problem, because all the services are completely independent. Couple of months later they were struggling hard, because, turns out, a business needs to be able to ask questions about all its data, not just per service.
Sure, a problem that can be fixed. But I got the impression that they haven't spent 10 minutes looking at potential downsides of their decision before making it. In a similar vain, people that equate monolith with spaghetti code and then end up with a spaghetti system almost immediately.
With all due respect, the argument that microservices can work is not an argument for doing microservices instead of a monolith.
By default a monolith is simpler, lower latency, has lower operational costs (RPCs are not actually free!), tends to be easier to refactor, leads to less duplication of code, and has better tools for traceability. (Do not underestimate the value of stack backtraces!) With best practices (that few do), all of these problems except the latency one are solvable with microservices. But you should not expect to solve them in most organizations.
Given this, you should only adopt microservices if they solve a real problem. For example if your codebase is too big for a single server to hold it, or you need extreme horizontal scalability, microservices can be wonderful. But most people using microservices do not actually have those problems. Most organizations that are trying to use microservices would be better off with monoliths. Eventually reality will settle in and they will realize it.
Incidentally this is not a new debate. At its heart microservices vs monolithic is the same as microkernel vs monolithic kernel. It is worth reading https://yarchive.net/comp/microkernels.html for Linus' criticism of microkernels - much if it applies directly to most microservices deployments.
Also having dev teams across time zones is itself a challenge. The devs are cheaper, but integration is worse. In Steve McConnell's book Software Estimation the industry average seems to be that development is over 40% longer, and defect rates also go up.
That, however, is a debate for another day.
Your customers do not care about your monolith. They don't see a monolith; all they see is features. Untangling it may or may not be the right choice.
In a certain set of situations, the path forward, instead of trying to untangle your monolith is --if you so desire-- create new services actually be true microservices, and keep your monolith as-is.
There are plenty of clusterfuck hybrids out there with services sharing database state etc. Anything can be an antipattern when you add people into the mix.
The author makes good points though, there are many places doing microservices because it's the hip thing to do and a monolith would easily suffice. But if you have independent software teams in your org that should be able to deploy code independently, then microservices makes a lot of sense.
As in all things engineering - it depends :)
But when you have chosen a cool new microservice architecture for your team to implement and you grab that small user story that spans 3-4 different services things suddenly went from. "Hey, easy implementation and refactor and the compiler will tell me if I fucked up" to something much more time consuming and error prone.
In an ideal world that would not happen of course. Just like it in an ideal world a monolith is built correctly as well.
the things that most people don't get is that: microservices are not free (now you're doing all this devops stuff N times and you have to think long and hard about changes that need to happen across api boundaries). The anti-pattern is that you take your monolith and you split it in 10 but apart from actually doing all this work you still treat it as a monolith (ie you still do mono-repo because it's convenient, the deployment still happens at the same time for all services, you centralize everything when it comes to logging and metrics and you even force people to do things in a certain way when it comes to their service). Everything grinds to a halt and now you're more concerned about "growing" the team to fix the issues that popped up and maybe chasing the new shiny thing to keep your resume up-to-date. Even worse people start feeling like they "own" their service and now the decisions that are made are maybe locally optimal but who cares about global optimization.
So my take is: start with a monolith and in 85% of the cases you'll be just fine forever. you don't need all the bells and whistles to get the job done. Introduce new things only so solve actual pain-points and when you do actually thing through what it means to introduce them (so go N->N+1 and never 1->N)
This is the greatest illusion of our time. We tend to think of all things in the world as "middle" like apples and oranges. If this is your viewpoint then you are biased, the truth has equal probability in being in all extremes just as well as the middle.
Data isolation. Allow individual teams/services to own their own data stores, and prevent any other team/package/service from reading or writing to their data store, and inadvertently breaking the associated invariants.
Performance isolation. Prevent one team/feature hogging too much memory/cpu/io, and negatively impacting every other team as well. Debugging performance hogs in a sufficiently large monolith becomes infeasible at a certain point.
Deployment isolation. Allow individual teams to made code updates and deployments whenever they want, without having to be tied down by a company-wide deployment process.
Language/dependency isolation. Allow different teams to use whatever language, dependencies, and dependency versions make most sense, for their use case.
At bigger companies that have hundreds or thousands of engineers, monoliths simply do not scale, and need to be broken down into more manageable pieces. It's unfortunate that smaller companies start cargo-culting these same practices without thinking critically about whether they actually need them.
But we keep reinventing the same solutions at each scale. At one time we had to invent functions to enforce segregation of responsibilities and create abstractions and shorthand. We had to group these together in modules and libraries. We had this clump of programs running on a computer that we had to organize into an operating system. Now an operating system is nearly a program or function and people are regurgitating the Unix philosophy and the end-to-end principle like it's a new thing. In the end, we're going to wind up with a well-architechted series of integrated microservices which present a comprehensible interface to users through a handful of abstractions that have proven useful over the years.
Computing is cheap enough that we can now talk about meta-computing, a higher level of abstraction from a computer, which is multiple layers of abstraction on top of eachother. Now we just have to build the next layer. And I think it will basically be a sort of meta-operating system. The same things, but we'll call it "orchestration" and "microservices" instead of a file/process/whatever manager and threads.
At the moment, however, we're still offering piecemeal services and products and so we don't have many fully formed concepts of what it is to build a cloud system. So things are still a bit chaotic, but at some point in near future we'll get there.
Trying to make everything work as microservives just for the sake of it, or because it sounds cool is just a terrible idea.
Start out with a monolith, and if you later see a need to create a microservive, then do it, when you have more knowledge about the bounderies of the service.
I love creating high performance services and playing with containers. It sure is cool with microservives that can scale linearly over a lot of machines. I also enjoy using the latest frameworks.
But guess what, my first ever service is just using a cheap dedicated server, serves an average of ~250 highly dynamic webpages each second while still using less than 7 % CPU, on PHP and MariaDB. Last 12 years have resulted in about 6 hours of downtime. A couple of hours planned, a couple as a result of denial of service attacks and a couple when there was a power issue at the datacenter.
So what I'm trying to say is that more complicated doesn't mean that it's better.
Some are best served with aggregates; some with monoliths.
For myself, I have always developed in a "layered," and "modular" manner, with discrete subprojects; each, given its own configuration management and lifecycle. The resultant applications tend to be "monolithic," but some are parts of a larger, loosely-connected architecture.
Works for me, but YMMV.
There are huge advantages to both patterns. For newer systems, if there's a clear enough split such as "backend" and "frontend" (where frontend is a statically-hosted SPA) then it could be advantageous to keep the codebases and deployments separate.
If data is shared between services, then keeping the code to interact with the data all in one service is likely most useful.
I like to use a few services, with one often ending up being the large "monolith" potentially with a few supporting microservices on the side as it makes sense. "As it makes sense" means that the service has a specific individual encapsulated concern. Billing could be a good example, depending on how it integrates with the rest of the system.
I find microservices very useful to encapsulate independent concerns and for experimentation (don't want to rewrite the whole app using some new tech, but the billing service is small enough to give it a shot). The main problem points are the glue that holds it all together, duplicating code shared between services, and changing apis / data schema.
Ultimately, it's best to know what you and your team is/will be most comfortable with managing based on everyone's skillsets and the product at hand. If you spend time to understand the differences between the patterns in practice, and remain realistic about the advantages and disadvantages of both, you can arrive at an informed decision that works well for your team.
And lastly, make sure you pick something and then build your product. These details don't mean anything to your customers. If you made the wrong choice, you'll know when it's the right time to switch.
At the end of the day if you don't architect your system correctly Monolith / Micro Services won't help you.
For me and my team now I have 1 ideaology regarding this topic. I don't care whether it's monolith or micro service. As long as I can have clear segregation of responsibility between the different modules. Our company now has a monolith (core banking app) that has modules that handle their own responsibility and communicate whether its over http or internal communication bus we developed it doesn't matter. We can easily move modules out into a separate service if we need.
What determines the factor of whether we move things into it's own service? A few things. If we need to deploy / scale something independently we will decide to take on the overhead and move things out into their own external service. Or if something has a specific security requirement that will increase the complexity of the overall system we will isolate that and deploy it separately. Otherwise we keep things as a monolith. For example in Banking there are many things like the ledger / transaction data that are highly sensitve that require certain security requirements like being hosted on a cloud that has certain standards. We will deploy this part on GCP. But
People seem to love to stereotype and find a one solution fits all. There is no such thing. Everything in engineering requires a deep level of understanding of the problem and making choices and the problems present itself.
I believe most apps can start their life out as a monolith, and can grow and divide as needed. There just isn't a one size fits all for anything in tech. That's what I've learned.
Microservices may be the solution to this problem right now, but I believe someone is going to come up with some other solution (tooling etc) that allows you to get the benefits of migrating to microservices without having to add a unnecessary network layer just to solve an organizational problem.
- Every change you made could break things elsewhere in a surprising way.
- Deploying changes was a nightmare - we had volounteer teams be on daily rotations of merging because merging was so incredibly difficult.
- Different teams and groups of teams would acquire this tribal knowledge of how to do things in their corner of the system. You needed to acquire the tribal knowledge before you could start to work in that region of code.
- Build times were atrocious! Sometimes folks would come up with a way to only build a part of the application and that would be considered innovation. "hey, 15 minute build instead of 1h!"
- Our QA's were STRESSED
Today I am convinced that the product was several applications masquerading as one. I am not willing to subject myself to that again.
One thing I've noticed is that big tech is taking advantage of giant mono-repos, while everyone else is stuck with 10s-100s of git repositories haphazardly connected and managed. For example - most off-the-shelf CI systems and VCS platforms smaller organizations are using are per-repository (GH, GH Issues, CircleCI, etc).
Managing micro-services would be a far easier task when all of the services (and infrastructure as code) live in the same repository, changes can be staged across multiple services at once, and tests are automatically ran for only the necessary dependencies.
Are there solutions for effective mono-repo management outside of FAANG? Am I wrong? :)
The big blocker most monolith faces as the application gets bigger and is deployed into more and more machines is that _releases becomes a bottleneck_. Scaling monolith's applications are difficult because partial rollout is usually not possible as "services" are often tightly coupled.
Micro-service architecture forces services behind a set of APIs. While the APIs may have breaking changes, each can be independently deployed. In other words, teams can do releases at their own pace.
The main cost-benefit analysis here is how important is independent releases vs the cost of operational overhead?
But there are problems that justify a service. Decoupling code is just not one of those problems.
Also I feel like all tech goes this way. Years ago to do "big data" you had drill, kafka, HDFS, pick your cloudera or hortonworks, roll up your HBase, your storm, spark - hire a team to install it.
It seems like now all we do is purchase Elastic cloud, and write a one-off spark script or pandas job and call it a f*cking night.
I guess things go in circles.
Just look at the US legislature right now. Anyway, it doesn't have to be Monoliths vs. Microservices. It can be a compromise. Perhaps the microservices are a bit less segmented than we have been imagining. It might be OK for a microservice to do more than one job. As the highlight shows, the fundamental ingredient is Engineering Discipline. If we strive for that it might work out in a Monolith, Microservice, or somewhere in between.
Our concept is to allow the decoupling of microservices, with the tooling of a monolith. Kinda hard to describe and we haven't done it yet, but basically give you the ability to write it as a monolith, but also have the separate scalability/deployment of a microservice.
It's true that a microservice doesn't magically create cleaner code, better designs, or anything like that. It can actually make all those things harder. Designing good remote APIs is hard, maintaining consistent code quality over lots of different codebases is hard.
All a microservice does is give you a way to independently release the code that lives behind a small chunk of your larger API (e.g. http://apis.uber-for-cats/v2/litter-boxes). This is why a good API gateway that's built for microservices is one of the first tools you actually need, and can get you surprisingly far.
It turns out that despite the complexity, this is an enormously valuable capability in a lot of different situations. Say you have a monolith that you can only release once every six months and you urgently need to get a new feature out the door. Or maybe half your code can't change very fast because it's mission critical for millions of users, but the other half wants to change really fast because you're trying to expand your product.
Of course the big bang refactor into microservices that he describes isn't really going to help you in any of these situations, but then again big bang refactors don't tend to help in much of any situation regardless of whether microservices are involved. ;-)
Microservices make a lot of sense when you release often, run multiple versions, have a lot of people working on independent components or have a lot of different scalability needs. A monolith can only scale in its entirity and often only vertically. That means that even if just one component cannot be locally optimised the whole application has to scale up.
If you only have a single application or task to build software for (i.e. a CRUD system for a CMS) then it makes no sense to split that out. Just like it makes no sense to build your own crypto, do your own CRM, do your own RDBMS, or do your own filesystem for that matter. That would just be adding overhear and engineering complexity where none is required.
while bad engineering will be bad engineering no matter how it's engineered, that doesn't make a whole pattern bad just because a lot of people apply it wrong. Goes for microservices as well as monoliths. (and XaaS)
Continuous deployment alleviates merge/coordination issues by integrating small changes frequently, which makes conflicts rare. Deploys are safer, again because you're deploying small changes often. And if something bad does go out, you can "roll forward" instead of rolling back, by reverting the bad commit. This is less harmful to velocity, because it doesn't require rolling back the other good commits in the deploy along with the bad ones.
I have less experience with microservices than with continuous deployment, but they seem to bring a lot of problems. Microservices take the fixed costs of deploying an application and multiply them by the number of services. Instead of centralizing one team to update dependencies and infrastructure for the whole application, every team has to spend 10-20% of their time doing that work. In the monolith case, everyone on the engineering team is familiar with the single codebase and architecture. But in microservices land, there are often more microservices than engineers. So when an engineer leaves, they pass off a whole pile of code, infrastructure, and architecture patterns that almost no one has any familiarity with. I do think you could avoid these problems, but overall microservices seem very high risk for little reward.
The one case I really see for services is when you have tasks with different load characteristics. But in that case, you can still have N monoliths (for small N), rather than the massive proliferation of microservices.
This requires quite a big application though for it to be worth it in my opinion.
You might want to always breaker up a monolith but there is indeed little reason to do it with micro services. You could just use modules. Or better breaker it into a number of libraries with well defined interfaces, which you then compose into one monolith binary.
But there are very good reasons to split out some code into services (which might or might not be micro services, just not in the same process).
One is that it (easier) allows you to use more than one programming language. Normally you should avoid that, but there are sometimes reasons for it for example if 80% if what you need is implemented in a library available in that language.
Another one is you can have different reliability constraints for different parts of the system. (Like number of instances handling load parallel).
Another one is reuse between different systems (e.g. sharing of user management by e.g. using OpenId Connect).
Another one is that you can upgrade part of the system without stopping other parts.
....(a bunch more)
So in the end I would brake it in parts and compose that parts into a number of services but I would not bother with the whole "micro" part and other cloud marketing bs (because that's what it degraded to).
But there are many other valid reasons for services - different deployment cycles, better resource utilization, faster and safe deploys, etc. It's just about using the right tool and thinking about implications.
I don't know if such a framework exists, but I really want a system that abstracts this to a certain degree - while the contracts between parts of the system are defined, whether any module works as a service with its own deployment policy over network or as part of a monolith is not expressed in application code but as a configuration, and code generation handles the underlying logic. So you can write your app as a modular monolith, but when you think that for operational reasons there is a reason to spin off some part of it as a service, you reconfigure your build rules instead of your code.
If it's important to know how many blue widgets are bought at night in Europe, vs. how many blue watcha-ma-call-its are bought in the evening in the US, and your location, orders and product data are in separate micro-services, you are kinda out of luck.
And, as mentioned by others, replication and API wrappers on micro-services suck for reporting.
If you built an eventing system, you'd be better off tapping into that to update the central reporting data store (warehouse/lake/etc.) I've used this myself to "good" effect. (some chance of failures, a little behind the times, etc.)
The central database may be "monolithic" in nature, but at least you'd be able to report on the data. If you expect to modify data in the feeding micro-service's databases, then yes, you do have a monolith. But, if it's "just" for reporting, it's like a dynamically-updated replica of the pertinent data for your reporting.
This, in a single sentence, captures all of my misgivings/discomfort/etc with the mad rush to micro-services in my organization. To whit...if we had the operational maturity to really effectively take on micro-services, we wouldn't need to rush into it.
It was only ever useful for massive deployments that only massive systems like facebook and others needed. However their engineering teams dominated the discourse and others followed, pretending if they too had the same requirements even though their engineering teams were small and their systems far simpler.
Cue Hadoop, AI/ML, block-chains
1. Easier for users to see "who owns what" (albeit a module pattern could fix this as well).
2. Different hardware resources or scaling for different parts of a monolith really isn't possible. If one module requires 16GB then everytime you scale a horizontally you must have at least 16GB, you're at the mercy of your worst module in the monolith.
3. Deploys are very difficult, and as you scale to over 10 developers increasingly becomes difficult to push up (it takes one persons bad commit to hold everyone in the organization from deploying).
4. Security boundaries are easier to define, each "module" in a monolith effectively has access to all resources for all modules.
5. Poly-languages are easier, albeit depending on the base language, you could do a lot of transpiling on a monolith but.. ew.
6. HTTP status codes and request paths can give you a clear view of how calls are happening in your system; in a monolith you'll only get stack traces generally on errors, not on successes, usually you need to invest more in static analysis and APM stuff for a monolith.
7. Microservices can be cheaper when you scale, you don't have the GCD of memory/CPU/disk requirements as you do in a monolith.
8. GCD of implementation details, if one request requires a sticky session, all of your requests require stick sessions...
9. More complex and long builds, most monoliths have component-based hot reloads, but even those can take 30s to a minute in my experience, and a full build, that's at least 20.
10. Harder to unit test, this can vary by language but without clear boundaries and resource definitions monoliths can be very tricky to unit test, microservices/distributed monoliths are inherently smaller with clearly declared resources so it becomes easier to find where and how data flows in them.
One more thing I really appreciate:
No more massive juggling acts to upgrade the language or libraries. I’ve spent way too much time having to worry about how to upgrade a monolith to the next major version or three of .NET and C++, worrying about major version incompatible in libraries etc.
With smaller services you will have to fight this battle many times, but each battle will be manageable and lower risk.
I feel that is a way too low number. If don't have good enough engineering practices with branching, pull requests/code reviews, unit/integration tests to handle 10 people then microservices will be painful as well.
I would say like 3+ teams at least to really justify it? You can do it earlier but I don't see it as a necessary benefit.
>7. Microservices can be cheaper when you scale, you don't have the GCD of memory/CPU/disk requirements as you do in a monolith.
Yes, but by default they are more expensive until you reach a certain scale and it needs to be a specific type of scaling.
> We’re gonna break it up and somehow find the engineering discipline we never had in the first place.
Indeed. I worked in a company that had a monolith, but the project was structured in modules. Every module had a Facade, which was the official way of communicating. Although in practice you could access other modules' entities, you weren't allowed to do that. As you can imagine, this rule was broke many, many times. Developers would look at the entity and see the data they want was there and ignore the facade right way, plain and simple.
If you split your project into separate services, and those services aren't in the same runtime application, there is no way to break this rule anymore. The team that didn't follow the rules has no other choice, it has to go through the APIs. Even better, they won't design the API themselves most of the time. Whoever maintain the service will want it to be cohesive, and will not care that much about the other team need to an urgent fix. Putting workaround becomes way harder, and this change alone improves design a lot.
The second point is that anyone that shared the same service/application with another team probably faced the situation where you couldn't deploy (or was too afraid to do so) because the other team pushed a lot of new code to master. You suddenly don't know if the deploy will break everything or not. When you see, you're spending a lot of time coordinating with many people about whether you can deploy it or not. Something that should be in production if a few minutes sometimes get delayed for days.
Of course that microservices are not a silver bullet, and there are teams that will benefit a lot from a monolith. With that said, I find hard to believe that monollith will come back in companies where the development team grew to be more than a few developers, because the trade-offs are not worth it.
Some of these are general distributed-system problems. Some would be less severe with a better but still microservice-based architecture. But in practice the microservice message that a lot of people get is that you should make every trivial bit of functionality its own service, and that road leads to disaster.
The underlying technology is a bigger deal than people give it credit I think. I've written frameworks and complex applications at prior workplaces to try to manage microservices well. Now (cloudsynth), we use go + grpc + typescript and everything feels like it can be isolated/sharded if and when it needs to. Golang and webpack have great tooling for splitting things off, isolating dependencies, etc.
Sometimes you don't have to live in the bimodal world of MicroServices vs Monolith.
1, back-end services with clear boundary, that decouple concerns based on dev teams' domain responsibilities, with less dependency among each other,and respected source of record. This is very much the "micro-service" is for.
2, middle tier services to consolidate or aggregate back-end APIs to serve the front-ends (especially the mobile apps) and take care of the business logic. Back-end guys all love micro-services, but someone must put them all together....GraphQL so far seems to fit this bill
3, Analytics and reporting, this is a totally different animal from the product development, and have almost opposite requirements. This is where whatever your ETL or Data Lake or Data Pipeline is used, along with your preferred BI or analytics tooling.
The fundamental problem a lot of companies have, especially fortune 500 type legacy shops, is that they haven't accepted(at the c-suite level) that they need to become tech companies to compete with startups that are eating their lunch .
Switching to micro-services to try and deploy more features while starving your development teams of talent and funding won't make you a tech company. If you want faster development + more features then you need lots of development teams, and large development teams means micro-services so that you don't have a slow to change, interdependent mess after a year or two.
Microservices, when done right (driven by well defined bounded contexts) are simpler to develop and iterate against; but that's not why we do Microservices!
You should not do Microservices without considerable experience in authoring integration tests, a clear understanding of the domain, observability tools, and a team that can handle debugging distributed system.
Bonus: You do not need a distributed system if you are working out of a single data center. You should not do Microservices if you think they're cool. You should not title your blog post claiming Monoliths are the future. If your future has a horizon of never scaling out then yes I guess they are ...
Many companies move to microservices so that they can evolve different parts of their platform at different rates, and invest differently in different business domains and product applications. Attracting talent for a problem in higher demand is one example of the lever you can pull, but so is writing a part of the application in R for data science or Java for stream processing, and hiring from a richer or different talent pool as a result.
If it's about monolith deployment architectures, the use case is really important.
If it's about monolith code bases, you need to define what a monolith code base even means, because that could mean anything. Are we talking storage size, custom written code lines, framework architecture, or just the number of people and/or teams building the underlying technology?
But it also requires a heavy investment in dev-ops and on-call issues. Because when one small thing fails, it becomes catastrophic in ways you can't imagine. So there's a huge tradeoff between engineering convenience and actually customer impact and uptime risk.
Monoliths aren't the future. They never left. Rather, they are still an option, along with microservices.
Blindly adopting anything is silly and error prone.
It gets tiring hearing the same advice preached every few years about a new techonology.
Note: no idea what I'm talking about, I'm genuinely curious if that's a valid solution.
For me though the most important thing is grokability. Our monolith is to a point literally no one on earth can understand the whole thing.
Even if the system is complex, the individual deployables being fully understood by some number of engineers is extremely valuable and drastically reduces search space for the cases where things don’t go as planned
Well, then you're really going to love it when the concerns are spread across different codebases connected by APIs!
Not that this isn't possible in a "well engineered" monolithic system, but design constraints are usually better than hoping for engineering discipline.
1. Use a package manager, and export interfaces to a given service in each of the consumers. It's great that GitHub now offers it for node. 2. Create a DAL library for I/O to a given database (e.g. PostgreSQL, or Mongo) that be consumed by other services. 3. Enforce styling company-wide with one source of truth. We use GitHub Actions, so we can enforce styling with a shared GitHub Action.
People look at what really big successful companies are doing and draw inspiration from that. Problems come when they mix up cause and effect, and then view things through the wrong lens.
As you say, microservices are a mostly organisational, partly technical, effect of having to scale a huge techincal org. But then when taking this end state and viewing it purely through a technical lens (as, naturally, technical people are wont to do) it's rather easy to convince yourself that it's actually the cause of this huge technical org's success.
Of course he looks at Netflix and Google as best practices but that discussion probably spent more money in salary than our server costs.
I think a lot of the principles in classic OOP design (SOLID) can be applied to microservice systems: Classes/Objects <> Services.
This debate of monoliths vs. micro services is like debating what integers you sum to arrive at 10. 10 + 0 vs. 4 + 4 + 2 ?
Everything has tradeoffs. Let’s focus our discussions on methods for understanding the problem set and weighing the tradeoffs of potential solutions.
So write code you can deploy today, monolith or micro serves, in the not-to-distant future we'll be able to cheaply refactor it at scale into any style you want.
By this principal, the right answer is probably along the lines of: a few services, carefully curated. Something like "macroservices (plural)".
So a container will have group of services rather than hosting single service. Similar will be happen for databases where different databases will on same host.
We've seen great success with a Mono-repo that enables the sharing of code across Micro-services, and enforcement of code and deployment processes.
However, using more services and being less concerned about the servers underneath is an opportunity that shouldn’t be ignored.
Both arguments are wrong. There is no substitute for engineering discipline and no paradigm will save you from a lack of it.
I care not whether you build microservices or monoliths, but please sir, do not blame the paradigm when your team can't do anything right.
If you have microservices then monoliths are the future.
There is also a clear distinction, in my mind, between microservices philosophy and 'macroservices', as I call it. Buying into a system with more services running than engineers is very different than having a number of teams, each working on their own single or handful of services.
I would argue that the organizational scaling derived from microservices resembles diminishing returns somewhere in the domain between a single service (monolith) and more services than engineers (microservices).
That's it.
tl;dr: If you don't understand the problem domain, build a monolith following sensible engineering principles to get going ASAP and then split it out when you understand where the functional lines actually are.
Life is a circle. Life is a circle. Life is a circle.
https://battlepenguin.com/tech/microservices-and-biological-...
I agree with the author on a lot of points. You shouldn't start out with bricks. You should build the house first, and one you get that figured out, only then should you turn the individual rooms into modular building blocks.
I got one: Macroservices!
just write bad infrastructure as bad code. Get on with the times