Modules, not microservices
blogs.newardassociates.com
blogs.newardassociates.com
There's two technical problems that microservices purport to solve: modularization (separation of concerns, hiding implementation, document interface and all that good stuff) and scalability (being able to increase the amount of compute, memory and IO to the specific modules that need it).
The first problem, modules, can be solved at the language level. Modules can do that job, and that's the point of this blog post.
The second problem, scalability, is harder to solve at the language level in most languages outside those designed to be run in a distributed environment. But most people need it a lot less than they think. Normally the database is your bottleneck and if you keep your application server stateless, you can just run lots of them; the database can eventually be a bottleneck, but you can scale up databases a lot.
The real reason that microservices may make sense is because they keep people honest around module boundaries. They make it much harder to retain access to persistent in-memory state, harder to navigate object graphs to take dependencies on things they shouldn't, harder to create PRs with complex changes on either side of a module boundary without a conversation about designing for change and future proofing. Code ownership by teams is something you need as an organization scales, if only to reduce the amount of context switching that developers need to do if treated as fully fungible; owning a service is more defensible than owning a module, since the team will own release schedules and quality gating.
I'm not so positive on every microservice maintaining its own copy of state, potentially with its own separate data store. I think that usually adds more ongoing complexity in synchronization than it saves by isolating schemas. A better rule is for one service to own writes for a table, and other services can only read that table, and maybe even then not all columns or all non-owned tables. Problems with state synchronization are one of the most common failure modes in distributed applications, where queues get backed up, retries of "bad" events cause blockages and so on.
Running code in a single process is MUCH lower overhead because you don't need to transit a network layer and you're generally just passing pointers to data around rather than serializing/deserializing it.
There are definitely some cases where using microservices does make things more CPU/memory efficient, but it's much rarer than people think. An example where you'd actually get efficiency would be something like a geofence service (imagine Uber, Doordash, etc.) where the geofence definitions are probably large and have to be stored in memory. Depending on how often geofence queries happen, it might be more efficient to have a small number of geofence service instances with the geofence definitions loaded in memory rather than having this logic as a module that many workers need to have loaded. But again, cases like this are much less common than the cases where services just lead to massive bloat.
I was working at Uber when they started transitioning from monolith to microservices, and pretty much universally splitting logic into microservices required provisioning many more servers and were disastrous for end-to-end latency times.
All of the catalog data was moved to a service which only served catalog data so its cache was optimized for catalog data and the load balancers in front of it could optimize across that cache with consistent hashing.
This was different from the front end web tier which used consistent hashing to pin customers to individual webservers.
For stuff like order history or customer data, those services sat in front of their respective databases and provided consistent hashing along with availability (providing a consistent write-through cache in front of the SQL databases used at the time).
I wouldn't call those areas where it makes things more efficient rare, but actually common, and it comes from letting your data dictate your microservices rather than letting your organiation dictate your microservices (although if teams own data, then they should line up).
Servers can only get so big. If your monolith needs more resources than a single server can provide, then you can chop it up into microservices and each microservice can get its own beefy server. Then you can put a load balancer in front of a microservice and run it on N beefy servers.
But this only matters at Facebook scale. I think most devs would be shocked at how much a single beefy server running efficient code can do.
On top of that, the tendency to get complacent with unstructured data in a lot of systems is really creating a very complicated lock-in when systems are developed for unique services on each cloud provider... Bowls of spaghetti.
Hire Solutions Architects for dev projects, make your apps future proof... I warn you. Too much microservice customization leads to vendor lock in and expensive operational costs... This is why a lot of apps get sunsetted early.
It might still be completely viable to rewrite something so that it's only 10% as efficient but you can scale it horizontally easily.
Aren't you still writing out to external data stores? If that's the case then its really not a comparison of in process to a RPC, it just two RPC hops now, no?
What happens if you need to redesign the architecture to meet new needs? That's right; it's not going to happen, because people will fight it tooth and nail.
You also cannot produce any meaningful end-user value without involving several teams.
Microservices is just "backend, frontend, and database team" reincarnated.
My take: do Microservices all you want, but don't organize teams around the services!
This isn't easily fixable, but I'd like technologists to at least be able to perceive the extent to which the surrounding business culture has permeated into technical culture. I'm probably considered old AF by a lot of you (41) but I'm becoming increasingly interested in tools/methodologies that enable fewer devs to do a lot more, even if it means that there are sharp edges.
hence, if you want a certain architecture, you likely need to execute a "reverse conway law" org first. To get the org into the target config, the software will follow.
Any organization that designs a system (defined broadly) will produce a design whose structure is a copy of the organization's communication structure.
— Melvin E. Conway
https://en.wikipedia.org/wiki/Conway's_lawExactly. In order to reap the benefits of modular/Microservices architecture, you need teams to be organized around product/feature verticals.
IMO it’s not about the internals of a product/service architecture, but about encapsulating, isolating, and formalizing inter-organizational dependencies.
Also - talking of redesign should be a major lift for any product/service that has real paying clients. The risk of breaking is huge, and the fact that during that time you won't be delivering incremental value is also going to look bad.
How often does major reorganization happen? For most companies the answer is never.
The "what if" argument is what leads to all sorts of premature optimizations.
This is actually the whole point (see: Reverse Conway).
Conway's Law is basically "You don't have a choice." You will cement the design of your app around your organizational design. (At least, beyond a certain org size. If you have only 4 engineers you don't really have the sort of structure in question here at all.) je42 narrowly beats me to the point that if this is a problem you can try to match your organization to your problem, but that takes a fairly agile organization.
"What happens if you need to redesign the architecture to meet new needs? That's right; it's not going to happen, because people will fight it tooth and nail."
Unfortunately, in the real world, this is not so much "a disadvantage to the microservice approach" as simply "an engineering constraint you will have to work around and take account in your design".
Despite what you may think after I've said that, I'm not a microservice maximalist. Microservices are a valuable tool in such a world but far from the only tool. As with everything, the benefits and the costs must be accounted for. While I've not successfully rearchitected an entire organization from my position as engineer, I have had some modest, but very real success in moving little bits and pieces around to the correct team so that we don't have to build stupid microservices just to deal with internal organizational issues. I don't mean this to be interpreted as a defeatist "you're doomed, get ready for microservices everywhere and just deal with it"; there are other options. Or at least, there are other options in relatively healthy organizations.
But you will have code that matches the structure of your organization. You might as well harness that as much as you can for the benefits it can provide because you're stuck with it whether you like it or not. By that I mean, as long as you are going to have teams structured around your services whether you like it or not, go in with eyes open about this fact, and make plans around it to minimize the inevitable costs while maximizing the benefits. Belief that there is another option will inevitably lead to suboptimal outcomes.
You can't engineer at an organizational level while thinking that you have an option of breaking the responsibility and authority over a codebase apart and somehow handing them out to separate teams. That never works long term. A lot of institutional malfunctioning that people correctly complain about on HN is at its core people who made this very mistake and the responsibility & authority for something are mismatched. Far from all of it, there are other major pathologies in organizations, of course. But making sure responsibility & authority are more-or-less in sync is one of the major checklist items in doing organization-level engineering.
That will always happen, and I've seen it happen in a monolith too (expressed as module ownership/ownership over part of the codebase).
It's inevitable in almost all organizations.
Look, once upon a time managers designed the system that workers implemented. The factory assembly line, or the policy manual.
But now the workers are CPUs and servers. The managers designing the system the workers follow are coders. I say coders are the new managers.
Now this leaves two problems. The first is that there are a lot of "managers" who are now doing the wrong job. (and explains perfectly why Steve Jobs gets involved in emails about hiring coders.) But that's a different post.
For me the second problem is microservices represent the atomic parts of a business. They should not be seen as a technical solution to anything, and because the first problem (managers arent managing workers anymore) there is no need to have the microservices built along Conways Law.
And so if you want to build a microservice it is not enough to consider the technical needs, it is not enough to automate what was there, you need to design an atomic building block of any / your company. They become the way you talk about the company, the way you think about the processes in a company.
An mostly the company becomes programmable. It is also highly likely the company becomes built mostly of interchangable parts.
This is a misunderstanding of Conway's Law. Your code _will_ reflect the structure of your organization. If you use microservices so will their architecture. If you use modules, so will the modules.
If you want a specific architecture, have your organization reflect the modules/microservices defined in that architecture.
The point ebing is that if workers are CPUs and coders are managers, then why worry about how the managers of the coders are arranged. Get rid of that management layer. Conway is not implying the financiers of the organisation affect the architecture.
This basically means that the microservices a company is built out of should more readily align to the realities of the business model. And instead of shuffling organisations around it would behove executives to look at the microservices.
One can more easily measure orders transferred, etc, if the boundaries are clearer.
Plus conway is just a ... neat idea, not a bound fate.
There is a caveat with the architecture reflecting the organisation of the software teams, but that usually follows the other to-be-automated structure as well.
There's a third technical problem that microservices solve, and it's my favorite: isolation. With monoliths, you provide all of your secrets to the whole monolith, and a vulnerability in one module can access any secret available to any other module. Similarly, a bug that takes down one module takes down the whole process (and probably the whole app when you consider cascading failures). In most mainstream languages, every module (even the most unimportant, leaf node on the dependency tree) needs to be scoured for potential security or reliability issues because any module can bring the whole thing down.
This isn't solved at the language level in most mainstream languages. The Erlang family of languages generally address the reliability issue, but most languages punt on it altogether.
> The real reason that microservices may make sense is because they keep people honest around module boundaries.
Agreed. Microservices, like static type systems, are "rails". Most organizations have people who will take shortcuts in the name of expedience, and systems with rails disincentivize these shortcuts (importantly, they don't preclude them).
Imagine a world where every pip/nuget/cargo package was a k8s service called out-of-process and needed to be independently maintained. We would have potentially hundreds of these things running, independently secured via mTLS, observability and metrics for each, and all calls run out of process. This is the abysmally slow hellscape that some are naively advocating for without realizing it.
It is relatively obvious what should be a library (json, numpy, NLog, etc), and what should be a service (Postgres, Kafka, Redis, NATS, etc.) when dealing with 3rd party components. It is also obvious that team scaling is not determined by whether code exists in a library or service from this example since all 3rd party code is maintained by others.
However, once we are all in the same organization, we lose the ability to correctly make this determination. We create a service when a library would do in order to more strictly enforce ownership. I think an alternate solution to this problem is needed.
Services can be continuously deployed by small owning teams.
So does modularity.
"The benefits expected of modular programming are: (1) managerial_development time should be shortened because separate groups would work on each module with little need for communication..."
On the Criteria To Be Used in Decomposing Systems into Modules, D.L. Parnas 1972.
He's right as far as Conway's Law goes, though.
Developers do a much better job with microservices. I think it's easy for them to respect the seriousness of designing and changing the API of a microservice. In contrast, they often don't respect or even realize the seriousness of making a change that affects the modular design of code.
Language-level support for defining and enforcing module interfaces might help developers invest the same level of care for module boundaries in a modular monolith as they do in a microservices architecture, but I've yet to work in a language that achieves this.
If you have a particular piece of the system that needs to be scaled, you can take that module out when it becomes necessary. You can alter your system such that you deploy the entire codebase but certain APIs are routed to certain boxes, and you can use batch processing patterns with a set of code outside the codebase.
You can have an admin application and a user application, and both of them are monoliths. They may or may not communication using the database or events or APIs.
However, you don't make this on the single bounded context guideline.
And I do agree that most people need less of those than they think.
I gave a talk [1] about scalability of Java systems on Kubernetes, and one of my recommendations is that Java-based systems - or any system on a runtime similar to the JVM, like CLR or even Go - should be scaled diagonally. Efficiency is the word.
While horizontal scaling can easily address most performance issues and load demand, in most cases it is the least efficient way for Java (and again, I risk saying .NET and Go), as these systems struggle with CPU throttling and garbage collectors.
In short, and exemplifying: one container with 2 CPUs and 2GB of RAM will allow the JVM to perform better, in general, than 2 containers with 1 CPU/1GB RAM each. That said, customers shouldn't be scaling horizontally to any amount more than what is adequately reasonable for resiliency, or unless the bottleneck is somewhere like disk access. For performance on the other hand, customers should be scaling vertically.
And Kubernetes VPA is already available. People just need to use it properly and smartly. If a Kubernetes admin believes a particular pod should double in number of replicas, the admin should consider: "would this microservice benefit even more from 1.5x more CPU/Memory than 2x more replicas?" and I bet to say that, in general, yes.
My default recommendation has always been to make instances "as big as possible, but no bigger". You may need 3 instances for redundancy, fault tolerance and graceful updates but past that you should probably scale up to near the size of your machine. There are obviously lots of complications and exceptions (for example maybe using most 3 instances uses 90% of your machines so you can't bin pack other processes there so it is better to use 4 instances at 70% as the machines will be used more efficiently) but bigger by default is generally a good option.
Even if Microservices are better for scale, most companies will never experience the level of scale of 2013 Twitter.
Are Microservices beneficial at much smaller levels of scale? Ex: 1M MAU
I fully agree with your argument. Then again, as mentioned elsewhere in this discussion, microservices are often not introduced to solve a scalability problem but an organizational one and there are many organizations that have more engineers than Twitter (had).
Personally, I still don't buy that argument because by solving one organizational problem one risks creating a different one, as this blog post[0] illustrates:
> […] Uber has grown to around 2,200 critical microservices
Unsurprisingly, that same post notes that
> […] in recent years people have begun to decry microservices for their tendency to greatly increase complexity, sometimes making even trivial features difficult to build. […] we experienced these tradeoffs first hand
I'm getting very strong Wingman/Galactus[1] vibes here.
[0]: https://web.archive.org/web/20221105153616/https://www.uber....
If a team wants to try out a different language, or hosting model, or even just framework/tooling, those things can be really hard to do within the same codebase; much easier when your only contract is handling JSON requests. And if your whole organization is totally locked into a single stack, it's hard to keep evolving on some axes
(I'm generally against microservices, but this is one of the more compelling arguments I've heard for them, though it still wouldn't mean you need to eagerly break things up without a specific reason)
Each team gets to distribute libraries over repos (COM, JAR, DLL, whatever), no way around that unless they feel like hacking binaries.
Creating an object-oriented API for a shared library which encapsulates a module to the degree a service boundary would is not trivial and it's very rarely done, never mind done well. Most OO libraries expose an entire universe of objects and methods to support deep integration scenarios. Maintaining that richness of API over time in the face of many consumers is not easy, and it (versioning) is a dying art in the eternal present of online services. The 90s aren't coming back any time soon.
If you own a library which is used internally, and you add a feature which needs a lot more memory or compute, how do you communicate the need to increase resource allocation to the teams who use the library? How do you even ensure that they upgrade? How do you gather metrics on how your library is used in practice? How do you discover and collect metrics around failure modes at the module boundary level? How do you gather logs emitted by your library, in all the places it's used? What if you add dependencies on other services, which need configuring (network addresses, credentials, whatever) - do you offload the configuration effort on to each of the library users, who need to do it separately, and end up with configuration drift over time?
I don't think binary drops work well unless the logic they encapsulate is architecturally self-contained and predictable; no network access, no database access, no unexpected changes in CPU or memory requirements from version to version.
There's plenty of code like this, but it's not usually the level of module that we consider putting inside a service.
For example, an Excel spreadsheet parser might be a library. But the module which takes Excel files uploaded by the user and streams a subset of the contents into the database is probably better off as a service than a library, so that it can be isolated (security risks), can crash safely without taking everything down, can retry, can create nice logs about hard to parse files, can have resource metrics measured and growth estimated over time, and so on.
The important part of microservices isn't just API boundaries, it's lifecycle management. This CAN be done with a DLL or JAR, but it's MUCH harder today.
At my last job, there were quite a few times where being able to scale some small "microservice instance" up from 2 -> 4 instances or 4 -> 8 or 8 -> 12 was a lot easier/quicker than investigating the actual issue. It'd stop production outages/hiccups. It was basically encouraged.
Not sure how that can be done with a "it's all modules in a giant monolith".
This happens a lot. Organisational problems conflated for technical ones. Technical solutions to organisational politics. Etc. It's often easier to admit technical challenges than organisational ones. Also, different people get to declare technical and organisational challenges, and different people that get to design the solutions.
There's also a dynamic where concepts are created by the technical 1% doing vanguard or challenging tasks. The guys responsible for scaling youtube or whatnot. Their methods and ideas become famous and are then applied to less technically demanding tasks, then in less technical organisations entirely.
I think if we can be honest at the actual problem at hand, 80/20 fixes will emerge. IE, the "real" value is not the architecture per se, but the way it lets you divide the responsibilities in the organisation.
https://ardalis.com/conways-law-ddd-and-microservices/#:~:te....
I would like someone to spell this out. It seems to me people are claiming that if a single binary serves some CPU-bound requests and some memory-bound requests, and you give it more memory, then the memory gets "wasted" on the CPU-bound part. Or if you give it more CPU, the CPU gets wasted on the memory-bound part. But this kind of assignment of resources to code paths seems to be a consequence of microservices. In a single computer, single binary situation resources should not get used up unless the workload actually wants to use them. A compute-heavy thread doesn't cost heap. A big heap doesn't slow down a compute-heavy thread. What am I missing?
My IDE should be able to easily traverse the call graph. My development environment should be simple to setup.
I’ve worked on microservices that required an insane amount of boilerplate to do simple things. Like 7 layers of controllers, clients, services, data services, etc, just to fetch a simple piece of data. And the developer experience of running dozens of services in a Kubernetes cluster running on my dev machine was awful.
Does anything like this exist?
I only dabbled many years ago but Erlang/OTP comes to mind.
And tRPC for TypeScript calls in browser and server.
> https://github.com/NathanRSmith/lib-courier-js
Interestingly, we're pursuing a monorepo & multi-monolith setup for the next version of our platform. So lib-courier is no longer necessary to stitch it all together. It was fun while it lasted though. Once you understood the routing algo & code patterns, lots of stuff "just worked".
It turns out that "transparent RPC" is basically a contradiction in terms. As soon as you start doing things across process boundaries, and even more so across network boundaries, it requires a very different approach for API design - something that's very cheap locally, like passing objects by reference, becomes expensive and full of footguns.
If you reduce the feature set to the point where it can be transparently mapped to either local or RPC - which is, basically, function calls processing and returning data organized into arrays & trees (but not graphs) - there's still the issue that RPC has so many more failure points that you have to handle that would never light up in local.
This is all still doable; I have my doubts about practical the end result would be, though.
Modularization allows development to scale.
Microservices allow operations to scale.
This breaks down when the database is essentially a generic graph. The worst solution I've seen to this is to have another service responsible for generic write operations and any service that wants to write data goes through that service -- you're essentially re-introducing the problem you're purporting to solve at a new layer with an added hop and most likely high network usage. The best solution I've seen is to obviously have the monolith. The enterprise solution I've seen, while not good by any means but nearly essential for promoting team breakdown and ownership, is to just let disparate services write as they see fit, supported by abstraction libraries and SOPs to help reinforce care from engineers.
I’ll go one step further and say that you should treat your data stores as services in and of themselves, with their own well-considered interfaces, and perhaps something like PostgREST as an adapter to the frontend if you don’t really need sophisticated service layers. The read/write pattern you recommend is a great fit for this and can be scaled horizontally with read replicas.
Been there: how do you handle schema changes?
One of the advantages that having a separate schema per service provides is that services can communicate only via APIs, which decouples them allowing you to deploy them independently, which is at the heart of microservices (and continuous delivery).
The way I see it today: everyone complains about microservices, 12 factor apps, kubernetes, docker, etc., and I agree they are overengineering for small tools, services, etc., but if done right, they offer an agility that monoliths simply can't provide. And today it's really all about moving fast(er) than yesterday.
Our data is mostly append-only, or if it's being changed, there is a theoretical final "correct" version of it that we should converge to. So to "get" data, you subscribe to messages about some state of things, and then each service is responsible for managing its own copy in its own db. This worked well enough until it didn't, and we had to start doing true-ups from time to time to keep things in sync, which was annoying, but not particularly problematic, as we design to assume everything is async and convergent.
The optimization (or compromise) we decided on, was that all of our services use the same db cluster, and that if the db cluster goes down, it means everything is down. Therefore, if we can assume the db is always up, even if a service is down, we consider it an acceptable constraint to provide a readonly view into other services db. Any writes are still sent async via MQ. This eliminates our syncing drifting problem, while still allowing for performant joins, which http apis are bad at and our system uses a lot of.
So then back to your original question, the way that this contract can break is via schema changes. So for us, since we use postgres, we created database views that we expose for reading. And postgres view updates are constrained that they must always be backwards compatible from a schema perspective. So then now our migration path is:
- service A has some table of data that you like to share
- write a migration to expose a view for service A
- write an update for service B to depend upon that view
- service B now needs some more data in that view
- write a db migration for service A that adds that missing data, but keeping the view fully backwards compatible
And any benefit of a microservice owning it's own rDB is still, that schema changes aren't easily reversible. Specially when new, non predefined, data has been flowing in.
Stateless microservices are great, in the sense that you don't have to build multiple versions of APIs... but stateful microservices are just a PITA.
Microservices allow your permissions to be clear and precise. Your database passwords are only loaded into the process that uses them. You can reason about "If there's an RCE in this service, here's what it can do". Trying to tie that to monoliths is hard and ugly.
And for that matter it can (but not necessarily) limit how much damage a bad code release can do.
Of course you don't get those benefits for free just by using microservices, but impleminting those kind if boundaries in a monolith is a lot harder.
Microservices aren't inherently more secure.
In the outside world, for an application that may truly need to scale, I'd go MySQL -> Vitess before I'd choose separate data stores for each service with state. But I'd also question if the need to scale that much really exists; you can go a long way even with data heavy applications with a combination of sharding and multitenancy.
It is about striking a balance. No reason should something be overly compounded or overly broken up.
It's invented by a software outsourcing firm to milk billable hours from contracts.
This experience has strongly impacted my view of microservices and for all personal projects I will develop in the future I will stick with a monolith until much later instead of starting with microservices.
For this reason, I've been trying to push for building a monolithic app first, then splitting into components, and introducing libs for common functionality. Only when this is all done, you think about the communication patterns and discuss how to scale the app.
Most microservice shops I've been in have instead done the naïve thing; just come up with random functionally separate things and put them in different micro services; "voting service", "login service", "user service" etc. This can come with a very very high price. Not only in terms of network traffic, but also in debuggability, having a high amount of code duplication, and getting locked into the existing architecture, cementing the design and functionality.
The main thing is that regardless of scaling, the app should always be able to run/debug/test locally in a monolithic thing.
Once people scale they seem to abandon the need to debug locally at their peril.
Scaling should just be a process of identifying hot function calls and when a flag is set, to execute a call as a network rpc instead.
[1] https://scholar.harvard.edu/waldo/publications/note-distribu...
And it is reassuring to hear that you seem to have success in avoiding these issues with a monolithic architecture, as I thought I was oldschool for starting to prefer monoliths again.
Not trying to be harsh here, but not expecting an increase of network call in a system where each component is tied together with... network calls sounds a bit naive.
> We are now attempting to solve this with caches and doing batch requests
So you have built a complex and bottle-necked application for the sake of scalability, then having to add caching and batching just to make it perform? That sounds like working backwards.
Obviously, I have no clue on the scale of the project you are working on, but it sure sounds like you could have built it as a monolith in half of the time with orders of magnitude more performance.
Scalability is a feature, you can always add it in the future.
Yes, if you can't scale fast enough as you need to, it can hurt your business. Not being able to keep up with demand is a (luxury) problem that every business faces, not just in tech. They would often be called 'growing pains' in a business, and though they are bad, they rarely contribute to the failure of a company.
Starting a startup/service/platform with microservices before you even understand the bottlenecks/market fit/customers is usually not a good idea. You can come a very, very long way with a monolith before you hit performance and scalability limits. And once you do, you can always start breaking things up into smaller services for scalabity. Obviously you need to make sure you are scaling on time to keep up with demand.
'Nail it, then scale it', and 'premature optimization is the mother of all f-ups' are popular sayings for a reason.
I wouldn't call that entirely unexpected. :-) It's a rather well-known issue:
Microservices
grug wonder why big brain take hardest problem, factoring system correctly, and introduce network call too
seem very confusing to grug
(from https://grugbrain.dev)How the hell... like... who decided to do Microservices in the first place if they didn't know this? This is such a rookie mistake. It's like somebody right out of high school just read on a blog the new way is "microservices" and then went ahead with it.
And both are way more common than they should be.
Microservices can have both design utility and simultaneously been a major fad.
Lurking reddit and HN you can watch development fashions come and go. It's really really profoundly hard to hold a sober conversation on splitting merits from limitations in the middle of the boom.
tl;dr: gartner_hype_cycle.png
Imagine if you are responsible, at runtime, for linking object files (.o) where each object is a just the compilation of a function.
Now why would anyone think this is a good idea (as a general solution)? Because in software organizations, the “linker’s” job is supposed to be done by the (“unnecessary weight”) software architect.
Microservices primarily serve as a patch for teams incapable of designing modular schemas. Because designing schemas is not entry level work and as we “know” s/e are “not as effective” after they pass 30 years of age. :)
> monolith
Unless monolith now means not-microservice, then be aware that there are a range of possible architectures between a monolith and microservices.
I first wrote a program which ran on a web server about 25 years ago. In that time, computers have experienced about ten doublings of Moore's law, i.e. are now over a thousand times faster. Computers are very fast if you let them be.
> I was told ~1200 RPCs independently by several engineers at Twitter, which matches # of microservices. The ex-employee is wrong.
> Same app in US takes ~2 secs to refresh (too long), but ~20 secs in India, due to bad batching/verbose comms.
RPCs are on the server side. Why would they app take longer to refresh in India than in the US?
Some more explanation:https://twitter.com/mjg59/status/1592380440346001408
Let each microservice owner figure out how to achieve their latency+reliability SLA's in every location - whether replicating datastores, caching, being stateless, or proxying requests to a master location.
> especially since some of these services are not even in the same data center.
I think you need to answer why? If you can't put all of the services in one data center, then by definition you can't write a monolith either. If the monolith would happily run in one datacenter, then you should have all instances of your microservices in that one datacenter.
It surprised me that you would conclude that this is a problem with microservices. It's like if a particularly architect always punches you in the groin in every meeting, and you've concluded that architects are bad people, rather than this one architect is a bad person.
At my last job we created a whole fleet of microservices instead of a single modular project/repo. Some of them required non-trivial dependencies. Some executed pretty long-running tasks or jobs for which network latency is insignificant and will remain so by design. Some were few pages long, some consisted of similar-purpose modules with shared parts factored out. But there was no or little processes like “ah, I’ll just ask M8 and it will ask M11 and it will check auth and refer to a database. I.e. no calls as trivial as foo(bar(baz())) but done all over the infra.
(Because this is at the same time one of the defining elements of this architecture... and the first one to be opted out when you actually start using it "for real").
Lots of ways to solve a problem, though.
The primary reason for this is that PDFs can contain executable code and the common tools used to process them are full of unpatched CVEs.
Either way, one of my biggest pet peeves is the near-ubiquitous use of HTTP & JSON in microservice architectures. There's always going to be overhead in networked servies, but this is a place where binary protocols (especially ones like Cap'n Proto) really shine.
Why is the other service in another data center? Does it need to be in another data center? If it does, how will a monolith help?
* they force alignment on one language or at least runtime
* they force alignment of dependencies and their versions (yes, you can have different versions e.g. via Java classloaders, but that's getting tricky quickly, you can't share them across module boundaries, etc.)
* they can require lots of RAM if you have many modules with many classes (semi-related fun fact: I remember a situation where we hit the maximum number of class files a JAR could have you loaded into WebLogic)
* they can be slow to start (again, classloading takes time)
* they may be limiting in terms of technology choice (you probably don't want ot have connections to an RDBMS and Neo4j and MongoDB in one process)
* they don't provide resource isolation between components: a busy loop in one module eating up lots of CPU? Bad luck for other modules.
* they take long to rebuild an redeploy, unless you apply a large degree of discipline and engineering excellence to only rebuild changed modules while making sure no API contracts are broken
* they can be hard to test (how does DB set-up of that other team's component work again?)
I am not saying that most of these issues cannot be overcome; to the contrary, I would love to see monoliths being built in a way where these problems don't exist. I've worked on massive monoliths which were extremely well modularized. Those practical issues above were what was killing productivity and developer joy in these contexts.
Let's not pretend large monoliths don't pose specific challenges and folks moved to microservices for the last 15 years without good reason.
On the RAM front, I am now approaching terabyte levels of services for what would be gigabyte levels of monolith. The reason is that I have to deal with mostly duplicate RAM - the same 200+ MB of framework crud replicated in every process. In fact a lot of microservice advocates insist "RAM is cheap!" until reality hits, especially forgetting the cost is replicated in every development/testing environment.
As for slow startup, a server reboot can be quite excruciating when all these processes are competing to grind & slog through their own copy of that 200+ MB and get situated. In my case, each new & improved microservice alone boots slower than the original legacy monolith, which is just plain dumb, but it's the tech stack I'm stuck with.
You are writing microservices and then running them on the same server??
How is this possibly a down-side from an org perspective? You don't want to fracture knowledge and make hiring/training more difficult even if there are some technical optimizations possible otherwise.
E.g. if you end up having a requirement to add some machine learning to your application, you might be better off using Tensorflow/PyTorch via Python than trying to deal with it in whatever language the core of the app is written in.
these are not maxims of development, there can be reasons that make these consequences worth it. Furthermore you can still use just a single language with microservices*, nothing is stopping you from doing that if those consequences are far too steep to risk.
*:you can also use several languages with modules by using FFI and ABIs, probably.
This is the Pendulum Swing all over again. If one language and runtime is limiting, forty is not liberating. If forty languages are anarchy, switching to one is not the answer. This is in my opinion a Rule of Three scenario. At any moment there should be one language or framework that is encouraged for all new work. Existing systems should be migrating onto it. And because someone will always drag their feet, and you can’t limit progress to the slowest team, there is also a point in the migration where one or two teams are experimenting with ideas for the next migration. But once that starts to crystallize any teams that are still legacy are in mortal danger of losing their mandate to another team.
A sane thing to do.
>they force alignment of dependencies and their versions
A sane thing to do. Better yet to do it in a global fashion, along with integration tests.
>they can require lots of RAM if you have many modules with many classes
You can't make the same set of features build in a distributed manner comsume _less_ RAM than the monolith counterpart. Given you're now running dozens of copies of the same java vm + common dependencies.
>they can be slow to start
Correct.
>they may be limiting in terms of technology choice
Correct.
>they don't provide resource isolation between components
Correct.
>they take long to rebuild an redeploy, unless you apply a large degree of discipline and engineering excellence to only rebuild changed modules while making sure no API contracts are broken
I think the keyword is the WebLogic Server mentioned before. People don't realise that monolith architecture does't mean legacy technology. Monolith web services can and should be build in Spring Boot, for example. Also, most of the time, comparisons are unfair. In all projects i've worked im yet to see a MS instalation paired feature-wise with his old monolith cousin. Legacy projects tends to be massive, as they're made to solve real world problems while evolving during time. MS projects are run for a year or two and people start to compare around apples to oranges.
>they can be hard to test
If other team's component break integration, the whole building stops. I think Fail-Fast is a good thing. Any necessary setup must be documented in whatever architectural style. It can be worse in a MS scenario, where you are tasked to fix a dusty, forgotten service with an empty README.
If anything, monolithic architecture brings lots of awareness. It's easier to get how things are wired and how they interact together.
imaging your application contains of two pieces - somewhat simple crud, that requires to respond _fast_ and huge batch processing infrastructure, that needs to work as efficient as possible, but doesn't care about single element processing time. And suddenly 'the sane thing to do' is not the best thing anymore. You need different technologies, different runtime settings and sometimes different runtimes. But most importantly they don't need constraints imposed by unrelated (other) part of the system.
>A sane thing to do. Better yet to do it in a global fashion, along with integration tests.
But brutally difficult at scale. If you have hundreds of dependencies, a normal case, what do you do when one part of the monolith needs to update a dependency, but that requires you update it for all consumers of the dependency's API, and another consumer is not compatible with the new version?
On a large project, dependency updates happen daily. Trying to do every dependency update is a non-starter. No one has that bandwidth. The larger your module is, the more dependencies you have to update, and the more different ways they are used, so you are more likely to get update conflicts.
This doesn't say you need microservices, but the larger your module is, the further into dependency hell you will likely end up.
This is incredibly subjective, and contingent on the size and type of engineering org you work in. For a small or firmly mid-sized shop? yea I can 100% see that being a sane thing to do. Honestly a small shop probably shouldn't be doing microservices as a standard pattern outside of specific cases anyway though
As soon as you have highly specialized teams/orgs to solve specific problems, this is no longer sane.
Isn't that exactly what's required when you're deploying microservices independently of each other? (With the difference of the interface not being an ABI but network calls/RPC/REST.)
You can then have a pretty flexible trade off between the convenience of having email be a rooted library against the trade off of keeping it a lead service (the implication being that leaf services can talk to one another over the network via service stubs, rest, what have you).
This is SOA (Service Oriented Architecture), which should be considered in the midst of the microservice / monolith conversation.
or maybe run redundant monolith fail over servers. should work the same as micro services.
You can have modules implemented in different languages and runtimes. For example you can have calls between Python, JVM, Rust, C/C++, Cuda etc. It might not be a good idea in most cases but you can do it.
Lots of desktop apps do this.
You can absolutely call js running in V8 vm from scala running in jvm. No networking needed, hell not even IPC is needed.
And when you deploy this you don't have to deploy all modules' http servers (for external requests into the system) and queue consumers in the same container, only a single module's. So no busy loops affect other modules, unless as a result of direct api call from module to module. If anything it encourages looser coupling as you are incenticised to use indirect communication through the queue over direct api calls.
Uh... what's the trick? I don't see how you can have V8 and the JVM communicate without something that's inter-process.
It's okay to not organize your code. It's okay to have files with 10,000 lines. It's okay not to put "business logic" in a special place. It's okay to make merge conflicts.
The overhead devs spend worrying about code organization may vastly exceed the amount of time floundering with messy programs.
Microservices aren't free, and neither are modules.
[1] Jonathan Blow rant: https://www.youtube.com/watch?v=5Nc68IdNKdg&t=364s
[2] Jon Carmack rant: http://number-none.com/blow/john_carmack_on_inlined_code.htm...
Games are highly stateful, with a game loop that iterates over the same global in-memory data structure as fast as it can. You have (especially in Carmack era games) a single thread performing all the game state updates in sequence. So shared global state makes a ton of sense and simplifies things.
Most web applications are highly stateless with request-oriented operations that access random pieces of permanently-stored data. You have multiple (usually distributed) threads updating data simultaneously, so shared global state complicates things.
That game devs gravitate towards different patterns for their code than web service devs should not be a surprise.
It’s easy to miss the point of what the OP is saying here and get distracted by the fact this file is ridiculously huge. This file used to be a paltry 1k file, a 10k file, a 20k SLOC file… but it is where it is today because of the OP suggested approach.
If you're a single dev making a game, by all means, do what you want.
If you work with me in a team, I expect a certain level of quality in the code you write that will get shipped as a part of the project I'm responsible for.
It should be structured, tested, malleable, navigable and understandable.
Also, the question of inlined code mostly applies to programming in the small, while modules are about programming in the large, so I don't think there's much relationship between the two.
The problem is when it's OK and for how long. If you have a team of people working with a codebase with all those "okays", then they have to be really good developers and know the code inside out. They have to agree when to refactor a business login out instead of adding a hacky "if" condition nested in another hacky "if" condition that depends on another argument and/or state.
I guess what I'm trying to say that if those "okays" are in place, then there's a whole bunch of unwritten rules that come in place.
But I agree that microservices certainly aren't free (I'd say they are crazy expensive) and modules aren't free either. But all those "okays" can end up costing you your codebase also.
It's not. This is the thing where you start thinking "YAGNI", yadayada, but you inevitably end up needing it. Layering with at least a service and a database/repositories is a no brainer for any non-toy app considering the benefits it brings.
> It's okay to have files with 10,000 lines
10.000 lines is a LOT. I consider files to become hard to understand at 1.000 lines. I just wc'd the code base I work on, we have like 5 files with more than 1.000 lines and I know all of them (I cringed reading the names), because they're the ones we have the most problems with.
It gets worse when the people who made the mess quit or move on, leaving the new hires to deal with it. I've seen this pattern enough times to wonder if it gets repeated with most companies or projects.
I do agree that microservices and/or modules aren't magical solutions that should be universally applied. But they can be useful tools, depending on the situation, to organize or re-organize a system for particular purposes.
Anecdotally, I've noticed that smart people who aren't good programmers tend to be able to write code quickly that can scale to a particular point, like 10k-100k lines of code. Past that point, productivity falls rapidly. I do believe that part of being a skilled developer is being able to both design a system that scales to millions of lines of code across an organization, and to operate well on one designed by someone else.
Ever since my time as a mathematician (I worked at a university) and using LaTeX extensively, I never understood the "divide your documents/code into many small files" mantra. With tools like grep (and its editor equivalents), jumping to definition, ripgrep et al., I have little problem working with files spanning thousands of lines. And yet I keep hearing that I should divide my big files into many smaller ones.
Why, really?
If you split up MegaFunction(){} to Func1(){} Func2(){}, etc, but it never makes sense to call Func2 except after Func1, then you haven't actually created two functions, you've just created one function in two places.
Refactoring should be about logical separation not about just dicing a steak because it's prettier that way.
>It's okay to not organize your code. It's okay to have files with 10,000 lines. It's okay not to put "business logic" in a special place. It's okay to make merge conflicts.
I absolutely disagree. It's "okay" if you're struggling to survive as a business and worrying about the future 5+ years out is pointless since you don't even know if you'll make it the next 6 months. This mentality of there being no need for discipline or craftsmanship leads to an unmanageable codebase that nobody knows how to maintain, everybody is afraid to touch, and which can never be upgraded.
You don't see the overhead of throwing discipline out the window because it's all being accrued as technical debt that you only encounter years down the road.
https://github.com/tinspin/rupy/wiki/Process
All input I have is you want your code to run on many machines, in fact you want it to run the same on all machines you need to deliver and preferably more. Vertically and horizontally at the same time, so your services only call localhost but in many separate places.
This in turn mandates a distributed database. And later you discover it has to be capable of async-to-async = no blocking ever anywhere in the whole solution.
The way I do this is I hot-deploy my applications async. to all servers in the cluster, this is what a cluster node looks like in practice (the name next to Host: is the node): http://host.rupy.se if you click "api & metrics" you'll see the services.
With this not only do you get scalability, but also redundancy and development is maintained at live coding levels.
This is the async. JSON over HTTP distributed database: http://root.rupy.se (2000 lines hot-deployable and I can replace all databases I needed up until now)
I easily find my way in messy codes with grep. With modules, I need to know where to search to begin with, and in which version.
Fortunately, I have never had the occasion to deal with microservices.
Couldn't disagree more. As usual, it's a tradeoff. You could spend an infinite amount of time refactoring already fine programs. But complex code decreases developers productivity by orders of magnitude. Maybe it's not always worth refactoring legacy code, but you're always much better off if your code is modular with good separation of concerns.
I find there are practical negative consequences to having a 10,000 line file (well, okay; we don't have those at work, we have one 30k line file). It slows both the IDE and git blame/history when doing stuff in that file (w.r.t. I'll look at the history of the code when making some decisions). These might not be a factor depending on your circumstances (e.g. a young company where git blame is less likely to be used or something not IDE driven). But they can actually hurt apart from pure code concerns.
You have to understand everything that calls that function before you change it.
This is not always obvious. It takes time to figure out where it's called. IDEs make this easier but not bullet proof. Getting it wrong can cause major, unexpected problems.
Unfortunately most engineers don't like nuance, they want one-size-fits-all solutions.
Microservices are the same modules. Though they force-add distributiveness, even where it can be avoided, which is fundamentally worse. And they make integration and many other things lot harder.
I.e. the idea that each microservice has direct access to its own, dedicated, maybe duplicated storage schema/instance (or if it needs to know, for example, the country name for ISO code "UK" it is supposed to ... invoke another microservice that will provide the answer for this).
I always worked in pretty boring stuff like managing reservations for cruise ships, or planning production for the next six weeks in an automotive plant.
The idea of having a federation of services/modules constantly cross-calling each other in order to just write "Your cruise departs from Genova (Italy) at 12:15 on May 24th, 2023" is not really a good fit for this kind of problems.
Maybe it is time to accept that not everyone has to work on the next version of Instagram and that Microservices are probably a good strategy... for a not really big subset of the problems we use computers for?
One strategy I've used is to designate one system as a source of truth (usually an ERP system) and periodically query its database directly to reload a cache. Every system works off their own periodically refreshed cache. Ideally, having all the apps query a read replica would prevent mistakes from taking down the source of truth.
I haven't done this, but I think using Postgres with a combination of read-only foreign tables and materialized views could neatly solve this problem without writing any extra code.
I don't know how far this would scale. I do know that coordination and procedures for using/changing the source of truth will fall apart long before technical limitations.
A micro service implementation of that would be CQRS, where the services to write updates, backend process (eg, notifying the crew to buy appropriate food), and query existing records are separated.
You might even have it further divided, eg the “prepare cruise” backend calls out to several APIs, eg one related to filing the sailing plans, one related to ensuring maintenance signs off, and one related to logistics.
The thing is that the framing of "the problems we use computers for" misses the entire domain of problems that microservices solve. They solve organisational problems, not computational ones.
And this is not because "my stuff is complicated and your stuff is a toy", either. It's more like "ERP or Banking Systems" were deployed decades ago, started as monoliths and nobody can really afford to rewrite these from scratch to leverage Microservices (or whatever the next fad will be). (I am also not sure it is a good idea in general for transactions that have to handle/persist lots of state, but this could be debated).
The problem, in fact, is that "new guys" think that Microservices will magically solve that problem, too, want to use these because they are cool and popular (this year), and waste (and make me waste) lots of time before admitting something that was clear from day 1: Microservices are not a good fit for these scenarios.
(I still think that "these scenarios" are prevalent in any company which existed before the 90s, but I might be wrong or biased on this).
Does anyone have a microservice "map" of Instagram? I feel that would be helpful here.
Or you can think of it as trunk/branch architecture. One main "trunk" service and other branch services augmenting it, which is a simpler thing to reason about.
Now imagine a small shop of 20 devs deciding to build something more complicated.
Binaries are deployed and scaled independently as thrift services.
Tons of rpc.
No, you don't use separate microservices for writing out that text message.
The idea is pretty simple instead of writing one big program, you write many smaller programs.
It's useful around the 3 to 4 separate developers mark.
It avoids you having to recompile for any minor change, allows you run the tests for just the part you changed and allows you to test the microservices in isolation.
If you a production issue, the exception will be in a log file that corresponds to a single microservice or part of your codebase.
Microservices are a hard form of encapsulation and gives a lot of benefits when the underlying language lacks that encapsulation. e.g. Python.
But in order to find out that Genua is the name of the port from where the cruise is departing, the appropriate time (converted to the timezone of the port, or the timezone of the client who will see this message, depending on what business rule you want to apply) and that Genua is in IT=Italy... how many microservices do I have to query, considering that port data, timezones, ISO country codes and dep.date/time of the cruise are presumably managed on at least four different "data stores"?
And the virtue of microservices is that they create hard boundaries no matter your skill and seniority level. Any unsupervised junior will probably dissolve the module boundaries. But they can't simply dissolve the hard boundary of having a service in another location.
The best place for such boundaries is at the team/division/org level, not team-internal or single dev-internal like microservices implies with its name. Embrace Conway's law in your architecture, and don't subdivide within a service.
"Hey our developers can't make modules correctly! Let's make the API boundaries using HTTP calls instead, then they'll suddenly know what to do!"
And that unsupervised junior? At one place I joined, that unsupervised junior just started passing data he "needed" via query string params in ginormous arrays from microservice to microservice.
And it wasn't a quick fix because instead of using the lovely type system that TELLS you where the stupid method has been used, you've got to go hunting for http calls scattered over multiple projects.
All you've done is make everything even more complicated, if you can't supervise your juniors, your code's going to go sideways whatever.
Microservices don't solve that at all and it's pure circular logic to claim otherwise. If your team can't make good classes, they can't make good APIs either. And worse still, suddenly everything's locked in because changing APIs is much harder than changing classes.
The rigor of maintaining resilient, backward-compatible, versioned internal APIs is too resource and time consuming to do well. All I see is hack and slash, and tons of technical debt.
It seems like in the last couple of years it started sinking in, that distributed systems are hard.
Look up a thing called "distributed monolith".
I've seen people doing that a few times already. They start changing due to some uninformed kind of "convenience", and then you look at the services API and it makes no sense at all.
Microservices are a solution to a problem. TDD is a solution to a problem, the same problem. Both are solutions that themselves create more, and worse, problems. Thanks to the hype driven nature of software development the blast radius of these 'solutions' and their associated problems expands far beyond the people afflicted by the original problem.
That problem? Not using statically typed languages.
TDD attempts to reconstruct a compiler, poorly. And Microservices tries to reconstruct encapsulation and code structure, poorly. Maybe if you don't use a language which gets hard to reason about beyond a few hundred lines you won't need to keep teams and codebases below a certain size. Maybe if you can't do all kinds of dynamic nonsense with no guardrails you don't have to worry so much about all that code being in the same place. The emperor has no clothes, he never had.
Edit: to reduce the flame bait nature of the above a bit. Not in all cases, I'm sure there are a very few places and scales where microservices make sense. If you have one of those and you used this pattern correctly, that's great.
And splitting out components as services is not always bad and can make a lot of sense. It's the "micro" part of "microservices" that marks out this dreadful hype/trend pattern I object to. It's clearly a horrible hack to paper over the way dynamically typed codebases become much harder to reason about and maintain at scale. Adding a bunch of much harder problems (distributed transactions, networks, retries, distributed state, etc) in order to preserve teams sanity instead of just using tooling that can enforce some sense of order.
TDD is a method of documenting your application in a way that happens to be self-verifying. You could use a Word document instead, but lose the ability for the machine to verify that the application does what the documentation claims that it should. Static typing does provide some level of documentation as well, but even if you have static typing available static typing isn't sufficient to convey the full intent that your documentation needs to convey to other developers.
A compiler can't check that your logic is correct. There may be a bit of overlap in the things verified by tests and a compiler but they don't solve the same problem.
How are static types a requirement for encapsulation? Dynamic languages are perfectly capable of providing encapsulation. Statically typed languages are also perfectly capable of having very poor encapsulation.
Most important thing for lots of startups and companies is Time to Market.
Traditional static typing based approachs are just a bad joke (3 times slower on average) in comparsion.
That's why we have the whole microservices and dynamic typing thing going on, because businesses that use it beat up businesses that don't. It's pretty simple really.
TDD was invented by a Java programmer. How does that fit into your world view?
Good LUCK getting that property with a monolithic or modular system. QE can never be certain (and let's be honest, they should be skeptical) that something modified in the same codebase as something else does not directly break another unrelated feature entirely. It makes their life very difficult when they can't safely draw lines.
Two different "modules" sharing even a database when they have disparate concerns is just waiting to break.
There's a lot of articles lately dumping on microservices and they're all antiquated. News flash: there is no universal pattern that wins all the time.
Sometimes a monolith is better than modules is better than microservices. If you can't tell which is better or you are convinced one is always better, the problem is with you, not the pattern.
Microservices net you a lot of advantages at the expense of way higher operational complexity. If you don't think that trade off is worth it (totally fair), don't use them.
Since we are talking about middleground, one I'd like to see one day is a deploy that puts all services in one "pod", so they all talk over Unix socket and remove the network boundary. This allows you to have one deploy config, specify each version of each service separately and therefore deploy whenever you want. It doesn't have the scalability part as much, but you could add the network boundary later.
1. Deployment. Being able to deploy code rapidly and independently is lost when everything ships as a monolith.
2. Isolation. My process is GC-spiraling. Which team's code change is responsible for it? Since the process is shared across teams, a perf bug from one team now impacts many teams.
3. Operational complexity. People working on the system have to deal with the the fact that many teams' modules are running in the same service. Debugging and troubleshooting gets harder. Logging and telemetry also tends to get more complicated.
4. Dependency coupling. Everyone has to use the exact same versions of everything and everyone has to upgrade in lockstep. You can work around this with module systems that allow dependency isolation, but IMO this tends to lead to its own complexity issues that make it not worthwhile.
5. Module API boundaries. In my experience, developers have an easier time handling service APIs than library APIs. The API surface area is smaller, and it's more obvious that you need to handle backwards compatibility and how. There is also less opportunity to "cheat", or break encapsulation, with service boundaries compared to library boundaries.
In practice, for dividing up code, libraries and modules are the less popular solution for server-side programming compared to services for good reasons. The downsides are not worth the upsides in most cases!
Why do we care about network latency, when we JUST established that microservices are about scaling large development teams? I have no problem with hackers ranting about slow, bloated and messy software architecture...but this is not the focus of discussion as presented in the article.
And then this conclusion:
> The key is to establish that common architectural backplane with well-understood integration and communication conventions, whatever you want or need it to be.
...so, like gRPC over HTTP? Last time I checked, gRPC is pretty well understood from an integration perspective. Much better than Enterprise Java Beans from the last century. Isn't this ironic? And where are the performance considerations for this backplane? Didn't we criticize microservices before because they have substandard performance?
> common architectural backplane with well-understood integration and communication conventions, whatever you want or need it to be
Regardless of the tech used to implement, this paradigm needs to be solved to have a good system. That backplane is not an implementation, but a set of understood guiderails for inter-module communication.
Even with gRPC, both sides need to know what to call and what to provide and expect in response. That's the "conventions" part. Having consistency is more important than the underlying tech. Just simple ReST over HTTP works just as well as gRPC.
Take for example Gartner's definition:
> A microservice is a service-oriented application component that is tightly scoped, strongly encapsulated, loosely coupled, independently deployable and independently scalable.
That's not too controversial. But... as a team why and when would you want to implement something like this? Again, let's ask Gartner. Here are excerpts from "Should your Team be using Microservice Architectures?":
> In fact, if you aren’t trying to implement a continuous delivery practice, you are better off using a more coarse-grained architectural model — what Gartner calls “Mesh App and Service Architecture” and “miniservices.”
> If your software engineering team has already adopted miniservices and agile DevOps and continuous delivery practices, but you still aren’t able to achieve your software engineering cadence goals, then it may be time to adopt a microservices architecture.
For Gartner, the strength of Microservice Architecture lies in delivery cadence (and it shouldn't even be the first thing you look at to achieve this). For another institution it could be something else. My point is that when people talk about things like Microservices they are often at cross-purposes.
Another way to put it is that teams that share parts of the same codebase introduce low level bugs that affect each other, and most organizations are clueless about preventing it and in some cases do not even detect it.
Part of the reason is that anything that leaves the application package increases the error potential exponentially.
Also, modules "bottleneck" functionality, and allow me to concentrate work into one portion of the codebase.
I'm in the middle of "modularizing" the app I've been developing for some time.
I found that a great deal of functionality was spread throughout the app, as it had been added "incrementally," as we encountered issues and limitations.
The new module refines all that functionality into one discrete codebase. This allows us to be super-flexible with the actual UI (the module is basically the app "engine").
We have a designer, proposing a UX, and I found myself saying "no" too often. These "nos" came from the limitations of the app structure.
I don't like saying "no," but I won't make promises that I can't keep.
BTW: The module encompasses interactions with two different servers. It's just that I wrote those servers.
Most likely the question is not well defined in the first place.
It's like we're stuck assembling applications by stacking pre-existing Lego sets together with home-made, purpose-built, humongously-sized Lego bricks made of cardboard acting as glue.
I've ranted a bit about that here: https://news.ycombinator.com/item?id=34234840
> ... the Fallacies of Distributed Computing.
I feel like I’m taking crazy pills, but at least I’m not the only one. I think the only reason this fallacy has survived so long this cycle is because we currently have a generation of network cards that is so fast that processes can’t keep up with them. Which is an architectural problem, possibly at the application layer, the kernel layer, or the motherboard design. Or maybe all three. When that gets fixed there will be a million consultants to show the way to migrate off of microservices because of the 8 Fallacies.
When a critical vulnerability comes out for whatever language you are using, you now have to patch, test and deploy X apps/repos vs much fewer if they are consolidated repositories written modularly. Same can be said for library/framework upgrades, breaking changes in versions, deprecated features, taking advantage of new features, etc.
Keeping the definition of runtimes as modular as the code can be instrumental in keeping a bunch of related modules/features in one application/repository. One way is with k8s deployments and init params where the app starts specific modules which then lends itself to be scaled differently. I'm sure there are home-grown ways to do this too without k8s.
I'm glad that this post was written so that we can look at widely accepted ideas a little more critically.
Monoliths are way slower to deploy than microservices, and when you have hundreds or thousands of changes going out every day, this means lots of changes being bundled together in the same deployment, and as a consequence, lots of bugs. Having to stop a deployment and roll it back every time a defect is sent to production would just make the whoe thing completely undeployable.
Microservices have some additional operational overhead, but they do allow much faster deployments and rollbacks without affecting the whole org.
Maybe I am biased, but I would love an explanation from the monolith-evangelist crowd on how to keep a monolith able to deploy multiple times a day and capable of rolling back changes when people are pushing hundreds of PRs every day.
Are those thousands of developers working on a single product? If so, then I'd argue that you have way too many developers. At that point, you'd need so many layers of management that the overall vision of what the product is gets lost.
You're likely right.
I think what most detractors of microservices are pointing to is that most companies don't reach thousands of devs in size. Or even hundreds.
People just read up on whatever seems to be the newest, coolest thing. The issue is that MS articles are usually coming from FAANG/ex-FANNG. These companies are solving problems that 99% of others do not.
As engineers we should be looking for the most effective solutions to a given business problem. Sadly, I see engineers with senior/staff titles just throwing cool tech terms/libs around. Boring tech club ftw
As a CTO, I couldn't agree more. For our internal product, we use 100% boring technologies. The most "modern" you'll find is a React SPA.
I sigh when clients want to go the microservices route for a team of just a few developers. When you want to use NextJs for their tables&forms app. When they choose to use Kubernates instead of a couple EC2 instances.
Don't get me wrong, these technologies are great for us because we can charge more for the wasted human time developing these overengineered solutions. But I always, for my peace of mind, try to talk them out of them. Sometimes works, sometimes doesn't, at the end of the day it's their money.
The article gets it right in ny opinion.
1. It has a lot to do with organisational constraints.
2. It has a lot to do with service boundaries. If services are chatty they should be coupled.
3. What a service does must be specified in regards to which data it takes in and what data it outputs. This data can and should be events.
4. Services should rely and work together based on messaging in terms of queues, topics, streams etc.
5. Services are often data enrichment services where one service enrich some data based on an event/data.
6. You never test more than one service at a time.
7. Services should not share code which is vibrant or short lived in terms of being updated frequently.
8. Conquer and divide. Start by developing a small monolith for what you expect could be multiple service. Then divide the code. And divide it so each coming service own its own implementation as per not sharing code between them.
9. IaaS is important. You should be able to push deploy and a service is setup with all of its infrastructure dependencies.
10. Domain boundaries are important. Structure teams around them based in a certain capability. E.g. Customers, Bookings, Invoicing. Each team owns a capability and its underlying services.
11. Make it possible for other teams to read all your data. They might need it for something they are solving.
12. Don't use kubernetes or any other orchestra unless you can't it what you want with cloud provider paas. Kubernetes is a beast and will put you to the test.
13. Services will not solve your problems if you do not understand how things communicate, fail and recovers.
14. Everything is eventually consistent. The mindset around that will take time to cope with.
A lot more...
Well, in our +20 product teams with all serving different workflows for 3 different user types, the separate micro services are doing wonders for us for exactly the things you've listed.
My comment should just stop here to be honest.
Microservice and modularity is orthogonal, it's not the same.
Modularity is related to business concept, microservice is related to infrastructure concept.
For example, i could have a module A which is deployed into microservide A1 and A2. In this case, A is almost abstract concept.
And of course, i could deploy all modules A, B, C using 1 big service (monothlic).
Moreover, i could share one microservice X for all modules.
All confusion from microservice, is made from the misconception that microservice = module.
Worse, most of "expert advice" which i've learnt actually relate Domain Driven Design to Microservice. They're not related, again.
Microservice to me, is to scale. Scale the infrastructure. Scale the team (management concept).
Then you can have a mono repo that deploys to multiple micro services / cloud functions / lambdas as needed depending on code changes and programmers don’t have to worry about RPC or json when communicating between modules and can just call the damn function normally.
Though there were many attempts to do it, I would just mention Erlang and Akka.
The answer to your question is close to the answer for "what's wrong with Akka"
Solving this in the general case for every function call is difficult. Pure functions are idempotent, so you can retry everything until nothing fails.
But once you add side effects and distributed state, we don’t know how to solve this in a completely generalized and performant way.
But author clearly avoided the real reasons why you actually need to split stuff into separate services: 1. Some processes shouldn't be mixed in the same runtime. Simple example batch/streaming vs 'realtime'. Or important and not important. 2. Some things need different stack, runtimes, frameworks. And is much easier to separate them instead of trying to make them coexist.
And regarding 'it was already in Simpsons' argument, I don't think it should even be considered as argument. If you are old enough to remember EJB, you don't need to be explained why it was a bad idea from the start. Why services built on EJB were never scalable or maintainable. So even if EJB claimed to cover the same features as microservices right now, I'm pretty sure EJB won't be a framework of choice for anybody now.
Obviously considering microservices as the only _right_ solution is stupid. But same goes for pretty much any technology out there.
If I run a monolith and one least-used module leaks memory real hard, the entire process crashes even though the most-used/more important modules were fine.
Of course it's possible to run modularised code such that they're sandboxed/resources are controlled - but at that point it's like...isn't this all the same concept? Managed monolith with modules vs microservices on something like k8s.
I feel like rather than microservices or modules or whatever we need a concept for a sliding context, from one function->one group of functions->one feature->one service->dependent services->all services->etc.
With an architecture like that it would surely be possible to run each higher tier of context in any way we wanted; as a monolith, containerised, as lambdas. And working on it would be a matter of isolating yourself to the context required to get your current task done.
It’s old. The examples are in PL/I. But his framework for identifying “Functional Strength” and Data Coupling is something every developer should have. Keep in mind this is before functional programming and OOP.
I personally think it could be updated around his concepts of data homogeneity. Interfaces and first class functions are structures he didn't have available to him, but don’t require any new categories on his end, which is to say his critique still seems solid.
Overall, most Best Practices stuff all seem either derivative of or superfluous to this guy just actually classifying modules by their boundaries and data.
I should note, I haven't audited his design methodologies, which I'm sure are quite dated. His taxonomy concerning modules was enough for me.
"The Art of Software Testing" is another of his. I picked up his whole corpus concerning software on thriftbooks for like $20.
Having said that, if you want to eke out another 3x throughput improvement, then by all means, grab your OpenSwoole or ReactPHP or AMPHP and go to town. But PHP already has Fibers, while OpenSwoole still has coroutines. Oh yeah, and the try/catch in OpenSwoole is broken so good luck catching stuff.
Queues are awesome, just use queues almost all the time and then either build micro services or don’t.
It is about where the shoe fits. If you become too heavily dependent on modules you risk module incompatibility due to version changes. If you are not the maintainer of your dependent module you hold a lot of risk. You don't get that with microservices.
If you focus too much on microservices you introduce virtualized bloat that adds too much complexity and complexities are bad.
Modules are like someone saying it is great to be monolithic. Noone should upright justify an overly complicated application or a monolithic one.
The solution is to build common modules that are maintainable. You follow that up with multi-container pods have them talk low level between each other.
Stricking that exact balance is what is needed not striking odd justifications for failed models. It is about, "What does my application do?" and answering with which design benefits it the most.
The way I usually describe my preferred heuristic to decide between modules and microservices is:
If you need to deploy the different parts of your software individually, and there’s a cost of opportunity in simply adopting a release train approach, go for microservices.
Otherwise, isolated modules are enough in the vast majority of cases.
A failed lookup of a function is greeted by "Do you want to import X so you can call foo()?". Having a battery of architectural unit tests or linters ensuring at module foo doesn't use module bar feels like a crutch.
Now, it might seem like making microservices just to accomplish modularization seems like a massive overkill and sa huge overhead for what should be accomlished at the language level - and you'd be right.
But that leads to the second largest appeal, which is closely related. The one thing that kills software is the big ball of mud where you can't really change that dependency, move to the next platform version or switch a database provider. Even in well-modularized code, you still share dependencies. You build all of it on react, or all the data is using Entity Framework or postgres. Because why not? Why would you want multiple hassles when one hassle is enough? But this really also means that when something is a poor fit for a new module, you shoehorn that module to use whatever all the other modules use (Postgres, Entity Framework, React...). With proper microservices, at least in theory you should be able to use multiple versions of the same frameworks, or different frameworks all together.
It should also be said that "modules vs microservices" is also a dichotomy that mostly applies in one niche of software development: Web/SaaS development. Everywhere else, they blur into one and the same, but sometimes surfacing e.g. in a desktop app offloading some task to separate processes for stability or resource usage (like a language server in an IDE).
The typical example is that when you have to explain something is the job of X and Y. Usually this means that X and Y are breaking those boundaries. Just make a semi-private (or even public) API only used for that thing and you have a broken boundary again. Or push it on a message queue, etc.
I think, it certainly helps, but then again it doesn't prevent it. Having modules you have the same effect.
So in the end you get more spots where things can get wrong and more operational complexity. If you stick to using them right, I think you can also stick to using modules right with less complexities, better performance, easier debugability, fewer moving parts.
Hackers will find a way to do hacky things everywhere. ;)
Also this whole discussion reminds me of Linus discussing how he thinks micro kernels add complexity a very long time ago. Not sure if they should be considered modules or microservices though.
Sharing global state and so on while maybe it shouldn't be done lightly, without thinking about it can and does make sense. And in most modern environments it's not like the most quoted issues can happen too easily.
Also I strongly agree with pointing out that this is actually a niche topic. It's mostly big because that niche is where probably the majority of HN and "startup" people spend their time.
But in general,just write some damn code. Presumably you have to write code, because building a software engineering department is one of the most difficult things you can do in order to solve a business problem. Even with the smartest engineers in the world (which you don't have), whatever you ship is inevitably going to end up an overly complex, expensive, bug-riddled maintenance nightmare, no matter what you do. Once you've made the decision to write code, just write the damn code, and plan to replace it every 3-5 years, because it probably will be anyway.
Someday I will find a way to untangle this wishlist enough to turn into a design.
Using hindsight, most of the systems that pretend that the network part of invoking RPCs is "easy" and "simple" and "local" end up being very complex, slow and error prone themselves.
See DCOM, DCE, CORBA and EJBs RMI.
You need instead a efficient RPC protocol that doesn't hide the fact it is a RPC protocol - like Cap'n Proto (it's "time-traveling RPC" feature is very interesting).
RPC vs local calls is trivial in comparison and you can get that level of transparency out of the box with functional programming - it's just data in data out.
1) A remote call – inherently – fail. However, some local calls never fail. Either because they are designed to never fail or because you have done the required checks before executing the call. A remote call can fail because the network is unreliable and there is no way around that (?).
2) A remote call – inherently – can be very slow. Again because of the unpredictable network. A local call may be slow as well but usually either because everything is slow or because it just takes a while.
So if you have a call that may or may not be local you still have to treat it like a remote call. Right?
I think having a certain set of calls that may or may not be executed locally is not that bad. Usually it will just be a handful of methods/functions that should get this "function x(executeLocally = false|true)" treatment - which is prob. an acceptable tradeoff.
Push the monolith to production, monitoring performance, and if and when performance spikes in an unpleasant way, "offload" the performance intensive work to a separate job server that's vertically scaled (or a series of vertically scaled job servers that reference a synced work queue).
It's simple, predictable, and insanely easy to maintain. Zero dependency on third party nightmare stacks, crazy configs, etc. Works well for 1 developer or several developers.
A quote I heard recently that I absolutely love (from a DIY construction guy, Jeff Thorman): "everybody wants to solve a $100 problem with a $1000 solution."
That said, server bottlenecks are not the only thing (micro)services are trying to address.
Ideally we could flexibly deploy services/components in the same way as WebLogic EJB. Discovery of where components live could be handled by the container and if services/components were deployed locally to one another, calls would be done locally without hitting the TCP/IP stack. I gather that systems like Kubernetes offer a lot of this kind of deployment flexibility/discovery, but I'd like to see it driven down into the languages/frameworks for maximum payoff.
Also, the right way to do microservices is for services to "own" all their own data and not call downstream services to get what they need. No n+1 problem allowed! This requires "inverting the arrows"/"don't call me, I'll call you" and few organizations have architectures that work that way - hence the fallacies of networked computing reference. Again, the services language/framework needs to prescribe ways of working that seamlessly establish (*and* can periodically/on-demand rebroadcast) data feeds that our upstreams need so they don't need to call us n+1-style.
Microservices are great to see, even with all the problems, they DO solve organizational scaling problems and let teams that hate each other work together productively. But, we have an industry immaturity problem with the architectures and software that is not in any big players' interest in solving because they like renting moar computers on the internet.
I have no actual solutions to offer, and there is no money in tools unless you are lucky and hellbent on succeeding like JetBrains.
One thing I've observed in a microservice-heavy shop before was that there was the Preferred Language and the Preferred Best Practices and they were the same or very similar across the multiple teams responsible for different things. It lead to a curious phenomenon, where despite the architectural modularity, the overall SAAS solution built upon these services felt very monolithic. It seemed counter-productive, because it weakened the motivation to keep separation across boundaries.
If I want to calculate the price of a stock option, that's an excellent candidate to package into a module rather than to expose as a microservice. Even if I have to support different runtimes as presented in the article, it's trivial.
A different class of problem that doesn't modularize well, in shared library terms, is something with a significant timing component, or something with transient states. Perhaps I need to ingest some data, wait for some time, a process the data, and then continue. This is would likely benefit from being an isolated service, unless all of the other system components have similar infrastructure capabilities for time management and ephemeral storage.
I'm sure I could gain many of the advantages of microservices through an OSGi monolith, however, an OSGi monolith is not the hot thing of the day, I'm likely to be poorly supported if I go down this route.
Ideally I also want some of my developers to be able to write their server on the node ecosystem - if they so choose, and don't want updating the language my modules run on (in this case the JVM) to be the biggest pain of the century.
Besides, once my MAU is in the hundreds of thousands, I probably want to scale the different parts of my system independently anyway - so different concerns come in to play.
Yes, the same often happens in microservices, but the extra complexity of a distributed system provides a slightly stronger nudge to decouple that means some teams at the margin do achieve something modular.
I’m something of a skeptic on modular monoliths until we as an industry adopt practices that encourage decoupling more than we currently do.
Yes, in theory they’re the same as microservices, without the distributed complexity, but in practice, microservices provide slightly better friction/incentives to decouple.
I've done monoliths and microservices. I've worked in startups, SMEs and at FAANGS. As usual, nothing in this article demonstrates that the person has significant experience of running either in production. They may have experience of failure.
In my experience, microservices are simply one possible scaling model for modules. Another way to scale a monolith is to just make the server bigger: the One Giant Server model.
If you have a fairly well defined product, that needs a small (<30) number of engineers, then the One Giant Server model might be best for you. If you have a wide feature set, that requires >50 engineers, then microservices is probably the way to go.
There is no noticeable transition from a well implemented monolith with a small team into a well implemented Giant Server with a small team. Possibly some engineers are worrying about cold start times for 1TB of RAM, but that's something that can happen well ahead of time, and hardware qualification is something one needs to do for microservices too. Some of the best examples of Giant Servers are developed by small teams of very well paid developers.
The transition from a monolith to a set of microservices, however is a very different affair. Unfortunately, the kind of projects that need to go microservices are often in a very poor state. Many such adventures one reads about are having to go to microservices because they've already gone to One Giant Server and those are now unable to handle the load. Usually these stories are accompanied by a history of blog posts about moving fast and breaking things, yolo, or whatever is cool. The transition to microservices is difficult because there are no meaningful modules: no modules, or modules that all use each other's classes and functions.
I don't believe that, once a particular scale is reached, microservices are a choice. They are either required or they are not. Either you can scale successfully with One Giant Server, or you can't.
The problem is that below a certain scale, microservices are a drag. And without microservices, it's very easy for inexperienced teams to fail to keep module discipline.
I like that architecture - the services are abstracted with clearly defined boundaries and they are easy to navigate / discover. Not sure if Java modules satisfy the concerns of the author or other HN users, but I liked it.
Going back to systems thinking, flow control (concurrency and rate limiting) and API scheduling (weighted fair queuing) are needed to make these architectures work at any scale. Open source projects such as Aperture[0] can help tackle some these issues.
Monolith is better in almost every way. The only thing that still bothers us is the build/iteration time. We could resolve this by breaking our gigantic dll into smaller ones that can be built more incrementally. This is on my list for 2023.
I have worked on a monolith that solved the exact same problem. And it was straightforward to maintain and upgrade.
I feel sorry for future developers who will have to take over and maintain micro-services systems created today. Teams that can’t create maintainable, well designed monoliths, will create an even bigger cluster f** using micro-services.
You only need to agree on an API between the modules and you're good to go!
Microservices suck dick and I hate the IT industry for jumping on this bandwagon (hype) without thoroughly discussing the benefits and drawbacks of the method. Debugging in itself is a pain with Microservices. Developing is a pain since you need n binaries running in your development environment.
We delivered many talks on that subject and implemented an ultimate tool for that: https://github.com/7mind/izumi (the links to the talks are in the readme).
The library is for Scala, though all the principles may be reused in virtually any environment.
One of the notable mainstream (but dated) approaches is OSGi.
https://michaelfeathers.silvrback.com/microservices-and-the-...
I've often noticed that these boundaries are not considered when carving out microservices.
Subsequently, workarounds are put in place that tend to be complicated as they attempt to implement two phase commits.
It's on the same page as "yes, we could have written this in assembler better" or "this could simply be a daemon, why is it a container?"
As if an agile, gitops based, rootlessly built, microservice oriented, worldwide clustered app will magially solve all your problems :D
If i learned anything it's to expect problems and build a stack that is dynamic enough to react. And that any modern stack includes the people managing it just as much as the code.
But yes, back when ASP.NET MVC came out i too wanted to rebuild the world using c# modules.
Mastery is substantially about figuring out what rules of thumb and aphorisms are in place to keep beginners and idiots from hurting themselves or each other, and which ones are universal (including some about not hurting yourself, eg gun safety, sharps safety, nuclear safety).
And this is how I work all the time.
In Fuchsia, the device driver stack can be roughly split into three layers:
* Drivers, which are components (~ libraries with added metadata) that both ingest and expose capabilities,
* A capability-oriented IPC layer that works both inside and across processes,
* Driver hosts, which are processes that host driver instances.
The system then has the mechanism to realize a device driver graph by creating device driver instances and connecting them together through the IPC layer. What is interesting however is that there's also a policy that describes how the system should create boundaries between device driver instances [2].
For example, the system could have a policy where everything is instantiated inside the same driver host to maximize performance by eliminating inter-process communication and context switches, or where every device driver is instantiated into its own dedicated driver host to increase security through process isolation, or some middle ground compromise depending on security vs performance concerns.
For me, it feels like the Docker-like containerization paradigm is essentially an extension of good old user processes and IPC that stops at the process boundary, without any concern about what's going on inside it. It's like stacking premade Lego sets together into an application. What if we could start from raw Lego bricks instead and let an external policy dictate how to assemble them at run-time into a monolithic application running on a single server, micro-services distributed across the world or whatever hybrid architecture we want, with these bricks being none the wiser?
Heck, if we decomposed operating systems into those bricks, we could even imagine policies that would also enable composing from the ground up, with applications sitting on top of unikernels, microkernels or whatever hybrid we desire...
[1] https://fuchsia.dev/fuchsia-src/development/drivers/concepts...
[2] https://fuchsia.dev/fuchsia-src/development/drivers/concepts...
with a good branching strategy
Just because you can, doesn't mean you should.