Disasters I've seen in a microservices world
world.hey.com
world.hey.com
Most microservices I've seen implemented are not actually microservices. Most teams (unless they are developing "simple" web or mobile applications) have no idea who is consuming their services, or how, or if what they've made is working correctly. They frequently don't have an understanding of the different models of releasing software, much less of a complete system architecture, failure modes, consistency models, reliability estimates, performance limits, etc. Mostly what I see time and again is a team of people who just write some code that seems to work for them, and then go home, without ever considering how or if it's working in the real world.
I don't know anything about modern computer science education. But it appears that new developers have absolutely no idea how anything but algorithms work. It's like we taught them how hammers and saws and wrenches work, and then told them to go build a skyscraper. There are only two ways I know of for anyone today to correctly build a modern large-scale computer system: 1) Read every single book and blog post, watch every conference talk, and listen to every podcast that exists about building modern large-scale computing systems, and 2) Spend 5+ years making every mistake in the book as you try to build them yourself.
It feels like the industry mostly just re-learns the same mistakes over and over like we're in Groundhog's Day (we're the extras, not Bill Murray). But it's equally possible that I just lack perspective and am expecting too much. Maybe the auto industry at the turn of the 20th century also spent decades re-learning the same lessons over and over, as the novelty of mass-producing complex systems continued to elude us. Hell, the new auto companies still don't get it right.
“There aren’t any new problems; only new engineers.”
- Quote I read online a long time ago
Agree with all your points, but there is a third way: apprenticeship with older engineers. I advise new engineers to find their balance between not accepting the status quo but also realizing software is way more complicated than they think.
For example: successfully versioning an API over time while supporting live deployments is (I think) not something you’re going to learn in school.As a further aside... I’ve heard older attorneys complain about newly minted lawyers. They work to pass school and then the bar. But when they start lawyering, new grads don’t know the first thing about how to actually file something at the courthouse. I’m saying, I don’t think it’s constrained to our industry.
That's unlikely to work: our industry has an ageism problem, combined with the low prospects of promotions/raises which results in engineers having to frequently job-hop to get ahead. This results in most organizations having very limited institutional memory. Other compounding problems are a discussion for another day (title inflation: e.g. "Senior engineer" with 2 years total experience)
CS was focused mostly on the broad underlying structure of computing, things like how software interacts with hardware (this was gutted a year or two before me), algorithms, programming paradigms (functional/etc), the theory behind databases (3rd normal form sound familiar?), and so on. The idea was to give you a really good foundation such that you could self-learn much much easier after college, but also unfortunately that you had to do some of that selfglearning right away because it didn't prep you for getting things done.
I only took one or two ITM courses out of curiosity, so I'm less confident in describing it in broad strokes, but the impression I got is that it was for teaching specific skills while not "wasting" time on better understanding how they work. Very get-things-done oriented. The web development one, for example, used a lot of jquery, but never touched on what was going on under the hood. I can see it working as preparation for being able to do a job, but it left you kind of stuck if you hit a roadblock and didn't know where to look for answers.
If I had to stereotype them, a CS student would spend time writing an algorithm for something instead of finding an existing library, while an ITM student would look for a library and if they can't find one say they can't do it. But the ITM student wouldn't get lost implementing and end-to-end project as long as they could use tech they already know, whereas the CS student has never done such a thing before.
I think you have a misconception about what manufacturing is like. Waterfall development is just as unrealistic in manufacturing as it is in software, the scales are just different. People create CAD models of the machines in a factory all the time but those are architectural diagrams - when it comes time to build, the people on the ground (the equivalent of the software engineers) have to figure out all of the details, often using no more than trial and error guided by a little intuition. Experienced individuals know more about the pitfalls and gotchas but they're still creating something new even if they're going off some grand design. Becoming a competitive manufacturer in something simple like ball point pen balls is a years long process of mastering the "craft," gathering real world data, and incrementally improving. Factors like a single critical employee's physical disability can decide the layouts of entire factories so there are no cookie cutter factories for them to build.
The real world is a lot messier than software. Just like we spend most of our time chasing bugs and implementing features, they spend most of their time building or troubleshooting interfaces between machines and processes, optimizing, and so on.
Is it? Aren't a lot of the largest systems build using waterfall, or they certainly were in the past. When i was taught waterfall at university you were allowed to reiterate steps, yet these is a constant straw man that you only gte to move down the waterfall.
This is true, but also in life. I'm so sad I can't just pass my knowledge to my daughter to give her a 30 year head start. I remember telling her 'dont touch that stove it's hot', she nods, and soon after is crying she burnt her hand. I wanted to be mad but it made me realize the one universal truth: advice and information is great, but humans just seem to learn most from mistakes. Made me realize how much time and energy and knowledge is lost, as every human learns a lot of the same things over and over. I'm sure there was a caveman telling his kid not to touch the fire thousands of years ago, too. And here we are. So don't be too cynical and depressed about this cycle, it's pretty natural.
IMO, there's a few ways to try and take a pragmatic approach to the problem above:
1. Make simple ways to do complex things. A library, a language, a framework, they all go after the same thing when designed well: batteries included complex functionality that the end user can't shoot themselves in the foot with. It seems like in this domain (microservices), we are still early and need better support for small companies to spin this up. I've found my current company runs microservices pretty well, but I know it took a TON of infrastructure, support, and education to get it to that point. Few companies can afford to have that.
The danger with this approach is that people who understand the internals become rare, and surface level only knowledge can lead to issues when you get to edge cases. I'll still take it, but wanted to call it out. The linux kernel is a great example of what that looks like down the line.
2. Make as many places as possible for developers in training to fail and make those 5 years of mistakes with low stakes consequences. In some ways, the way the industry hires already does this, but I'm sure there are better ways we can come up with versus the current sandboxes. If your frustration boils over, this is a place to channel it positively.
I think some of the problem is that there's no real standardized training for developers. Many have 4 weeks of 'bootcamp' and are cut into the wild.
And I can't be too negative about it, I didn't even have that. I would just say I learned from 20 years of reading and mistakes.
One thing that's crazy to me is that when people make mistakes, in most cases it's forgotten or swept under the rug. Whereas if they do a full post-mortem and share it with the entire engineering organization, they're teaching everyone a valuable lesson (possibly multiple lessons) for free. If that lesson prevents the same situation from happening again even one time, they'll have "paid off" the first mistake, and it cost them nothing. Free money!
That's not to say it's a waste. Once a person gets bitten, they look for resources and learn from that. But until then, we're basically all stupid apes making the same mistakes in perpetuity.
I don't think this is some sort of innate human thing, I think it's purely down to ego. You think you know more or you're better than the person who made the mistake, and therefore you don't listen to them because you would never do that.
A few months ago, i engaged to help fix a troubled system. The thing was built around kubernetes and the architecture was basically aligned with how a bunch of contractors were hired. Nobody really understood how it worked and it performed poorly. So poorly in fact that I reimplemented a core component in a shell script, mostly to quickly understand how it worked, and to my surprise, the thing ran faster on my laptop at moderate load than the actual system!
Sounds like a bad IT project, but to be honest whole companies run this way, saving pennies while setting dollars on fire.
Channeling a post from earlier today — I’m not sure if companies are younger and dumber today, but I’m sure that you are older and wiser now than you were then. And I’m sure you are better at seeing the idiocy in companies. Bad projects have always been around. But it isn’t until you have more experience that you can fully appreciate a bad project — particularly when you’re in the middle of it!
Possibly a case of teaching to the test (interview, in this case)?
My CS education is a couple decades back, but while we had an algorithm course, most credit hours were spent on more practical things like building compilers, building an OS, building a shell, SQL, networking stacks, encryption, software engineering complex systems and so forth.
Is the current CS landscape more algorithm-focused because the FAANGs have made that the only stick to measure by?
[1]they later added a Software Engineering degree which focused more on architecture and designing large scale software system, software project management and so on.
There is a big bubble in tech salaries that is going to get popped as more and more people realize you just don't need a developer for most things and can get away with no-code for the majority of your needs.
This will act as a force multiplier for some, but I think for most developers it will see a correction in their salaries to more reasonable levels.
Someone should coin a new law for how programmers have to rediscover Brooks's law every 5-10 years. The issue with microservices, as it has always been, is that you need an enormous company to brute force the communication pathways and maintenance overhead for it to all work. And by work, I don't mean function efficiently (as the Multics vs Unix video shows). I mean just function. Just work at all. The Multics team had all the devs and the Unix team was two guys doing laps around the Multics team. Because they had the mathematics on their side.
Remember the bad old days of memory thrashing? That's what happens to teams that do not have enough bandwidth to properly maintain the dozens of services they are responsible for. Your organization gets frozen.
This is what we all get for taking advice from the Googles and Facebooks of the world. Google has like a billion lines of code in a monorepo. They do not do things remotely like 99% of the businesses out there. They are sitting on huge piles of money that lets them be incredibly inefficient for decades.
Google has a dedicated organisation for maintaining all the infrastructure required for and around microservices. We are talking about 500-1000 engineers just improving and maintaining the infrastructure.
I don't have to worry about things like distributed tracing, monitoring, authentication, logs, logs search, logs parsing for exceptions, anamoly detection on logs, deployment, release management etc. They are already there and is glued pretty well that 90 percent of engineers don't even have to spend more than a day to understand all that.
1) something has different run time needs than something else. Think CPU, memory, network. Breaking stuff up allows you to make different choices here. Although I have dodged this by simply deploying the same monolith and configure it to do different things.
2) Something needs to be developed by a different team and for whatever reasons you don't want those teams to be too dependent. It's a bad reason but it's valid in a lot of companies where certain teams just need to be engineered around or where there is a fundamental lack of trust between different parts of the org chart. Conway's law is a thing. It's the most common reason to do micro services.
3) You have two things depending on each other (cyclical dependency) and you want to reuse that thing. Extracting it to a third thing is a common way out. It's true for almost any component technology. If you have two components, you'll find a reason to create a third. And a fourth. And so on. However, consider using something less dramatic. E.g. code libraries are a valid choice. Or having an extra module in your source tree.
Everything else is just needlessly/prematurely increasing overhead, deployment friction, etc. You get more things to monitor, deploy, manage the roadmap off, worry about, specialize in, etc. Big bloated organizations do micro services because they are big and bloated. Many smart startups keep this nonsense to a minimum. Of course some startups start out being over funded and bloat too early. VC money is great and sometimes requires over engineering like this (i.e. impress the suits). I've heard more than a few CTOs boast their multi cloud strategy and micro service architecture. In my mind that translates to we funnel a lot of VC money to Amazon and pay people full time to do just that. Ridiculous monthly bills and no users or traction is a common pattern in that world.
We split up our services when the majority of changes made to a service can be made independently of everyone else. So then each of our services is a different codebase, and they each have a completely different rate of change from each other. You very rarely would ever make changes across all services as once, only when changing what's being communicated between some set of services. Some services are large, some are small, and many of them fall into the category of -- developed for a while then stabilized and now basically rarely every touched just works.
So then the advantage is that you always have a small mental working set. You're focusing on a service at a time, have less to worry about breaking everything else with your changes. And then when you deploy, even if everything goes horribly wrong, it's just your one service that is down, and everything else is up and running, and you'll just have to process your queued messages once you came back online.
And then of course each service is smaller, so less code, less tests, faster to compile, faster to run through the pipeline, faster to release. And you're only ever doing rolling deploys on a small sub-section of your infrastructure, never the whole thing at once.
Unless you have the specific need of that piece of code not being in the same process/container/computer/network.
Unless your linking time is prohibitively high (but even very large software takes minutes, not hours), the benefit to compile time should be the same. But maybe "redeploying" is more disruptive (I don't know, I haven't done web development for something approaching 10 years now), and depending on the language, problems in one part can crash the whole process. (But depending on the actual process model used, shouldn't it just restart? Unless the language is not sufficiently memory safe, e.g. it's C, and you get way worse problems with memory smashers.)
You would have a hard time to make a resource efficient distributed in-process key-value store.
You could have an in-process library but this would cause your two programs to have either:
1. Store different data
2. All writes would need to be distributed to all processes (looping back to needing a network connection).
Yep, exactly. I've built more than one microservice in my time that loads up a few GB of data into RAM and efficiently queries it for a lot of services. At that point the line between DB and service gets blurry.
But, essentially, the TCP connection between your services isn't necessarily a big issue. If there is some logic/operation/access that is complex it makes sense to chuck a language-independent interface in front of it for most use cases. There are some where performance trumps maintenance but where it doesn't I think this abstraction is helpful.
Again, my, and likely OPs point, was that distinct, modular development (and for example the mentioned low compile times) is not contingent on separate services and IPC or even RPC, which come with a whole other can of worms.
(BTW your scope of shared library is also pretty narrow. For persistence for example, SQLite is essentially a persistence layer that is a shared library. And some libraries do certainly spawn threads for their work. But that’s not really the point.)
SQLite is a great library, but that's more of a business-agnostic tool. How do you encapsulate a dedicated portion of business logic, let's say, managing image datasets for building machine learning models. e.g. I want some apis for basic CRUD on datasets, i want to automatically create new versions of existing datasets as images get published, i want to generate stats in the background for these datasets, i want a ui to review and manage datasets for model building, and then i want all this packaged in a single library?
Also, to clarify, I'm not playing a game here. I don't care about theoreticals, I want to hear about how people have actually solved their problems, so I can then evaluate and use those approaches. So I'm trying to get at something a little bit more specific and less handwavy.
In a monolithic architecture these would run in the same address space, be separated into modules by the language's own means of abstraction, and the communication between them would be done by function or method calls.
I don't see what HTML files have to do with it, what do you think microservices are?
1. Multiple languages/tech stacks. Have something like protobufs and now any language can "import" your library instead of it just being things that can handle cffi.
2. State management. For example, redis couldn't be an in-process library. There are in-process kv stores but that's not the idea of redis.
3. Transparent monitoring for all languages. I can passively detect all retries, latencies, request counts, etc since there is a protocol.
4. Centralized fixes. Similarly to dynamic linking I can now update a single component and "fix the world" rather than redeploying every binary.
5. Canaries: I can deploy version N+1 to a mirror of production traffic and make sure it doesn't explode. Then I can route 1%, 10%, 50%, 100% of traffic and check those metrics as well.
6. If something goes horribly wrong I can rewrite the entire service for performance or maintenance or whatever within a tightly bounded amount of time. Discord has many blog posts about doing this for some of their services.
Great talk: https://www.youtube.com/watch?v=-UKEPd2ipEk
This is what I'm excited about--enabling a team to use the most powerful tools that address their own problems, rather than whichever Blub is approachable enough for the entire staff.
Say you run a factory, and everyone needs a hammer. When there's only two teams, two kinds of hammer is fine. But fast forward to 10 teams. Can you always get all of those hammers when you need them? What if team #3 needs extra support? You need to find a new engineer who knows hammer #3, or take an engineer off of team #6 and train them on how to use hammer #3. And if all the hammers need a slight modification specific to that company, you need to make 10 different modifications. And how will all the teams benefit from shared hammer tips-n-tricks if they all use different hammers? There's a lot more overhead than you think at first.
You don't need microservices architectures for that. Even a clean architecture on a monolith ensures that you can compartmentalize the problem domain.
At most, just refactor out that responsibility into a library/component/module.
Important to note here is that a single microservice can actually be quite large with anywhere from 3-12 people working on just one or a few services. A single service could be larger than a startups monolith.
I would argue that this is really the only reasonable use of microservices. If you can fully understand your monolith, you shouldn’t change to microservices. If you fully understand or even know about all microservices in your organization, chances are that you’re doing it wrong.
That has a huge impact because now you need to set up the data store and potentially perform migrations, which means you may need to change how you do deployments. It may introduce a bottleneck such as number of connections to that store, so you now have limitations to think about when scaling.
If you're using libpng to edit an image on your disk you are working with both state and a data store. Yet plenty of applications do this every day without feeling a need to call it over a network.
And if you have to set up and migrate a client/server DBMS for the library, wouldn't you also have to do the exact same steps as if it was running as its own process somewhere the network?
You would not. The HTTP (say) API objects are versioned separately from the DB (or any internal implementation detail) of the service. This is abstraction: your API interfaces don’t change despite any internal changes.
To illuminate this a bit more suppose you do share such a library anyways. Having 2 independent services let’s you upgrade each of them independently. You aren’t forced to update all of them at once.
This is exactly the same as how libraries are supposed to be written. A stable public API so the user doesn't have to bother with internal implementation details.
> Having 2 independent services let’s you upgrade each of them independently. You aren’t forced to update all of them at once.
You aren't forced to do that with libraries either. You don't have to statically link them at compile time, you can load them whenever you please.
> And if you have to set up and migrate a client/server DBMS for the library, wouldn't you also have to do the exact same steps as if it was running as its own process somewhere the network?
No only the team doing it has to. Users of a service only have to worry about the API that it exposes, which should ideally never change in backwards incompatible ways and if it does, you make sure everyone calling the API understands it and migrates and you break backwards compatibility only after everyone has migrated.
As a user the implementation details of other services I don’t work on, don’t concern me.
Note that a service equivalent to libpng should probably not exist and be a library instead. A more reasonable service would be a wishlist service that keeps track of every customers wishlist and sends notifications to customers when products on their wishlist get discounted or are back in stock.
No, it does not mean that at all. Where are you getting this idea from?
> Users of a service only have to worry about the API that it exposes
Users of a library only have to worry about the API that it exposes too. What kind of libraries have you been using?
> Note that a service equivalent to libpng should probably not exist and be a library instead. A more reasonable service would be a wishlist service that keeps track of every customers wishlist and sends notifications to customers when products on their wishlist get discounted or are back in stock.
What's the difference between the two? The only difference that I can see is that the libpng and libwishlist wouldn't waste a network round-trip and lots of serialization and deserialization compared to a png service or whishlist service.
Same problem applies, if your library suddenly wants to access another service (like S3 or some other AWS service).
You can of course maintain a stable API in a library and you should but that restricts the changes you can do to that library. A service I could rewrite in another language without users noticing.
If you're going to rewrite it in another language, what difference does it make if it was a library to begin with? There's no point in making all the upfront work in both the service and all the clients making them talk on the network on the off chance you need to do a rewrite in another language, when you need to rewrite the API layer too. The only thing that would over after the rewrite is the client code. Doing it risks having wasted time doing a network API level on both client and server for nothing, not doing it just risks having to do it when you need it. Doing a rewrite "without users noticing" has no real benefit here.
Where does the redis instance come from? How does the library get access to it?
These are problems the library author cannot solve and end up being problems for the user. Much like monitoring of that redis instance will also become a problem for the user.
Same place as it would if it was a microservice. Someone has to provision it, this is not automatic just because you use microservices.
> How does the library get access to it?
How does the microservice get access to it? A common way seems to be reading a host and port from an environment variable. Nothing is preventing a library from reading from an environment variable too.
> Much like monitoring of that redis instance will also become a problem for the user.
Again, this is something you still need someone to do for the microservice. None of these things are automatically solved by microservices. There are all kinds of automation tools for these things, but they work the same whether you use microservices or not.
This seems quite reasonable In some companies it seems it's more like 1 developer managing 3-12 services :-(
However, the one thing that's not a reason to use microservices and which the article brought up a lot, is to increase resilience. Microservices do not increase resilience: at best they can avoid reducing it.
One of the most compelling reasons for me to use microservices is to limit the potential damage of bad decisions.
In a monolithic application, one developer can have a bad day and write some code that eg. leaks a database object into an area of the program which should deal with business logic only. Or perhaps they try out some new technology that turns out to have problems.
If that manages to get through code review, and is then copied elsewhere by other developers who are just following the example, you can quickly end up in a situation where fixing the problem requires everyone to stop feature work so that a significant rewrite or at least refactoring can take place.
In most cases, the business cannot afford to stop all feature work in this way, and so the problem will persist forever, becoming a permanent drop in productivity. It will also make your developers miserable.
In a microservices architecture, a single mistake can only grow to the size of a single service. In the worst case you have to stop feature work on that one service for a period of time, but work on other areas of the product can continue uninterrupted.
Of course, you could still make a mistake whilst defining the boundaries between services. Luckily those decisions are much less frequent and involve many more people, so there's less chance of a freak bad decision. And even if you do get it wrong, you're only looking at two or three affected services rather than your whole product.
But that's normal and expected on any non-trivial codebase that has been evolving over some amount of time. We call that 'maintenance'. You'll never have 100% of your development resources pushing new features out. And if you do, well... we call that 'incurring technical debt'. You should be investing in this kind of maintenance continuously, because as your product evolves, new problem areas will emerge.
>In most cases, the business cannot afford to stop all feature work in this way, and so the problem will persist forever, becoming a permanent drop in productivity.
Yes.. That's called 'technical debt'. If your business does not invest in maintenance today, then it'll will cost them in the future. This isn't limited to software. If you're maintaining physical infrastructure, same deal. You're always fixing things. Microservices are not a solution to this.
>In a microservices architecture, a single mistake can only grow to the size of a single service.
Not necessarily. Microservices tend to be tightly coupled to other microservices. I've worked on a system that suffered from sporadic 'failure storms' where a failure in one service propagated across the entire system. It took days for us to track down the root cause. Microservices aren't a panacea. In fact, everything is easier with monolithic systems.
Having to stop all feature work is not normal. A healthy business should be able to have a certain proportion of its engineers working on tech debt at all times, it shouldn't have to stop all feature work to address tech debt.
Like you said, part of your dev resources will (should) be dedicated to maintaining your system.
In a microservice, in the absolute worst case (let's say a foundational technology suddenly becomes unsuitable) you have to rebuild just that service.
Work on other services can continue uninterrupted in the meantime.
You're creating a strawman. That's now how software is written. Monolithic software is modularized as well, but the boundaries are behind interfaces and modules, instead of http services. That makes them easier to work with.
>In a microservice, in the absolute worst case (let's say a foundational technology suddenly becomes unsuitable) you have to rebuild just that service.
No. Not necessarily.
>Work on other services can continue uninterrupted in the meantime.
Again, it's a weird hill you're dying on. You can refactor, or rewrite an module without pausing feature development. It happens all.the.time.
I don't know what your professional experience is but clearly it is not indicative of the wider reality.
But you are right in that it makes the code resilience worse, as it makes logic discovery and overview harder.
This scenario is only possible if the project has no automated testing or invests any effort in QA or no one pays any attention whatsoever to whatever is running in prod.
Meanwhile, one of the basic features of CICD is automated rollbacks.
If this sort of problem happens on a project, the problem was not caused by a developer having a bad day. The problem is caused by an entire team enduring a culture that allows problems to fester without anyone doing anything about it.
This is supposedly the #1 technical reason to move to microservices: horizontal scaling + distributed system (in the rare cases the app needs low latency across the world)
The other one is being able to reuse services, whether provided by third parties or that you can download from docker hub/GitHub and run yourself.
> 2) Something needs to be developed by a different team and for whatever reasons you don't want those teams to be too dependent.
This is supposedly the single core reason to adopt microservices architecture. This is what says on the tin. If you want to have multiple teams work independently on separate components, they need to be loosely coupled deployed independently.
This is not brain surgery.
> Everything else is just needlessly/prematurely increasing overhead, deployment friction, etc.
What else is there? The rationale for microservices was from the start organizational, with a lesser scope on operations/technical aspects.
I've been on teams where we can't even reliably deploy code over time to one service. The idea of our team maintaining multiple services is madness. That's not a knock on any individual's technical merits. Deploying code to the cloud is just hard, period—and that's even when talking about a traditional monolith!
I'm glad the pendulum is swinging back. The microservices pattern is useful for the areas in which it's useful…it just so happens that problem space is way smaller than the hype cycle cared to admit a few years ago.
If devops has a good philosophy here it is - if something is hard, you should do it more often. When something is intermittent pain it can be avoided or considered a one-off; if you are say deploying your web app once a week or more people will start to figure out how to automate the manual processes and optimize the slow ones, rewrite troublesome components and come up with strategies for things like incremental database migrations.
More generally - everything is a trade-off, and you shouldn't blindly accept more complexity when you don't need the benefits that it is supposed to provide. But sometimes you need to embrace the complexity of doing things the right way when you are hitting up against limitations of doing it the wrong way (in this example - slow, painful, error-prone deployments)
Microservices are useful when the monolith becomes too complex for local reasoning and management, so you instead compartmentalize things into components with contracts and reason about those, while you are responsible for managing just your own component. You're taking on complexity (and latency, and additional resource utilization) because your system started to hit up against the limits of everyone working within the same project.
If you haven't hit the point where your database is falling over in production because of concurrent loads, or where development is hindered by the infrastructure/dependency or time requirements of doing a local test, or various other problems - you likely don't need to invest in doing microservices.
People draw the boundries out in very odd places when they're following microservices like a religion, but here's an example where it might not suck.
Pull out the auth service.
Should a startup with a single app do this? Probably not.
But as soon as your big enough that someones yelling at you to power multiple services and you have compliance requirements that requires you solve protective monitoring, you probably have a project for a small team of experts, a reason to run this indepedently, and a stable enough API that you can do it well.
Alternatively you could pay Okta & co in blood. :)
This is real: I've worked in projects where purely transformational code was offloaded into a "service". Refactoring it into a library reduced lines of code, computational cost, and code complexity dramatically.
But wouldn't it be cool if there were a framework where the developer didn't have to demarcate where services started and ended? In principle, any pure asynchronous function could be abstracted out to a service. It would be neat if the compiler did that for me and deployment of the application was more like "deploy the cluster" instead of deploying each individual service.
Erlang/Elixir, for JVM Akka, for .NET Akka.NET.
The question was why not eliminate the difference between in-process async call and a service, and that's basically what Erlang and actors do.
> `left-pad.io` is 100% REST-compliant as defined by some guy on Hacker News with maximal opinions and minimal evidence.
Maybe they had something right! It did cause problems when users put something on a tight loop that was actually a remote call but not easy to see.
Anyho... what's old is new!
One at a time, please wait your turn...
Personally I think the Persistency-as-a-library and Consistency-as-a-library architecture is a saner alternative.
I feel this kind architecture of especially appeals to people who like to write only new code instead of understanding existing code. They don't like to read old code to see where new functionality may fit in so they spin up new services.
We are complaining about maintenance of old COBOL code but God be with the poor people who in 20 years will have to maintain the monstrosities we are creating today.
I'm hoping for some major tax changes right before my retirement so I can dust off my COBOL and get phat lewt!
I wonder how many people that have actually written COBOL will be left by 2030? No one mentions it here, but I'm sure there must be programmers deep in the bowels of banks and insurance companies that are learning COBOL even now?
Oh - And RPG... wonder if there's any of that left ?
Imagine living in Cincinnati making $250k maintaining COBOL? Some people are totally fine with that paradigm and will keep the system running as long as it needs to.
Still, I wouldn't be surprised about insane wages for COBOL. You make money at the bleeding edge of tech, or way back on the tail end - scarcity breeds profit :-)
At a previous employer I was responsible for a critical service that was starting to show strain as traffic ramped up. It used Hystrix as the circuit breaker for calls to backend services, including the DB, and at peak times the thread pool would fill and start rejecting additional requests. I was tasked with fixing that.
There's a very simple formula for tuning the number of threads:
> requests per second at peak when healthy × 99th percentile latency in seconds + some breathing room
The catch is that getting good RPS and latency numbers is in a distributed system deployed across three geographically separated datacenters is the opposite of simple. In particular, the legacy of the system meant that we had one write instance of the DB, in one datacenter, meaning that latency was different depending on which DC was the source of the call, so there was no one setting that worked for all instances.
He then argued that our banking solution didn't need logging, because it was so well tested that failure rates would be extremely low.
I'm not making this up.
Making it the entire basis for application abstraction is lunacy, the sort of extremely clever idiocy that can only occur in the tech world.
I can only assume it's because he saw his clients doing it and for a consultant the client is always right. (And I think the client was doing it to indulge their hard to get developers who wanted complexity and the freedom to use their programming language of the week)
It feels like a cowboy industry when this is the 'leadership' we have.
Look at the REST/JSON API debacle that has unfolded over the last two decades. It is trivially obvious that REST is both difficult to implement and largely pointless outside of a hypermedia system, but the thought leaders never said a thing about it.
I don’t have a strong opinion on alternatives. GraphQL seems ok for server to server stuff.
Also, in things like kube, it's possible to schedule jobs next to each other so your TCP connection is essentially going over localhost.
There are many tricks you can do to have this even out perform a single process application.
Which reaffirms my overall opinion that most people writing services have no idea what's a service and what it encapsulates (it encapsulates its own state, for example, it's basically distributed OOP).
BTW, remember when "your service should be 100 lines of code top" was considered a best practice?
Why is it so hard for most to resist this hype nonsense wave when its oncoming and it's so hard to resist the anti-hype wave that inevitably follows it? Because hype and anti-hype are simply the oscillation of a sea of empty minds in look for a solution to a problem they don't understand.
I've been writing service oriented apps for decades. In my "world" nothing has changed.
Because then how will one do career development? We are building now for a client and the client wants it, I can feel it in my bones that what we are building is far more complex and will be difficult to maintain. I am no expert in microservices but this hunch is just coming from common sense.
But the really bad moment is when the developers themselves make those bad choices entirely on their own.
The client is to blame. The client let go of all the people who knew the technical side of the product and has hired an architect to re-architect everything. And the client is in a hurry.
Once you have thousands of engineers, then you either need extreme discipline and a huge team maintaining the "devops" pipeline that everything goes through, or, it's basically everyone for themselves and a "devops" team trying to help others setting standards, best practices and whatnot.
Now as with many things that nothingburger got suddenly overhyped and ended up being shoved into every hole disregarding of any technical rationale.
There are so many developers who think microservices are the answer to every ill in the world it is infuriating. It’s like one of those ‘nosql is webscale we should use it’ conversions.
Massive optimism detected at the part of "to realize later".
Also, if you have "edge case" with high (close to 1 but lower than 1) probability of successful outcome, with more and more tries probability of all outcomes being successful moves to zero (multiplication of lower than 1 values) and probability of at least single failed outcome goes to 1 (addition of lower than 1 values). You still have to handle low probability edge cases.
Your organization and product goals may be totally different and not need that level of testing. But if it is needed, you can do it. It's up to your business to decide how much to invest in it.
> ”At the heart of our Engineering Strategy is a rather simple document. A manifesto if you’d like. We call it Core Engineering Principles.”
> Lists 7 fairly complex rules
> ”There are about 30 more touching stacks, architecture, and org structure.”
I’m dying. This could only be written with a straight face by a microservices enthusiast.
Settling on 3 programming languages is a good compromise between "one language to rule them all", and "anything goes". Being limited a single language's talent pool can be quite limiting, as is only having one language to problems in. Having no limits on language choice also has obvious drawbacks.
The "7 fairly complex rules" are not _that_ complex considering they are rules for trying to tame a domain that is inherently complex. If you try and scale any software engineering team much beyond 30 engineers you will need to have similar rules regardless of whether you have a majestic monolith, microservices, or some mix of the two.
Strong service boundaries helped us, they didn't hold us back.
Why not both? You can set things up so your go-to-def can understand your API calls and head to the right place. This is very easy to do with a monorepo + protobuf setup.
> How much does it cost to spin 200 services in a cloud provider? Can you do it? Can you also spin up the infrastructure needed to run them?
Assuming 256mb RAM/service you're still well within 1 machine territory. Once you get above 1 machine territory you can set things up so that you can:
1. Build really good integration testing tooling so that devs don't really need to interface with all services. In a test spin up everything you need + deps, run an API call, tear everything down. This can be cached if your build system does that. You can run into issues if you have situations where 1 API call hits every service but if you've done that you've already messed up. In those cases the best you can do is mock a step in the chain there, run a few test hits the entire chain before release, but then have devs run against the mock in their integration tests.
2. Hybrid environments. You run a dev cluster that has all of your basic infrastructure that doesn't change much and provide a way for developers to launch new tasks that don't get routed to unless the driver has a feature flag flipped. Essentially you have a "dev" cluster that is continuously delivered from your master repo, each developer has the ability to launch new tasks in this cluster, and they can say "all traffic from alice for FoobarService should go to `{namespace=bob,service=Foobar}`.
> As you can imagine, end-to-end tests have similar problems to development environments. Before, it was relatively easy to create a new development environment using virtual machines or containers. It was also fairly simple to create a test suite using Selenium to go through business flows and assert they were working before deploying a new version.
Why is it not simple anymore? I've implemented this at more than one company.
> Aside from being an obvious single-point-of-failure, defeating some of the service-oriented architecture's principles, there's more. Do you create a user per service? Do you have fine-grained permissions so service A can only read or write from specific tables? What if someone removes an index unintentionally? How do we know how many services are using different tables? What about scaling?
Come up with a convention for your company and stick to it. If you can automate it that's better. If you build some way for your task you are running to know "who" it is they can inject that information into other libraries. For example you can inject the following environment variables into a container:
FOOBAR_DB_ABCD=pg
FOOBAR_DB_ABCD_PASSWORD=...
FOOBAR_DB_ABCD_HOST=...
FOOBAR_DB_ABCD_PORT=...
You can then have some library you write expose a `OpenDatabase("abcd")` that connects in and injects everything. A security operator can then provision accounts and everything transparently. If you generate those env vars from some automated config management tool you don't even have to see the passwords.> Instead of having a monolith getting all of the traffic, now you have a home-made Spring Boot service getting all of it! What could go wrong? Engineers quickly realize this is a mistake, but as there are many customizations, sometimes they cannot substitute this piece for stateless, scale-friendly ones.
I don't think this is a single point of failure. This is a single point of failure for a specific subset of your infrastructure and at that it should be a very simple (mostly-pass-through) component. If your mobile gateway dies, your backend one shouldn't. If all of your API gateways die then your integrations to third parties that are required for legal compliance should stay up, etc.
> I've seen teams using circuit breakers and then increase the timeouts of an HTTP call to a service downstream.
You should always decrease timeouts for operations if you're attempting to retry calls. You can also use a load balancer that already knows the liveness state of all of your instances.
Aside from this bit I mostly agree with this section about timeouts, retries, etcs is actually correct. If you are tackling a single problem that's simple don't break it into a distributed system. If you are saying "I want to have X but we should implement a Y" where Y is a completely different thing that doesn't need to talk to X directly, then why not implement it in a separate binary? There's no reason they can't share code to make the burden of operation low.
If you have never gone through a significant dependency upgrade on a large monolithic codebase, you might not appreciate the value of microservice architecture. I’ve been at companies where the number one technical achievement for a tech organization for an entire calendar year was ‘successfully moved the app from .net 2 to .net 4’. Have you ever gone through the process of integrating an acquired company’s monolith into an acquiring company’s monolith? It’s futile. With a microservice architecture though I’ve seen acquisitions integrate significant systems (things like billing and auth) in a matter of weeks where monolith projects dragged on for years.
Sure, there’s no silver bullets. But there are a lot of problems with monoliths which microservices eliminate. Not without trade offs, naturally.
Though I think mostly that is because the discussion there has been well trodden already. Microservice Architectures are relatively new, less completely explored, and less familiar to many.
There are new and interesting mistakes to make (the biggest two being to use this model where it is not at all appropriate, or cargo-culting and not really understanding what its advantages are so not actually taking advantage of them) where most of the "interesting" problems to be had in the monolithic world have been well documented for quite some time (which is why new models pop up, to try remove the potential for these problems in some, many, or even most, cases (anyone who says method X is better in all cases should be treated with deep suspicion)).
Disorganized management won't magically get fixed if the root problems of ASAP, priority, techinal debt, etc features aren't fixed.
The fact that the organization was trying to combine two large monolithic codebase into one is an obvious smell that technical decisions are made by someone non-technical.
For better or worse, nontechnical people sometimes make decisions. One of the jobs of a software architecture is to be robust in accommodating the consequences of them. Monolithic architectures frequently aren’t.
We already had services and service oriented architecture. Finding the correct size and establishing boundaries between services is a huge part of a correct architecture.
Building a monolith was a first step, then companies were branching out different parts in other services, when and if needed (eg. when doing acquisitions, as you mentioned).
Microservices just pushed companies to write services that were too small (I still remembers thought leaders praising 20 lines microservices) and declare dinosaurs everyone doing monoliths.
Your ‘successfully moved the app from .net 2 to .net 4’ becomes `updated all versions of the logger in all microservices`.
I have worked in both. Until you realize that both worlds have their pros and cons, just stop.