Goodbye Microservices: From 100s of problem children to 1 superstar
segment.com
segment.com
>> A huge point of frustration was that a single broken test caused tests to fail across all destinations. When we wanted to deploy a change, we had to spend time fixing the broken test even if the changes had nothing to do with the initial change. In response to this problem, it was decided to break out the code for each destination into their own repos
They also introduced tech debt and did not responsibly address it. The result was entirely predictable, and they ended up paying back this debt anyway when they switched back to a monorepo.
>> When pressed for time, engineers would only include the updated versions of these libraries on a single destination’s codebase... Eventually, all of them were using different versions of these shared libraries.
To summarize, it seems like they made some mistakes, microed their services in a knee-jerk attempt to alleviate the symptoms of the mistakes, realized microservices didn't fix their mistakes, finally addressed the mistakes, then wrote a blog post about microservices.
That seems... appropriate?
This is the general problem with the microservices bandwagon: Most of the people touting it have no idea when or why it's appropriate. I once had a newly hired director of engineering, two weeks into a very complicated codebase (which he spent nearly zero time looking at), ask me "Hey there's products here! How about a products microservice?" He was an idiot that didn't last another two months, but not before I (and the rest of the senior eng staff) quit.
I'm fully prepared to upvote more stories with the outline of Microservices were sold as the answer! But they weren't.
It's almost like engineering techniques aren't magic pixie dust that you can sprinkle over your project and get amazing results...
Micro services is an extremely powerful pattern which solves a bazillion critical issues most important ones being:
* separation of concerns
* ability of different teams maintain, develop and reiterate on different subsystems independently from each other
* loose coupling of the subsystems
Do you have auth server that your API accesses using auth.your.internal.name which does not share its code base with the API? You have a micro service.
Do you have a profile service that is responsible for the extra goodies on a profile of a user that the rest of the API business logic does not care about?
You have a micro service. Do you spin up some messaging layer in a cloud that knows a few things about the API but really is only concerned with passing messages around? You have a micro service.
The alternative is that you have a single code base and a single app that starts with ENV_RUNMODE=messaging or ENV_RUNMODE=API or ENV_RUNMODE=website and ENV_RUNMODE=auth ( except in the case of auth it only implements creation/changes of the new entries and a change of passwords but not validation as the validation is done by the code in any ENV_RUNMODE by accessing the auth database directly with read-write privilege and no one ever implemented deletion of the entries from the authentication database. Actually, even that would be an good step - there's no auth database because that would require knowing the mode we are running in and managing multiple sets of credentials so instead it is simply another table in a single database that stores everything )
That is the alternative to micro services. So I would argue that unless Segment has that kind of architecture it does not have a monolith. It implements a sane micro services pattern.
Should the engineering be lead by a blind squirrel that once managed to find a nut, in a winter, three years ago, the sane micro services pattern would be micro serviced even more -- I call it nanoservice pattern aka LeftPad as a service. We aren't seeing it much yet but as Go becomes bigger and bigger player in shops without excellent engineering leadership I expect to see it more and more due to Go giving developers tools to gRPC between processes.
This made me lol
They become a bucket of clichés and abstract terms. Clichéd descriptions of problems you're encountering, like deployments being hard. Clichéd descriptions of the solutions. This let's everyone in on the debate, whether they actually understand anything real to a useful degree or not. It's a lot easier to have opinions about something using agile or microservice standard terms, than using your own words. I've seen heated debates between people who would not be able to articulate any part of the debate without these clichés, they have no idea what they are actually debating.
For a case in point, if this article described architecture A, B & C without mentioning microservices, monoliths and their associated terms... (1) Far fewer people would have read it or had an opinion about it. (2) The people who do, will be the ones that actually had similar experiences and can relate or disagree in their own words/thoughts.
What makes these quasi-ideological in my view is how things are contrasted, generally dichotomously. Agile Vs Waterfall. Microservices Vs Monolithic Architecture. This mentally limits the field of possibilities, of thought.
So sure, it's very possible that architecture style is/was totally besides the point. Dropping the labels of microservices architecture frees you up to (1) think in your own terms and (2) focus on the problems themselves, not the clichéd abstract version of the problem.
Basically, microservice architecture can be great. Agile HR policies can be fine. Just... don't call them that, and don't read past the first few paragraphs.
The problem, as your identify, is that once a pattern has been identified people too easily line up behind it and denigrate the "contrasting" pattern. The abstraction becomes opaque. We're used to simplistic narratives of good vs evil, my team vs your team, etc. and our tendency to embrace these narratives leads to dumb pointless conversations driven more be ideology than any desire to find truth.
I just think there can be downsides to them. These are theories as well as terms and they become parts of our worldview, even identity. This can engage our selective reasoning, cognitive biases and our "defend the worldview!" mechanisms in general. At some point, it's time for new words.
Glad people seem ok with this. I've expressed similar views before (perhaps overstating things) with fairly negative responses. I think part of it might be language nuance. The term "ideology" carries less baggage in Europe, where "idealist" is what politicians hope to be perceived as while "ideologue" is a common political insult statesside, meaning blinded and fanatic.
Concepts like microservices synthesize a bunch of tradeoffs and patterns that have been worked on for decades. They’re boiled down to an architecture fad, but have applicability in many contexts if you understand them.
Similarly with Agile, it synthesizes a lot of what we know about planning under uncertainty, continuous learning, feedback, flow, etc. But it’s often repackaged into cliche tepid forms by charlatans to sell consulting deals or Scrum black belts.
Alan Kay called this one out in an old interview: https://queue.acm.org/detail.cfm?id=1039523
“computing spread out much, much faster than educating unsophisticated people can happen. In the last 25 years or so, we actually got something like a pop culture, similar to what happened when television came on the scene and some of its inventors thought it would be a way of getting Shakespeare to the masses. But they forgot that you have to be more sophisticated and have more perspective to understand Shakespeare. What television was able to do was to capture people as they were.
So I think the lack of a real computer science today, and the lack of real software engineering today, is partly due to this pop culture.”
I will take issue with one thing though... Shakespeare's plays were for something like a television audience, the mass market. The cheap seats cost about as much as a pint or two of ale. A lot of the audience would have been the illiterate, manual labouring type. They watched the same plays as the classy aristocrats in their box seats. It was a wide audience.
Shakespeare's stories had scandal and swordfighting, to go along with the deeper themes.
A lot of the best stuff is like that. I reckon GRRM a great novelist, personally, with deep contribution to the art. Everyone loves game of thrones. It's a politically driven story with thoughtful bits about gender, and class and about society. But, its not stingy on tits and incest, dragons and duels.
The one caveat was that Shakespeare's audience were all city slickers, and that probably made them all worldlier than the average Englishman who lived in a rural hovel, spoke dialect and rarely left his village.
What is an elitist pursuit is not really Shakespeare, it's watching 450 year old plays.
We really like to think in silos, categorize everything to make them feel familiar and approachable. Which is useful, but sometimes we need to shake them off so we can actually see the problems.
After a while... it's like the cliché about taxi drivers investing in startups... Sign it's time to get out. When people I know have no idea start talking about the awesomeness of some abstract methodology... I'm out.
You try to remove the critique from microservices, but for me these issues are actually good arguments against microservices. It's hard to do right.
This is correct; I'd argue doing microservices right is even harder than doing a monolith right (like, keeping the code base clean).
Then there have been quite a few warnings to not use shared code in microservices.
Even so, imagine the chaos if frequently engineers/devs need to add code to one lib(the shared one), wait for PR approval, then use that new version in a different lib to implement the actual change that was needed? Thats seems to be introducing a direct delay into getting anything productively done...
In reality this process has absolutely nothing to do with the structure of the organisation. It's true purpose is to shuffle out people who are in positions where they're performing poorly, and move in new people. It just provides cover (It's not your fault, it's an organisational change).
This is exactly the same, they couldn't say "You've solved this problem badly, go spend 6 months doing it properly". So instead they say they need a new paradigm to organise how they build their solution. In the process of that they get to spend all the time they need fixing the bad code, but it's not because it's bad code, it's because the paradigm is wrong.
The problem is the same problem with the organisational structure- if you don't realise the real purpose, and buy into the cover you end up not addressing the issue. You end up with a shit manager managing a horizontal and then managing a verticle, then managing a horizontal. You end up with a bad monolithic-service instead of bad micro-services.
We went through this painful period. Kept at it devoting a rotating pair to proactively address issues. Eventually it stabilized, but the real solution was to better decouple services and have them perform with more 9s of reliable latency. Microservices are hard when done improperly and there doesn't seem to be a short path to learning how to make them with good boundaries and low coupling.
>>To summarize, it seems like they made some mistakes, microed their services in a knee-jerk attempt to alleviate the symptoms of the mistakes, realized microservices didn't fix their mistakes, finally addressed the mistakes, then wrote a blog post about microservices.
I read the article a few days ago and was struck by what a poor idea it was to take a hundred or so functions that do about the same thing and to break them up into a hundred or so compilation and deployment units.
If that's not a micro-service anti-pattern, I don't know what is!
In particular, as to their original problem, the shared library seems to be the main source of pain and that isn't technically solved by a monolith, along with not following the basic rule of services "put together first, split later".
I feel prematurely splitting services like that is bound to have issues unless they have 100 developers for 100 services.
The claim of "1 superstar" is misleading too, this service doesn't include the logic for their API, Admin, Billing, User storage etc etc, it's still a service, one of a few that make up Segment in totality.
Reading about their setup and comparing with some truly large scale services I work with, I'm left with the idea that Segment's service is roughly the size of one microservice on our end.
Perhaps the takeaway is don't go overboard with fragmenting services when they conceptually fulfill the same business role. And regardless of the architecture of the system, there are hard state problems to deal with in association with service availability.
Some rules of thumb I just came up with:
Number of repos should not exceed number of developers.
Number of tests divided by number of developers should be at least 100.
Number of lines of code divided by number of repos should be at least 5000.
Your tests should not run faster than the time it takes to read this sentence.
A single person should not be able to memorize the entire contents of a single repo, unless that person is Rain Man.
I'd say you've never had good tests.
I have a test-suite for a bunch of my frameworks that dates to the mid 90s, with tests added regularly with new functionality.
It currently takes 4 seconds total for 6 separate frameworks and 1000 individual tests. Which is actually a bit slower than it should be, it used to take around 1-2 seconds, so might have to dig a little to see what's up.
With tests this fast, they become a fixed part of the build-process, so every build runs the tests, and a test failure is essentially treated the same as a compiler error: the project fails to build.
The difference goes beyond quantitative to qualitative, and hard to communicate. Testing becomes much less of a distinct activity but simple an inextricable part of writing code.
So I would posit:
Your tests should not run slower than the time it takes to read this sentence.
Either way, the idea of considering slow tests a feature was novel to me.
If you're using a tape drive or SD cards, sure. But even a 10 year old 5400RPM on an IDE connection should be able to satisfy your tests' requirements in a few seconds or less.
I suspect your tests are just as monolithic as you think microservices shouldn't be. Break them down into smaller pieces. If it's hard to do that, then redesign your software to be more easily testable. Learn when and how to provide static data with abstractions that don't let your software know that the data is static. Or, if you're too busy, then hire a dedicated test engineer. No, not the manual testing kind of engineer. The kind of engineer who actually writes tests all day, has written thousands (or hundreds of thousands) of individual tests during their career. And listen to them about any sort of design decisions.
If you need to access a database in your tests you're probably doing it wrong. Build a mock-up of your database accessor API to provide static data, or build a local database dedicated for testing.
I tend to use my integration tests also as characterization tests that verify the simulator/test-double I use for any external systems within my unit tests.
See also: the testing pyramid[1] and "integrated tests are a scam"[2], which is a tad click-bait, but actually quite good.
It's not ridiculous. It's good.
I work on an analysis pipeline with thousands of individual tests across a half dozen software programs. Running all of the tests takes just a few seconds. They run in under a second if I run tests in parallel.
If your tests don't run that fast then I suggest you start making them that fast.
I'd be willing to bet that if you learned (or hired someone with the knowledge of) how to optimize your code, you could get some astounding performance increases in your product.
unless they have 100 developers for 100 services.
That cure is worse than the disease. Every service works differently and 80% of them are just wrong, and there’s nothing you can do because Tim owns that bit.Tom (real guy) was too busy all the time to do anything other than the 80/20 rule. He was too busy because he didn't share. So of course he was a fixture of the company...
This is a business decision, reality has no influence here.
That's why they invented the term "Business reality".
Now all the developers are going to the CTO or CEO and undermining the other developers, trying to persuade the CTO that so-and-so's code is shit.
It's not just code coverage that matters. It's the code path selection that matters. If you have a ton of branches and you've evaluated all of them once then yeah you sure might have 100% "coverage". But you have 0% path selection coverage since a single invokation of your API might choose true branch on one statement, false branch on another statement, and a second invokation might choose false branch on the first and true on the second.
While the code was 100% tested, the scenarios were not. What happens if you have true/true or false/false? That's not tested.
There's a term for this but I forgot what it is and don't care to go spelunking to find it.
Happy path?
sqlite calls this "branch coverage"
https://www.sqlite.org/testing.html#statement_versus_branch_...
I felt this article is more about how to use microservices right way vs butchering the idea. It is not right to characterize this as microservices vs monolith service. Initial version of their attempt went too far by spinning up a service for each destination. This is taking microservices to extreme which caused organizational and maintenance issue once number of destinations increased. I am surprised they did not foresee this.
The final solution is also microservice architecture with a better separation of concerns/functionalities. One service for managing in bound queue of events and other service for interacting with all destinations.
Agreed. I treat services like an amoeba. Let your monolith grow until you see the obvious split points. The first one I typically see is authentication, but YMMV.
Notice I also do not say 'microservices'. I don't care about micro as much as functional grouping.
Changing one "shared library" shouldn't mean deploying 140 services immediately.
They had one service to begin with forked for each destination. Of course that was a nightmare to maintain!
They also just might have had too many repos.
Is this rule mentioned or discussed somewhere? A quick google search links to a bunch of dating suggestions about splitting the bill. Searching for the basic rule of services "put together first, split later" reveals nothing useful.
I've worked with microservices a lot. It's a never-ending nightmare. You push data consistency concerns out of the database and between service boundaries.
Fanning out one big service in parallel with a matching scalable DB is by far the most sane way to build things.
Sing that from the rooftops. That is exactly my observation as well. All the vanilla "track some resource"-style webapps I've worked on were never designed to cope with a consistency boundary that spans across service boundaries. Turning a monolith into distributed services is hard for that reason - you have to redesign your data access to ensure that consistency boundaries don't span across multiple services. If you don't do that, then you have to learn to cope with eventual consistency; in my experience, most people just don't think that way. I know I have trouble with it. Surely I'm not the only one.
I've never quite understood why people think that taking software modules and separating them by a slow, unreliable network connection with tedious hand-wired REST processing should somehow make an architecture better. I think it's one of those things that gives the illusion of productivity - "I did all this work, and now I have left-pad-as-a-service running! Look at the little green status light on the cool dashboard we spent the last couple months building!"
Programmers get excited about little happily running services, these are "real" to them. Customers couldn't care less, except that it now takes far longer to implement features that cross multiple services - which, if you've decomposed your services zealously enough, is pretty much all of them.
That is not completely on the developer, either. Pre 4.0 Mongodb, for example, does not do transactions. On the other hand, I've seen some pretty flagrant disregard for it just because there are not atomicity guarantees.
Microservices makes reasoning on that harder.
I'd argue that's on the developer, if he was the one to choose a database that doesn't support transactions, and then didn't implement application-level transactions (which is very hard to do correctly).
If two microservices have to share databases, they shouldn't be microservices.
One microservice should have write access to one database and preferably, all read requests run through that microservice for exactly the reason you mentioned.
>I've never quite understood why people think that taking software modules and separating them by a slow, unreliable network connection with tedious hand-wired REST processing should somehow make an architecture better.
If you're running microservices between regions and communicating with each other outside of the network it is living in, you're probably doing it wrong.
Microservices shouldn't have to incur the cost of going from SF to China and back. If one lives in SF, all should and you can co-locate the entire ecosystem (+1 for "only huge companies with big requirements should do microservices")
>ustomers couldn't care less, except that it now takes far longer to implement features that cross multiple services - which, if you've decomposed your services zealously enough, is pretty much all of them.
Again, that is an example of microservices gone wrong. You'll have the same amount of changes even in a monolith and I'd argue adding new features is safer in microservices (No worries of causing side effects, etc).
I will give you +1 on that anyway because I designed a "microservice" that ended up being 3 microservices because of dumb requirements. It probably could've been a monolith quite happily.
Problem domains (along with organizational structures) inherently create natural architectural boundaries... certain bits of data, computation, transactional logic, and programming skill just naturally "clump" together. Microservices ignore this natural order. The main driving architectural principle seems to be "I'm having trouble with my event-driven dynamically-typed metaprogrammed ball-of-mud, so we need more services!".
Which, strangely enough, ends up looking almost the same thing as running services on a microkernel.
The "natural" order is very often bad for reliability, speed and efficiency. It forces the "critical path" of a requests to jump through a number of different services.
In well built SOA you often find that the problem space is segmented by criticality and failure domains, not by logical function.
Then before you know it there are a dozen more shared libraries and you have the distributed monolith.
Then either have to Stand up every micro service every time integration test or make changes and hope for the best.
You can't have consistent microservices without distributed transactions. If a service gets called, and inside that call, it calls 3 others, you need to have a roll back mechanism that handles any of them failing in any order.
If you write to the first service and the second two fail, you need to write a second "undo" call to keep consistent.
Worse, this "undo state" needs to be kept transactionally consistent in case it's your service that dies after the first call.
In reality, nobody does this, so they're always one service crash away from the whole system corrupting the hell out of itself. Since the state is distributed, good luck making everything right again.
Microservices are insane. Nobody that knows database concepts well should go near them
I'd venture a guess that most applications have all sorts of race conditions that could cause data corruption. The fact of the matter is that almost nobody notices or even cares.
Aside from that, I've found that even in non-mission-critical scenarios ("it's just porn!") it's incredibly convenient to have a limited number of states the system can be in. It makes debugging easier and reduces the number of edge cases ("why is this null??") you have to handle.
I think you'd be surprised/alarmed at how little transactions actually get used in the software world. Not just on small systems where it doesn't matter but I've seen a complete absence of them in big financial ones handling billions of dollars worth transactions (the real world kind) a day. Some senior, highly paid people even defend this practice for performance reasons because they don't realize the performance cost of implicit transactions. And this is just the in process stuff where transactions are totally feasible, it get's even worse when you look at how much is moved around via csv files to FTP and excel sheets attached to emails. I've spent the last 2 weeks being paid to fix data consistency issues that should never have been issues in the first place.
Maybe when we're teaching database theory we shouldn't start at select/join but at begin transaction/commit/rollback?
However, I would argue that transactions are overkill. What's the worst case scenario if I book a ride for Grab and my request gets corrupted? I'm guessing I'll see an error message and I'll have to re-request my ride.
No, you just need loosely-coupled services, where inconsistency in this circumstance doesn't manifest as a problematic end-state.
If your business model is to be cheap with high volume sales, then corrupting say 1 out of 10,000 customer transactions may be worth it. If you give customers a good price and/or they have no viable alternative, you can live with such hiccups, and the shortcuts/sacrifices may even make the total system cheaper. You are like a veterinarian instead of a doctor: you can take shortcuts and bork up an occasional spleen without getting your pants sued off. But most domains are NOT like that.
No architecture principle will help you if you design things the wrong way.
The point of Micro services architecture is to design large working systems, from smaller working systems.
(Further below, I'll go into in which contexts I'd agree with your assessment and why. But for now the other side of the coin.)
In the real world, current-day, why do many enterprises and IT departments and SME shops go for µservice designs, even though they're not multimillion-user-scale? Not for Google/Netflix/Facebook scale, not (primarily/openly) for hipness, but they do like among other reasons:
- that µs auto-forces certain level of discipline in areas that would be harder-to-enforce/easier-to-preempt by devs in other approaches --- modularity is auto-enforced, separation of concerns, separation of interfaces and implementations, or what some call (applicably-or-not) "unix philosophy"
- they can evolve the building blocks of systems less disruptively (keep interfaces, change underlyings), swap out parts, rewrites, plug in new features to the system etc
- allows for bring-your-own-language/tech-stack (thx to containers + wire-interop) which for one brings insights over time as to which techs win for which areas, but also attracts & helps retain talent, and again allows evolving the system with ongoing developments rather than letting the monolith degrade into legacy because things out there change faster than it could be rewritten
I'd prefer your approach for intimately small teams though. Should be much more productive. If you sit 3-5 equally talented, same-tech/stack and superbly-proficient-in-it devs in a garage/basement/lab for a few months, they'll probably achieve much more & more productively if they forgoe all the modern µservices / dev-ops byzantine-rabbithole-labyrinths and churn out their packages / modules together in an intimate tight fast-paced co-located self-reinforcing collab flow. No contest!
Just doesn't exist often in the wild, where either remote distributed web-dev teams or dispersed enterprise IT departments needing to "integrate", rule the roost.
(Update/edit: I'm mostly describing current beliefs and hopes "out there", not that they'll magically hold true even for the most inept of teams at-the-end-of-the-day! We all know: people easily can, and many will, 'screw up somewhat' or even fail in any architecture, any language, any methodology..)
Do they actually force a discipline? Do people actually find swapping languages easier with RPC/messaging than other ffi tooling? And do they really attract talent?!
You make some amazing claims that I have seen no evidence of, and would love to see it.
I'm just relaying what I hear from real teams out there, not intending to sell the architecture. So these are the beliefs I find on the ground, how honest and how based-in-reality they are are harder to tell and only slowly over time at any one individual team.
A lot of this is indeed about hiring though, I feel, at least as regards the enterprise spheres. Whether you can as a hire really in-effect "bring your own language" or not remains to be seen, but by deciding on µs architecture for in-house you can certainly more credibly make that pitch to applicants, don't you think?
Remember, there are many teams that have suffered for years-to-decades from the shortcomings and pitfalls of (their effectively own interpretation of / approach to) "monoliths" and so they're naturally eagerly "all ears". Maybe they "did it wrong" with monoliths (or waterfall), and maybe they'll again "do it wrong" (as far as outsiders/gurus/pundits/coachsultants assess) with µs (or agile) today or tomorrow. The latter possibility/danger doesn't change the former certainties/realities =)
Regardless of whether you are a monolith or a large zoo of services, it works when the team is rigorous about separation of concerns and carefully testing both the happy path and the failure modes.
Where I've seen monoliths fail, it was developers not being rigorous/conscientious/intentional enough at the module boundaries. With microservices... same thing.
The disadvantage is obviously that creating such's a 'perfect architecture' is hard to do because of different concerns by different parties within the company/organisation.
I think you get at two very good points. One is that realistically you will never have enough time to actually get it really right. The other is that once you take real-world tradeoffs into account, you'll have to make compromises that make things messier.
But I'd respond that most organizations I see leave a lot of room for improvement on the table before time/tradeoff limitations really become the limiting factor. I've seen architects unable to resolve arguments, engineers getting distracted by sexy technologies/methodologies (microservices), bad requirements gathering, business team originated feature thrashing, technical decisions with obvious anticipated problems...
The only teams I had to spend time on were the ones which were on a common DB before we moved off of it.
The legacy concerns I don’t see being true, as it’s mainly a requirements/documentation problem and you can achieve the same effect with feature toggles.
You easier to debug end-to-end tests of a microservice architecture that monolith? That's not my experience. How do you manage to put side by side all the events when they are in dozen of files?
I use Serilog for structured logging. Depending on the log destination, your logs are either stored in an RDMS (I wouldn’t recommend it) or created as JSON with name value pairs that can be sent directly to a JSON data store like ElasticSearch or Mongo where you can do adhoc queries.
https://stackify.com/what-is-structured-logging-and-why-deve...
IOrderService -> OrderService in production.
IOrderService ->FakeOrderService when testing.
But if you are testing an artifact, why isn’t the artifact testing part of your CI process? What you want to do is no more or less an anti pattern than swapping out mock services to test a microservice.
I’m assuming the use of a service discovery tool to determine what gets run. Either way, you could screw it up by it being misconfigured.
>But if you are testing an artifact, why isn’t the artifact testing part of your CI process?
It is and it shall be part of the CI process. Commit gets assigned build number in tag, artifact gets the version and build number in it's name and metadata, deployment to CI environment is performed, tests are executed against specific artifact, so every time you deploy to production you have a proof, that the exact binary that is being deployed has been verified in its production configuration.
>I’m assuming the use of a service discovery tool to determine what gets run.
Service discovery is irrelevant to this problem. Substitution of mock can be done with or without it.
If you are testing a single microservice and don’t want to test the dependent microservice - if you are trying to do a unit test and not an integration test, you are going to run against mock services.
If you are testing a monolith you are going to create separate test assemblies/modules that call your subject under test with mock dependencies.
They are both going to be part of your CI process then and either way you aren’t going to publish the artifacts until the tests pass.
Your deployment pipeline either way would be some type of deployment pipeline with some combination of manual and automated approvals with the same artifacts.
The whole discussion about which is easier is moot.
Edit: I just realized why this conversation is going sideways. Your initial assumptions were incorrect.
you may want to test only A: with monolithic architecture you'll have to produce another build of the application, that contains mock of B (or you need something like OSGi for runtime module discovery).
That’s not how modern testing is done.
https://www.developerhandbook.com/unit-testing/writing-unit-...
> Your initial assumptions were incorrect. With nearly 20 years of engineering and management experience, I know very well how modern testing is done. :)
What is an app at the system boundaries if not a piece of code with dependencies?
If you have a microservice - FooService that calls BarService. The "system boundary" you are trying to test is FooService using a fake BarService. I'm assuming that you're calling FooService via HTTP using a test runner like Newman and test results.
In a monolithic application you have class FooModule that depends on BarModule that implements IBarModule. In your production application you use create FooModule:
var x = FooModule(new BarModule)
y = x.Baz(5);
In your Unit tests, you create your FooModuleL
var x = FooModule(new FakeBarModule) actual= x.Baz(5) Assert.AreEqual(10,actual)
And run your tests with a runner like NUnit.
There is no functional difference.
Of course FooModule can be at whatever level of the stack you are trying to test - even the Controller.
For a certain class of applications and organizational constraints, I also would prefer it. But it requires a much tighter alignment of implementation than microservices (e.g., you can't just release a new version of a component, you always have to release the whole application).
https://www.developerhandbook.com/unit-testing/writing-unit-...
It works similarly in almost every language.
For a certain class of applications and organizational constraints, I also would prefer it. But it requires a much tighter alignment of implementation than microservices (e.g., you can't just release a new version of a component, you always have to release the whole application).
Why is that an issue with modern CI/CD tools? It’s easier to just press a button and have your application go to all of your servers based on a deployment group.
With a monolith, with a statically typed language, refactoring becomes a whole lot easier. You can easily tell which classes are being used, do globally guaranteed safe renames, and when your refactoring breaks something, you know st compile time or with the correct tooling even before you compile.
It's not so much about the deployment process itself (I agree with you that this can be easily automated), but rather about the deployment granularity. In a large system, your features (provided by either components or by independent microservices) usually have very different SLAs. For example, credit card transactions need to work 24x7, but generating the monthly account statement for these credit cards is not time-critical. Now suppose one of the changes in a less critical component requires a database migration which will take a minute. With separate microservices and databases, you could just pause that microservice. With one application and one database, all teams need to be aware of the highest SLA requirements when doing their respective deployments, and design for it. It is certainly doable, but requires a higher level of alignment between the development teams.
I agree with your remark about refactoring. In addition, when doing a refactoring in a microservice, you always need a migration strategy, because you can't switch all your microservices to the refactored version at once.
That’s easily accomplished with a Blue-Green deployment. As far as the database, you’re going to usually have a replication set up anyway. So your data is going to live in multiple databases anyway.
Once you are comfortable that your “blue” environment is good, you can slowly start moving traffic over. I know you can gradually move x% of traffic every y hours with AWS. I am assuming on prem load balancers can do something similar.
If your database is a cluster, then it is still conceptually one database with one schema. You can't migrate one node of your cluster to a new schema version and then move your traffic to it.
If you have real replicas, then still all writes need to go to the same instance (cf. my example of credit card transactions). So I also don't understand how your migration strategy would look like.
blue-green is great for stateless stuff, but I fail to see how to apply it to a datastore.
https://www.goodreads.com/book/show/36308520-building-evolut...
1. Is auto-enforced modularity, separation of concerns, etc actually better than enforcing these things through development practices like code review? Why are you paying people a 6 figure salary if they can't write modular software?
2. Is the flexibility you gain from this loose coupling worth the additional costs and overhead you incur? And is it really more flexible than a modular system in the first place? And how does their flexibility differ? With an API boundary breaking changes are often not an option. In a modular codebase they can easily be made in a single commit as requirements change.
3. Is bring-your-own-language actually a good idea for most businesses? Is there a net benefit for most people beyond attracting and retaining talent? What about the ability to move developers across teams and between different business functions? Having many different tech stacks is going to increase the cost of doing this.
I do see the appeal of some of these things, but IMO the pros outweigh the cons for a smaller number of businesses than you've mentioned. And the above is only a small sample of that. Most things are just more difficult with a distributed system. It's going to depend on the problem space of course, but most backend web software could easily be written in a single language in a single codebase, and beyond that modularization via libraries can solve a lot of the same problems as microservices. I'm very skeptical of the idea that microservices are somehow going to improve reliability or development speed unless you have a large team.
I take your point, but it saddens me that there aren't better ways of achieving this modularity nowadays.
You really do have to modularize. In some languages, you can even use separate compilation units for separate modules to enforce the separation.
You can do all of that but get simultaneous deployment, which cuts out whole classes of integration nightmares.
The real friction in the system is always in the boundaries between systems. With microservices it's all boundaries. Instead of a Ball of Mud you have Trees and No Forest. Refactoring is a bloody nightmare. Perf Analysis is a game of finger pointing that you can't defuse.
In terms of the common code divergence why not just use private NPM and enforce latest? Have a hard-fast rule that all services must always use the latest version of common.
If Linux tried to be an entire computing system all in one code base, (sed, vim, grep, top, etc., etc.) what do you think that would look like code base/maintainability wise? Sounds like a nightmare to me.
You mean like BSD?
Busybox is a slightly better example (even though it’s also a userspace program).
they're part of the same repo and built at the same time than the kernel. Run a big "make universe" here : https://github.com/freebsd/freebsd and see for yourself. That they are different binaries does not matter a lot, it's just a question of user interface. See for instance busybox where all the userspace is in a single binary.
That's news to me, and seems insane.
Unless you mean "their own database tables", not "database servers". But that's just the same as having multiple directories and files in a Unix filesystem.
If you can afford all your components sharing the same database without creating a big dependency hell, then your problem is _too small_ for microservices.
If your problem is so large that you have to split it up to manage its complexity, start considering microservices (it might still not be the right option for you).
If it could be avoided, these systems were not touched anymore. Instead, other applications where attached to the front and sides.
So I would say: it applies to systems where proper modularization was neglected. In the anecdotical cases I referred to, one major element of this deficiency was a complex database, shared across the whole system.
You would have to have an oddly disconnected schema if modifications to the program don't result in programs accessing parts of the database that other programs are already accessing. If this isn't a problem it means you're using your database as nature intended and letting it provide a language-neutral, shared repository with transactional and consistency guarantees.
so maybe not microservices, but fine nonetheless.
EDIT: two more comments:
- this is exactly what relational databases were designed for. If people can't do this with their micro-services, maybe their choice of database is the issue.
- "micro-service" as the original post suggests, is not synonymous with good. "monolith" is only synonymous with bad because it got run-over by the hype-train. If you have something that works well, be happy. Most people don't.
This video is strangely relevant in this thread... https://www.youtube.com/watch?v=X0tjziAQfNQ
Also I can pipe these tools together from the same terminal session, like
tail -f foo | grep something | awk ...
You don't have that in general with Microservices. Unix tools are Lego, Microservices aren't. They are Domino at best.Probably one could come up with an abstraction to do Lego with Microservices but we're not there yet.
There is the "base" system which is the OS itself and common packages (all the ones you mentioned), then there is the "ports" repo which contains many open-source applications with patches to make them work with OpenBSD.
Here is a Github mirror of their repos: https://github.com/openbsd
I think OpenBSD has reaped many of the same advantages described by Segment with their monorepo approach, such as easily being able to add pledge[1] to many ports relatively quickly.
The system is defined by a graph of these nodes, since you can pipe messages wherever they're needed; all nodes can be many-many. Each node is a self-contained application which communicates to a master node via tcp/ip (like most pub/sub systems a master node is requried to tell nodes where to send messages). So you can do cool stuff like have lots of seprate networked computers all talking to each other (fairly) easily.
It works pretty well and once you've got a particular node stable - e.g. the node that acquires images - you don't need to touch it. If you need to refactor or bugfix, you only edit that code. If you need to test new things, you can just drop them into an existing system because there's separation between the code (e.g. you just tell the system what your new node will publish/subscribe and it'll do the rest).
There is definitely a feeling of duct tape and glue, since you're often using nodes made by lots of different people, some of which are maintainted, others aren't, different naming conventions, etc. However, I think that's just because ROS is designed to be as generic as possible, rather than a side effect of it running like a microservice.
I get where you are coming from. But you will go async sooner or later if you need any reasonable error recovery or reliability.
The only question is how much pain you will suffer before you do so.
gRPC has no queuing and the connection is held open until the call returns. All of Google's cloud databases are immediately consistent for most operations
But strictly speaking, even inside a single multicore CPU, there is no such thing as immediate consistency. The universe doesn't allow you to update information in two places simultaneously. You can only propagate at the speed of light.
Oh, and the concept of "simultaneous" is suspect too.
Our hardware cousins have striven mightily and mostly successfully for decades to create the illusion that code runs in a Newtonian universe. But it is very much a relativistic one.
You mean GNU coreutils?
Right took for the job should be the goal as opposed to chasing fashion. Microservices are definitely overused, but they do have many legitimate use cases. Your CRUD web application probably doesn’t need microservices, but complex build systems might.
Imagine the complexity of Amazon.com or Netflix were a monolith. But something like Basecamp is probably better as a monolith.
There is also the issue of scale. An ML processor might need more (and different) hardware than a user with system.
Right tool (and architecture) for the job.
That was called Obidos and it's shortcomings were why Jeff pushed through the services mandate at Amazon.
Gotta put that CS Masters to work somehow. Can't just sit here doing plumbing all day every day.
Maybe there could be a software development corollary to the Politician's Syllogism[1], or even just a webdev one.
Services are not micro services. Most large scale applications can and should be split into multiple services. However, when approaching a new problem you should work within the monolith resisting the service until you absolutely can't any longer. Ideally this will make your services true services, that could capture an entire business unit. When it's all said and done you should be able to sell off the service as a business.
The other use case, which should be obvious, is compliance. If you are thinking about implementing anything that would require PCI or SOX you should do that in a service to shield the rest of the dev org from the complexities. So, any webapp that takes payment and interacts directly a payment processor.
That said, you're correct in that you should not be rolling out a new service to avoid sharding.
And even kernels have kernel threads, which are basically local microservices. Anything which needs to scale beyond a single system is more deserving of microservices than a kernel.
This is apples and oranges. The production profile of operating system kernels and web applications are so dissimilar that the analogy is not useful. It may be true that most web applications don't need to be split into multiple services, but the Linux kernel provides no evidence either way.
Agreed but I also prefer to keep any changing storage data as a separate concern (s3 or similar).
So the trinity of services would be DB, storage, application.
Similarly it’s made with make. If anyone has a project more complex than the Linux kernel or GCC I’ll gladly listen to why they need some exotic build system... never met anyone yet...
This exactly. Me too. Data consistency concerns in sufficiently large real world projects can be practically dealt with only 2 ways IMO: transactions or spec changes.
Microservices, when constructed from a well-designed model, provides a level of agility I’ve never seen in 33 years of software development. It also walls off change control between domains.
My take from the Segment article is that they never modeled their business and just put services together using their best judgment on the fly.
That’s the core reason for doing domain driven design. When you have a highly complex system, you should be focused on properly modeling your business. Then test this against UX, reporting, and throughput and build after you’ve identified the proper model.
As for databases, there are complexities. Some microservices can be backed by a key-value store at a significantly lower cost, but some high-throughput services require a 12-cylinder relational database engine. The data store should match the needs of the service.
One complexity of microservices I’ve seen is when real-time reporting is a requirement. This is the one thing that would make me balk at how I construct a service oriented architecture.
See Eric Evans book and Vaughn Vernon’s follow up.
I think the key is doing something that works, for you, and care a little less about what other people are doing.
Most paradigms have their advantages and disadvantages and it’s really just about working around that to best utilize the stuff you have.
Presumably someone with 33 years of software development experience is around 50 years old. How old do you think someone needs to be before they are qualified to comment on trends in software development practices? 70 years old?
I also would totally agree that companies should never just adapt "best practices" because those leads to super complex enterprise systems which are not necessary for most companies.
SOA and microservices are the same thing; microservices is just a new name coined when the principles were repopularized so that it didn't sound like a crusty old thing.
It's certainly a truism that technology can be cyclical, but that's not relevant in this case.
The OP's statement and article "Goodbye Microservices" is anecdotal and incorrect.
I have been on teams developing microservice architectures for about six years and this particular paradigm shift has proven to be a dramatic leap forward in efficiency, especially between the business, technical architecture, and change control management.
When you develop a domain model, the business can ask questions about the model. The architects can answer them by modifying the model. The developers can improve services by adopting model changes in code. This is the fundamental benefit of domain driven design and works fluidly with a microservice architecture.
There's still a pervasive belief in technology circles that software should be developed from a purely technical perspective. This is like saying the plumber, electrician, and drywaller should design a house while they're building it. They certainly have the expertise to build a house, and they may actually succeed for a time, but eventually the homeowner will want to change something and the ad hoc design of the house just won't allow for it. This is why we have architects. They plan for change within the structure of a house. They enable modification and addition.
Software development is no different. The Segment developers have good intentions, but they needed to work with the business to properly model everything, then build it. Granted, it sounds like they're a fast moving and successful business, so there are trade-offs. But once the business "settles", they really should go back to the drawing board, model the business, then build Segment 3.0.
From this reality it's good to design everything as if it's a (micro|macro)service part of a larger landscape of apps.
Reality is also that you can never have transactions for everything across all your systems, so transaction alternatives like compensations are always something to deal with.
Transactions are a nice and convenient shortcut when they're applicable but they're far from mandatory.
Plus of course options around Paxos etc:
Although personally, I've never felt the need to try and apply it specifically, but the idea is interesting.
The only nit I have on that video is that after a great motivation and summary, their example application at the end (processing game statistics in Halo) didn’t seem to need Sagas at all. Their transactions (both at the game level and at the player level) were fully idempotent and could be implemented in a vanilla queue/log processor without defining any compensating transactions, unless there were additional complexities not mentioned in the talk.
When she's discussing Compensations she mentions that the Transaction (T_i) can't have an input dependency on T_i-1. What are some things I should be thinking about when I have hard, ordered dependencies between microservice tasks? For example, microservice 2 (M2) requires output from M1, so the final ordering would be something like: M1 -> M2 -> M1.
Currently, I'm using a high-level, coordinating service to accomplish these long-running async tasks, with each M just sending messages to the top-level coordinator. I'd like to switch to a better pattern though, as I scale out services.
In general the separation of the domain should be such that you don't need a transaction across services.
* Old technology is deemed by people too troublesome or restrictive.
* They come up with a new technology that has great long-term disadvantages, but is either easy to get started with short-term, or plays to people's ego about long-term prospects.
* Everyone adopts this new technology and raves about how great it is now that they have just adopted it.
* Some people warn that the technology is not supposed to be mainstream, but only for very specific use cases. They are labeled backwards dinosaurs, and they don't help their case by mentioning how they already tried that technology in the 60s and abandoned it.
* Five years pass, people realize that the new technology wasn't actually great, as it either led to huge problems down the line that nobody could have foreseen (except the people who were yelling about them), or it ended up not being necessary as the company failed to become one of the ten largest in the world.
* The people who used the technology start writing articles about how it's not actually not that great in the long term, and the hype abates.
* Some proponents of the technology post about how "they used it wrong", which is everyone's entire damn point.
* Everyone slowly goes back to the old technology, forgetting the new technology.
* Now that everyone forgot why the new technology was bad, we're free to begin the cycle again.
This should be the top post on articles like this. Just to keep it fresh always in everyone's mind.
JS has improved massively over the last few years, it's very nearly an entirely different language than what it used to be.
There's a reason everyone was desperate to avoid writing JS not all that long ago be it the form of coffeescript, silverlight, flash, or GWT. ES5 JS sucks. Like... a lot.
As does most of Hacker News with every technology they hate on.
The issue I have with Node JS is that the entire ecosystem moves way too fast. However, when I need realtime interaction via WebSockets fe. it really feels like a good choice.
Node is no less stable, today, than Ruby or PHP, and is only arguably less stable than Python because Python has largely ossified.
SQL database usage is still an ongoing bummer, though.
Same happened with a Rails/React project that was setup last year, i tried to update one package that was marked as vulnerable, but it ended up requiring the same thing. I opted to leave that package.
I've been working with node for a long time, but i feel like is the same level of stability since the start. Thats very different with Ruby or PHP, you can have old setups working just fine, maybe requiring some extra steps to build certain dependencies but overall working.
2) Can you explain to me how updating a Gemfile or composer.json file is not going to result in a similar dependency cascade? 'Cause, from experience, it certainly will if the project isn't dead. About the only environment I've ever worked in where keeping up on your dependencies on a regular basis isn't required is a Java one--and that's assuming you don't care that much about security patches.
As an arbitrary example, consider the deps of two similar packages, delayed_job (Ruby, https://github.com/collectiveidea/delayed_job) and Kue (JS, https://github.com/Automattic/kue).
Delayed_job has two dependencies if you're running on Ruby (rather than JRuby): rake and sqlite3. Neither rake nor sqlite3 have any dependencies of their own - in production mode, of course.
On the other hand, Kue has nine direct dependencies, each of which have their own. The full dependency tree of Kue has one hundred and eighty two separate dependencies.
And, moreover, it basically doesn't matter. Pick what you like, change it later if you care. (You probably won't.)
Has it? I'm pretty sure if I were to ask what those best practices and tools are, I'd get a number of different answers.
But I bet they'd be the same ones those people used two years ago, or at least a really close variation.
webpack is six years old.
npm is eight years old.
nodejs is nine years old.
Anecdotal, but I personally feel like things have stabilized a great deal.
Likewise, yarn in place of npm.
I actually see Vue as more filling the gap for Angular. React uses JavaScript as the templating language instead of trying to use HTML attributes to turn HTML into a programming language like both Vue and Angular do. For that reason alone I have zero interest in Vue and learning yet another templating language. I'm pretty sure React will be here for a few more years. User base is still growing in fact.
And Yarn is a drop in replacement for npm. Took me 10 minutes to learn. If you already know npm, there is almost zero learning curve.
The code seems fine. But I don't really trust anybody who is that insistent on owning the changes made to my database schema--I am certain it is well-intentioned but it makes me itch. Although I will say that there's an interesting project[1] that creates entities from a database that I need to examine further and see if it's worth using to get around TypeORM's unfortunate primary design goals.
Objection.js is an ORM for Node.js that aims to stay out of your way and make it as easy as possible to use the full power of SQL and the underlying database engine while keeping magic to a minimum.
^^ copy+pasted from github
- I was worried at first because I progressively trust non-TypeScript projects less and less, but the test suite looks fine and they have official typings so there's that mitigation at least.
- I really don't love the use of static properties everywhere in order to define model schema. Which is probably a little hypocritical, because one of my own projects[0] does the same thing, but IMO decorators are a cleaner way to do it that reads better to a human.
- Require loop detection in model relations is cool. I like that.
- In general it's a little too string-y for my tastes. I think `eager()` should take a model class, for example, rather than a string name. Maybe behind the scenes it uses the model class and pulls the string name out of it? But I think using objects-as-objects is a better way to do things than using strings-as-references.
Overall, though, it seems very low-magic and I did understand most of it from a five minute peek, so I kind of like this. I think it's a little too JavaScript-y (rather than TypeScript-y) for my tastes, but maybe that can be addressed by layering on a bit of an adapter...I'll need to look deeper.
E.g. "Let's get really good at microservices so we don't need a monolith." IOW there's no consensus about when to dig in and go deeper. And even if there was, some people would leverage that information and try to pull ahead of the others digging in.
For example: web UI started as declarative (HTML -> PHP), and then transitioned towards more imperative (jQuery, Backbone, Angular), and is now moving back towards declarative (React, Polymer)
Each HTML document describes what it wants to have rendered, but does not describe how it should be rendered.
As far as PHP, it's an ugly language from a language-only perspective, but easy to deploy and has lots of existing web-oriented libraries/functions. Think of PHP has a glue language for its libraries and some front-end JavaScript to improve the UI. Python and server-side JavaScript may someday catch up, but still have a lot of ground to make up.
And then there's those altogether dropping Elixir for Go...
And Ruby was a major leap forward ~ in productivity.
If it was easy to figure this stuff out, they wouldn't pay us to try :)
Also, by merging Merb in Rails 3, Rails got a fresh lease of life. This is similar to how Struts lived on by rebranding WebWork as itself.
Finally, Ruby shows up prominently as a language developers dislike: https://stackoverflow.blog/2017/10/31/disliked-programming-l...
"deemed" is the keyword here. Old tech is "deemed" bad, new one is "deemed" good. Without any numbers attached, just by way of hand-waving and propaganda. And it's all "deemed" Computer Science :)
With how far mankind has come, it's a little silly to think that "natural" systems will remain the only things that science is concerned with.
...uh, what computer science are you talking about? Formal verification is a huge part of CS, and provability is a tiny part of what makes science science - systematic study through observation and experimentation. Science is a discipline, not in itself a fact to be proved.
Also, what parts of CS do you think are inapplicable to general purpose computing?
"I'd like to welcome you to this course on Computer Science. Actually that's a terrible way to start. Computer science is a terrible name for this business. First of all, it's not a science. It might be engineering or it might be art. We'll actually see that computer so-called science actually has a lot in common with magic. We will see that in this course. So it's not a science. It's also not really very much about computers. And it's not about computers in the same sense that physics is not really about particle accelerators. And biology is not really about microscopes and petri dishes. And it's not about computers in the same sense that geometry is not really about using a surveying instruments."
In Portuguese it is never called Computer Science as such, rather Informatics, Computation or Informatics Engineering if I do a literal translation.
And those that have Engineering in their name, are only allowed to be called that way if recognised by the Engineers country organization as such.
My degree is simply "MEng Computing". It's recognized[1] by the engineering association, although that's so irrelevant in most IT that I had to look up the organization: the "BCS (BCS - Chartered Institute for IT) and the IET (Institute of Engineering and Technology)."
[0] My job title includes the word "informatician".
[1] https://www.imperial.ac.uk/computing/prospective-students/ac...
I personally believe nested blocks produce a visual structure (indenting) that helps one understand the code by its "shape". Go-to's have no known visual equivalent.
Computers don't "care" how you organize software, they just blindly follow commands. Thus, you are writing it for people more than machines, and people differ too much in how they perceive and process code. For the most part, software is NOT about machines.
Dijkstra spent many, many pages on that. And it's still not a clear cut "this one is always better" case, as there are some obvious exceptions.
I've written a lot on the fortran-inspired MS Basic when a child. I know quite well how bad they can become.
Do note that some actual coders have claimed that if you establish goto conventions within a shop, people handle them just fine.
Yes, for example I love dense code. A 1 line regex is a lot simpler to me than the 50-60 lines of code equivalent.
Yet some programmers find the 1 line regex significantly more difficult even if they know regex.
Are stuck using the new tech, because implementing it burned bridges with the old tech.
Thesis --> Antithesis --> Synthesis
Antithesis: microservices in separate repos
Synthesis: microservices in a monorepo
Coming to this topic, I see Microservices as a solution to the problem of Continuous Delivery which is necessary in some business models. I can't see those use cases reverting back to Monolith architecture. For such scenarios, the problems associated with Microservices are engineering challenges and not avoidable architecture choices.
But still, how is the initial comment an example of the Hegelian Dialectic?
For these projects it seemed to make a ton of sense given the rest of the stack, the nature of the projects, the deadlines I'm facing, and I've really loved working with it so far.
Why did you end up moving off of it?
I thankfully don't have to worry about power loss (large company, teams dedicated to critical systems infrastructure in a big way), and a cron job can handle any archival concerns in the projects I'm running here.
It does sound like they've improved much since 2011.
Anyway thanks for the insight
I think the same could be said of the design world. There was a time not too long ago when designs actually felt polished and had real shapes, shadows, gradients. When you clicked on a button you actually knew you were clicking on a button. Then iOS 7 came along and everything became white and flat and buttons were replaced with text with no borders. I think we are slowly moving back to where we were ten years ago.
PS How long until people start ditching React for jQuery?
The problem you've identified is that most code, in general, is terrible. The code written by people who chase trends tends to be worse than average.
This. Instead of learning a handfull of technologies well they learn a lot of technologies very poorly. If I am hiring I now look at it as a red flag when people have too many frameworks listed.
SQL to me is a huge design flaw despite it's ubiquity. On the web bottlenecks happen at IO and algorithmic searches. Databases are essentially the bottlenecks of the web and how do we handle such bottlenecks? SQL; A high level almost functional language that is further away from the metal than a traditional imperative language. A select search over an index is an abstraction that is too high level to be placed over a bottleneck. What algorithm does a select search execute? Why does using select * slow down a query? Different permutations of identical queries causes slow downs or speed ups for no apparent reason in SQL. SQL is a leaky abstraction that has created a whole generation of SQL admins or people who memorize a bunch of SQL hacks rather than understand algorithms.
Part of the reason SQL has stood the test of time is the very fact that it allows such a high level of abstraction. The big problem that it solved, compared to much of what existed at the time, was that it allowed you to decouple the physical format of the data from the applications that used it. That made it relatively easy to do two things that were previously very hard: Ask a database to answer questions it wasn't originally designed to answer, and modify a database's physical structure without having to change the code of every application that uses it.
A lot of "easier" technologies - including, arguably, ORM on top of relational databases - make things easier by sacrificing or compromising those very features that allow for such flexibility. Which speaks to the grandparent's point about technologies that make it easy to get started in the short term, at the cost of having major disadvantages in the long term.
Many of the problems you mention above occur because the database handles stuff for programmers. Sure. you could create a custom solution around your biggest bottlenecks, but do you want to create a custom solution for every query, or do you want the database to do it for you. The generation of SQL admins is a replacement for a much larger group of programmers that would be needed if they weren't here, and more importantly, an army of people to deal with security, reliability, etc. that people using a good RDBMS get to take for granted.
Having worked on a couple projects that used Mongo at the core, I wish I could say the same.
which to you would be more clear?
SELECT * FROM TABLE WHERE id = 56
binary_search(column_name=id, value=56, show_all_columns=True)
x = binary_search(column_name=id, value=56, show_all_columns=True)
y = dictionary_search(column_name=id, value=56, show_all_columns=True)
z = join(x, y, joinFunc=func(a,b)(a==b))
The example above could be a join. With the search function name itself specifying the index. If no such index is placed over the table it can throw an exception: "No Dictionary Index found on Table"
This is better api design. However, years of optimization and development on SQL implementations makes it so that most no-sql api's will have a hard time catching up to the performance of SQL. It's like googles V8. V8 is a highly optimized implementation of a terrible language (javascript) which is still faster than wasm.
However, I think specifying the algorithms in the queries is really not a good idea. Your performance characteristics can change over time (or you might now know them at all yet when you start the project). With your solution, if you, e.g. realize later that it makes sense to add a new index, you'd have to rewrite every single query to use that index. With SQL, you simply add the index and are done.
Declarative yes, but unlike SQL my example is imperative. An imperative language is easier to optimize then a functional or even a expression based language (SQL) because computers are basically machines that execute imperative assembly instructions. This means the abstractions are less costly and have a better mapping to the underlying machine code.
>However, I think specifying the algorithms in the queries is really not a good idea. Your performance characteristics can change over time (or you might now know them at all yet when you start the project). With your solution, if you, e.g. realize later that it makes sense to add a new index, you'd have to rewrite every single query to use that index. With SQL, you simply add the index and are done.
Because the DB sits over a bottleneck in web development you need to have explicit control over this area of technology. If I need to do minute optimizations then am api should provide full explicit control over every detail. You should have the power to specify algorithms and the language itself should never hide from you what it's doing unless you tell choose to abstract that decision away...
What I mean by "choose to abstract that decision away" is that the language should also offer along with "binary_search" a generic "search" function as well, which can automatically choose the algorithm and index to use... That's right, by changing the nature of the API you can still preserve the high level abstractions while giving the user explicit access to lower level optimizations.
Or you can memorize a bunch of SQL hacks and gotchas and use EXPLAIN to decompile your query. I know of no other language that forces you to decompile an expression on a regular basis just to optimize it. Also unlike any other language I have seen literally postgres provides a language keyword EXPLAIN that allows users to execute this decompilation process as if they already knew SQL has this flaw. If that doesn't stand out like a red flag to you I don't know what will.
SQL on the otherhand... EXPLAIN is used on a regular basis, it's built in to the programming language and rather then just mark lines of code with execution time deltas it literally functions as a decompiler to deconstruct the query into another imperative language. This is the problem with SQL.
> Why does using select * slow down a query?
Because the database first has to perform a translation step to do an initial read from its system tables in order to enumerate the rows to be returned as a result of the final query.
> Different permutations of identical queries causes slow downs or speed ups for no apparent reason in SQL.
The key word there is "apparent", and again, just because it's not apparent to you, doesn't mean that it's not knowable and apparent to someone else. I also take exception to the concept of "permutations" of "identical" queries. Because if your query is permuted, it's no longer identical. The way you write your SQL has an impact on how it's evaluated. Just because you don't understand the rules, doesn't make it a mystery.
As a side note, I'd highly recommend reading up on the Relational Algebra that underpins SQL and other relational databases. https://en.wikipedia.org/wiki/Relational_algebra
A good abstraction only requires you to know the abstraction not what lies underneath. What we have with SQL is a leaky abstraction. My argument that a high level leaky abstraction placed over a critical bottleneck in the web is a design mistake.
Back at my first big tech company, I remember reading the best document I have ever read related to software engineering. It was entirely devoted to choosing your database/storage system. The very first paragraph of the document was entirely devoted to engraining in your head that "choosing a database is all about tradeoffs". They even had a picture where it just repeated that sentence over and over to really engrain it in you.
Why? Because every database has different performance characteristics such as consistency, latency, scalability, typing, indexing, data duplication and more. You really need to think about each and every one because choosing the wrong database/not using it correctly usually cause the biggest problems/most work to solve that you will ever have to face.
You aren't responding to my argument, everything you said is something I already know. So lol to you. You're making a remark and extending the conversation without addressing my main point. I'm saying that the fact that you need "a general understanding of how a database works on the inside" is a design flaw. It's a leaky abstraction.
A C++ for loop has virtually the same performance across all systems/implementations; if I learn C++ I generally don't need to understand implementation details to know about performance metrics. Complexity theory applies here.
For "SELECT * FROM TABLE", I have to understand implementation. This is a highly different design decision from C++. My argument is that this high level language is a bad design choice to be placed over the most critical bottleneck of the web: the database.
The entire reason why we can use slow ass languages like php or python on the web is because the database is 10x slower. Database is the bottleneck. It would be smart to have a language api for the database to be highly optimize-able. The problem with SQL is that it is a high level leaky abstraction so optimizing SQL doesn't involve using complexity theory to write a tighter algorithm. It involves memorizing SQL hacks and gotchas and understanding implementation details. This is why SQL is a bad design choice. Please address this reasoning directly rather than regurgitating common database knowledge.
And I just said that "choosing a database is all about tradeoffs" which you need to understand (aka: the leaky abstractions).
> A C++ for loop has virtually the same performance across all systems/implementations
> For "SELECT * FROM TABLE", I have to understand implementation.
No you don't, it has the same performance: a for loop. However, by grouping all of your data onto 1 server, for loops are much more costly than the likely orders of magnitude more regular servers you have than a database. Fortunately, your SQL database supports indexes which speed up those queries. Granted, I'm no database expert, but adding the right indexes and making sure your queries utilize them have solved pretty much every scaling problem I have thrown at them.
> It would be smart to have a language api for the database to be highly optimize-able. The problem with SQL is that it is a high level leaky abstraction so optimizing SQL doesn't involve using complexity theory to write a tighter algorithm. It involves memorizing SQL hacks and gotchas and understanding implementation details.
It is optimizable and 90% of those optimizations I have made simply involve adding an index and then running a few explains/tests to make sure you are using them properly.
If you'll only answer me this though, what database would you recommend than? I'm dying to know since you think you know better and google, a company that probably has more scaling problems than anyone else, doubled down on SQL with spanner which from what I have read, requires even more actual fine tuning.
And I'm saying the tradeoff of using a leaky abstraction is entirely the wrong choice. A hammer vs a screwdriver each have tradeoffs but when dealing with a nail, use a hammer, when dealing with a screw use a screw driver. SQL is a hammer to a database screw.
>No you don't, it has the same performance: a for loop. However, by grouping all of your data onto 1 server, for loops are much more costly than the likely orders of magnitude more regular servers you have than a database.
See you don't even know what algorithm most SQL implementations use when doing a SELECT call. It really depends on the index but usually it uses an algorithm similar to binary search off of an index that is basically a binary search tree. It's possible to index by a hash map as well, but you don't know any of this because SQL is such a high level language. All you know is that you add an index and everything magically scales.
>Fortunately, your SQL database supports indexes which speed up those queries. Granted, I'm no database expert, but adding the right indexes and making sure your queries utilize them have solved pretty much every scaling problem I have thrown at them.
Ever deal with big data analytics? A typical SQL DB can't handle the million row multi dimensional group bys. Not even your indexes can save you here.
>It is optimizable and 90% of those optimizations I have made simply involve adding an index and then running a few explains/tests to make sure you are using them properly.
I don't have to run an EXPLAIN on any other language that I have ever used. Literally. There is no other language on the face of this planet where I had to regularly go down into the lower level abstraction to optimize it. When I do it's for a rare off case. For SQL it's a regular thing... and given that SQL exists at the bottleneck of all web development this is not just a minor flaw, but a huge flaw.
>If you'll only answer me this though, what database would you recommend than? I'm dying to know since you think you know better and google, a company that probably has more scaling problems than anyone else, doubled down on SQL with spanner which from what I have read, requires even more actual fine tuning.
I don't know if you're aware of CSS or javascript and the glaring flaws everybody complains about front end web development but it's a good analogy to what you're addressing here. Javascript and CSS are universally known to have some really stupid flaws yet both technologies are ubiquitous. No one can recommend any alternative because none exists. SQL is kind of similar. The database implementations and domain knowledge have been around so long that even alternative NO-SQL technologies have a hard time over taking SQL.
Which brings me full circle back to the front end. WASM is currently an emerging contender with javascript for the front end. Yet despite the fact that WASM has a better design then JS (not made in a week) current benchmarks against googles V8 javascript engine indicate that WASM is slower then JS. This is exactly what's going on with SQL and NO-SQL. Google hiring a crack team of genius engineers to optimize V8 for years turning a potato into a potato with a rocket booster has made the potato faster then a formula one race car (WASM)
> For "SELECT * FROM TABLE", I have to understand implementation
This is not even remotely an apples-to-apples comparison. One is a fairly simple code construct that executes locally. The other is a call to a remote service.
It doesn't matter if the language you use to write it is XML, JSON, protocol buffers or SQL, any and all calls across an RPC boundary are going to have unknown performance characteristics if you don't understand how the remote service is implemented. If you are the implementer, and you still choose not to understand how it works, that's your choice, not the tool's. Every serious RDBMS comes with a TFM that you can R at any time. And there are quite a few well-known and product-agnostic resources out there, too, such as Use the Index Luke.
Alternatively, feel free to write your own alternative in C++ so that you can understand how it works in detail without having to read any manuals. It was quite a vogue for software vendors to sink a few person-years into such ventures back in the 90s. Some of them were used to build pretty neat products, too. Granted, they've all long since either migrated to a commodity DBMS or disappeared from the market, so perhaps we are due for a new generation to re-learn that lesson the hard way all over again.
>It doesn't matter if the language you use to write it is XML, JSON, protocol buffers or SQL, any and all calls across an RPC boundary are going to have unknown performance characteristics if you don't understand how the remote service is implemented.
Dude, then put your database on a local machine and execute it locally or do an http RPC call to your server and have the web app run a for loop. Whether it is a remote call or not the code gets executed on a computer regardless. This is not a factor. RPC is a bottleneck but that's a different type of bottleneck that's handled on a different layer. I'm talking about the slowest part of code executing on a computer not Passing an electronic message across the country.
So whether you use XML, JSON, or SQL it matters because that is the topic of my conversation. Not RPC boundaries.
>If you are the implementer, and you still choose not to understand how it works, that's your choice, not the tool's. Every serious RDBMS comes with a TFM that you can R at any time. And there are quite a few well-known and product-agnostic resources out there, too, such as Use the Index Luke.
I choose to understand a SQL implementation it because I have no choice. Like how a front end developer has no choice but to deal with the headache that is CSS or javascript.
Do you try to understand how C++ compiles down into assembler? For virtually every other language out there in existence I almost never ever have to understand the implementation to write an efficient algorithm. SQL DBs are the only technologies that force me to do this on a Regular Basis. Heck they even devoted a keyword called 'EXPLAIN' to let you peer under the hood. Good api's and good abstractions hide implementation details from you. SQL does not fit this definition of a good API.
If that doesn't stand out like a red flag to you, then I don't know what will.
>Alternatively, feel free to write your own alternative in C++ so that you can understand how it works in detail without having to read any manuals. It was quite a vogue for software vendors to sink a few person-years into such ventures back in the 90s. Some of them were used to build pretty neat products, too. Granted, they've all long since either migrated to a commodity DBMS or disappeared from the market, so perhaps we are due for a new generation to re-learn that lesson the hard way all over again.
In the 90s? Have you heard of NOSQL? This was done after the 90s and is still being done right now. There are alternative implementations to database API's that DON'T INVOLVE SQL. The problem isn't about re-learning, the problem is about learning itself. Learn a new paradigm rather than remark to every alternative opinion with a sarcastic suggestion: "Hey you don't like Airplanes well build your own Airplane then... "
EXPLAIN does a very good job of explaining why one query is faster than another.
> SQL is a leaky abstraction that has created a whole generation of SQL admins or people who memorize a bunch of SQL hacks rather than understand algorithms.
This just smacks of sound bite material.
The high-level was the point because, in the original idea, there was a separation of concerns assumed: The dev writes, in SQL, what the DB should do and the DBA decides how the DB does it.
Of course that assumes there is a competent DBA...
Putting a high level leaky abstraction over the bottleneck of the web is a mistake. A language that is a zero cost abstraction is a better design choice.
Yeah, but JS buys you a lot. There are certain things that you can accomplish with JS that you absolutely cannot accomplish without it.
OTOH, anything that you can accomplish with React, you can accomplish without React. I'm with the GP on that one.
If something works for you and makes life easier then you should use it. There is no right answer. You just need to be honest with yourself when planning things out - am I using this technology because it's new and shiny or because it is the right tool for the job right now.
I am well aware. It was mostly a joke :]
You say that, but from time to time I still discover slight variations in browser behavior or bugs that were opened 8 years ago that would've been avoided if I had just used jQuery. Most modern frameworks will abstract away these differences, but sometimes you'll need to access the DOM directly.
$_max = 1 - (H_c / H_h)
That is, the amount of money you can extract via consulting is proportional to the ratio of "hot" to "cold", representing the minimum and maximum hype for a given technology during the cycle.We can very easily break the cycle by training a deep learning TensorFlow brain in the cloud, that will be fed the daily mouse gestures and key presses of all developers in the world. It's an awesome new technology that can solve any problem.
Pretty soon the global brain will start to see patterns emerging, for example when developers post hype phrases on forums with unsubstantiated claims about the potential of some awesome new technology. As soon as a hype event is detected, a strong electric shock is commanded via the device the developer is using, thereby stopping the hype flow and paralyzing the devellllllllllllllllllllllllllll
"Maybe he was taking dictation?!"
It's amazing how many stupid to the point of crazyness situations seem perfectly natural nowadays. Thank computers!
The OP at the time suggest serverless should instead be re-branded Function as a Service which is a considerable improvement.
[0] https://www.reddit.com/r/sixwordstories/comments/1xp4wt/worl...
> In the beginning the Universe was created.
> This has made a lot of people very angry
> and been widely regarded as a bad move.”The new way to do it is not AI, but NI -- natural intelligence.
You take human babies, and you send them to schools and colleges where they learn programming.
Then you make them use daily mouse gestures and key presses to solve problems.
Its hard to explain, but it is the new leap forward.
The thing with design is that it isn't a science or formal field of logic. I can use math to determine the shortest distances from point A to point B but I can't prove why a design for product A is definitively better than a design for product B.
With no science we are doomed to iterate over our designs without ever truly knowing which design was the best.
Technology, in my experience, never seems to reward too much optimism or too much cynicism.
Something that does surprise me is that Panic's Coda and Transmit apps still seem quite successful, so maybe my perception is out of whack.
Many would call that bad practice, I call it static typing benefits
At the moment serverless vendors make that hard, and the frameworks (like serverless.com) are still emerging to make that simple.
For me the problem is most people don't think through the cost/benefits. Amazon don't ever get tired of saying "never pay for idle" in their sales pitches, but quite a lot of applications out there are never idle and can be quite accurately managed in terms of scaling, and therefore you're actually paying a premium for something you don't need.
GraphQL could wither like Backbone and Angular and nobody would really notice or care. An industry-wide shift away from SPAs would be something else entirely.
Sure, it's easy to work with at first but then you start to realize that every framework and every language has different ways of doing all the niceties that you now expect. Asset pipelining, layouts, conditional rendering, and template helpers all end up becoming stuff that every language has to individually develop with varying levels of success.
Even with those features baked in, you probably still want to modify the page using JavaScript, anyway, so then you have to re-render parts of the page without the aid of the expansive view system that rendered your page. And, of course, the more JS you put in your app, the more you have some bastardized hybrid of a SPA and a server rendered page.
> * Everyone slowly goes back to the old technology, forgetting the new technology.
This step is just as misguided as cult-y as "Everyone adopts this new technology and raves about how great it is now that they have just adopted it."
In some cases the technology WAS the right idea, just implemented incorrectly or not sufficiently broadly, and the baby ends up getting thrown out with the bathwater.
I think that regardless of whether microservices works for anyone or not, they came about to address a real issue that we still have, but that I’m not sure anyone has fully solved.
I think that microservices are an expression of us trying to get to a solution that enables loose coupling, hard isolation of compute based on categorical functions. We wanted a way to keep Bob from the other team from messing with our components.
I think most organizations really need a mixture of monolithic and microservices. If anyone jumps off the cliff with the attitude that one methodology is right or wrong, they deserve the outcome that they get. A lot of the blogs at the time espoused the benefits without bothering to explain that Microservices were perhaps a crescent wrench and really most of the time we needed a pair of pliers.
Does the service need an independent and dedicated team to manage its complexities, or is it a "part time" job? Try a Stored Procedure first if its the second.
Is the existing organization structure (command hierarchy) prepared and ready for a dedicated service? (Conway's law) Remember, sharing a service introduces a dependency between all service users. Sharing ain't free.
Do you really have a scalability problem, or have you just not bothered to tune existing processes and queries? Don't scrap a car just because it has a flat tire.
Does that make it the fault of the technology/pattern? I don't think so. I think it just means that there's no magic bullets in tech and people who don't know what they're doing will always cause problems no matter what models they follow.
See: The Lindy effect https://en.wikipedia.org/wiki/Lindy_effect
People really drink the koolaid that is written on these sites and it is extremely detrimental to their companies. PostgreSQL with a nice boring Java/.NET layer would blow this stuff out of the water performance wise (for their actual real life usecase), would be far easier to manage, deploy, find people for etc. I mean; using these stacks is good for my wallet as advisor, but I have no clue why people do it when they are not even close to 1/100000th of Facebook.
When the legacy systems started to hurt us (because they were written by the founder in a couple of weeks in the most hacky way), we decided against microservices and went to improve the actual code into something more performing and more maintenable,also moving from PHP5 to PHP7.
As much as we all wanted to go microservices and follow the buzz, we were rational enough to see that it didn't make any sense in our case.
Experience helps a lot. When you know that complexity and too much diversity breeds tech debt, you learn to say "No" decisively.
I witnessed someone that wanted to leverage their service into a promotion so they started pushing for an architecture where everything flowed through their service.
It was the slowest part of our stack and capped at 10tps.
No one above my pay grade seems to see a problem with this. But hey! REST! JSON! HTTPS! Pass the Kool-Aid!
[1] NAPTR records---given a phone number, return name information; RFC-3401 to RFC-3404
Get ready for the surprise twist: It wasn't going well. I was hired as an expert JS consultant to advise them on which JS framework to use.
My advice? Get ready for surprise #2: "don't use javascript [or use it sparingly as needed]."
But none of this implies that we are losing knowledge, just that the curve of engineers is fatter at the inexperienced level.
The problem is that the all the best knowledge has clearly not made it out. For example, this design introduces a "Centrifuge" process that redirects requests to destination specific queues... congratulations you've just reinvented a message bus, a technology that goes back to the 80s. There is absolutely nothing new about virtual queues as described here but unfortunately the authors are likely not at all aware of the capabilities of real enterprise messaging systems (even free, open-source ones like Apache Artemis) and certainly not aware of the architecture and technologies and algorithms that underlie them and the (admittedly much more expensive) best-of-breed commercial systems.
(I won't even go into the craziness of 50+ repos. That's just pure cargo cult madness.)
Watching the web/javascript reinvent these 30 year old technologies is a little disheartening but who knows they may come up with something new. (Then again, recently the javascript guys have discovered the enormous value of repeatable builds. Unfortunately the implementations here all pretty much suck.) Still, we ought to perhaps ask ourselves why this situation has come about...
Maybe someone needs to create a cloud based serverless SaaS SPA web app developed in F# to help track this stuff and prevent it happening in the future.
It seems like the problem here was bad testing and micro repos, not microservices.
>However, we weren’t set up to scale. We lacked the proper tooling for testing and deploying the microservices when bulk updates were needed. As a result, our developer productivity quickly declined.
My impression after reading this post was that microservices were symptoms of problems in how their organization wasn't set up to implement them effectively, rather than the actual cause of those problems.
It's amazing to me how people still "meh" away testing as a secondary concern, and then regret it later. Over and over again.
WRITING software is easy, anyone can do it. CHANGING software is extremely difficult. THAT is why we have tests. Also, if you are smart about it, you can get documentation out of the deal for relatively little additional cost.
My go to example is on-boarding new developers:
New dev: "OK, i'm here! How do I start?"
with tests: Clone the repo, install dependencies, and run the test suite. As you develop new features, be sure to write additional test.
They are up and going in a matter of minutes.
without tests: Clone the repo, install deps, download testing database, achieve homeostasis with your dev environment, learn the entire system, build up the state you require to write your feature, iterate on it by hand over and over again.
My first job out of college was like that. I had been doing professional-ish (I was paid and employed but I basically worked alone with no other engineers around) work for two years but this still didn't raise any flags.
You basically quoted exactly the setup.
So the initial problem was a single queue? Well, then split the queue, no need to go all crazy splitting all the code.
Switching to 100+ microservices? There is no need to switch to 100+ repos too, runtime services don't need to have one repo per service, just use a modular approach, or even feature flags.
100+ microservices, some of them with much lower load than others? Then consolidate the lower load ones, no need to consolidate "all" of the microservices at once.
Library inconsistencies between services? No, just no, always use the same library version for all services. Automate importing/updating the libraries if you need to.
A single change breaks tests in a way you need to fix unrelated code? WTF, don't you have unit tests to ensure service boundary consistency and API contracts?
Little motivation to clean up failing tests? Yeah... you're doing it wrong.
Only then you figure out to record traffic for the tests? HUGE FACEPALM, that's the FIRST thing you should do when dealing with remote services!
This answer drove me crazy in the article. "We had trouble keeping the libraries up to date and fixing breakages, so our solution was... to update them all and fix the breakages." And per your second point, tests can go both ways - if libX is used by serviceY, write a libX integration test for serviceY instead of / in addition to a serviceY integration test for libX.
I've never understood the false dichotomy of microservices vs monolith... just split things when it makes sense. ¯\_(ツ)_/¯
If you did this then you'd have to go and update all services whenever you wanted to introduce a breaking change. I don't think what you're suggesting is as easy as it sounds.
Otherwise, just update to the latest library versions whenever you touch a codebase.
A few services > monolith
monolith > 100's of services.
The big trick with any technology is to apply it properly rather than dogmatically and if you are breaking up your monolith into a 100's(!) of microservices you are clearly not in control of your domain. That's a spaghetti of processes and connections between them rather than a spaghetti of code. Just as bad, just in a different way.
Engineer 1: "We'll have one repo per downstream service."
Engineer 2: "But we have hundreds of those, so now we have to manage hundreds of github repos???"
Anyone sane: "That doesn't sound right, we should rethink this..."
A couple of hundred VMs is nothing in a scenario like that. Good luck trying to debug anything.
I though this was an honest and interesting look at a decision which in retrospect was a bad idea. Hopefully it'll stop some other people making similar mistakes (too many repos, too many services, fast changing libraries shared between many services, etc...).
It'd be better if it wasn't framed as having found that the one true way is the monolith, but there are some lessons here for most devs.
Actually worse in many ways:
- Harder to test
- Harder to debug
- Harder to deploy
- Harder to monitor
- Harder to reason about, refactor, change/add functionality
- You've (basically) turned a lot of the operations your services need to perform into RPCs, probably killing performance on top of everything else, leading to
- More complex/demanding (or just MOAR) infrastructure requirements
- Higher dev, maintenance and infrastructure costs
- Slower delivery of value to the business and customers
- Potentially crippling opportunity costs
It surprises me how often people don't see this coming. Seriously: keep your systems as simple as you possibly can. Unless you're Netflix, dozens or hundreds of microservices probably isn't as simple as you possibly can.
This strikes me as the core of their problem, and every step taken was a way to bandaid this limitation. Would the cost of moving to faster-scaling infrastructure been so high as rearchitecting the entire system?
> When we wanted to deploy a change, we had to spend time fixing the broken test even if the changes had nothing to do with the initial change.
This seems like a separate and even larger problem. Changes are breaking tests for unrelated code areas? Is the code too tightly coupled? Sounds like it. The unit tests are doing exactly what they're designed to do. Hard to feel sympathy for the person who's breaking them and then trying to figure out a way to sidestep them rather than fix the underlying issues.
It seems like the primary problem causes were flaky/unreliable tests, and difficulty making coordinated changes across many small repositories.
Having worked on similar projects before (and currently), with a small team driving microservices oriented projects, I would probably recommend:
1) single repository to allow easy coordinated changes.
2) a build system that only runs tests that are downstream of the change you made (Bazel is my favorite here, but others exist). This means all services use the HEAD version of libraries, and you find out if a library change broke one of the services you didn't think about. This also allows for faster performance.
3) Emphasis on making tests reliable. Mock out external services, or if you must reach out to dependencies use conditional test execution, like golang's Skip or junit's Assume if you can't verify a working connection.
If you still can't build a reliable service with those choices, then it's time to think about changing the architecture.
But by reading the first paragraphs of the article you see that the guys from Segment made a series of grave mistakes on their "microservices architecture", the most important one being the use of a shared library on many services. The goal of microservices is to achieve isolation, and sharing components with specific business rules between them not only defeats the purpose, but results in increased headaches.
Without deep knowledge of the solution, it's hard to judge, but it seems this was never the case for microservices. They needed infrastructure isolation when the first delay issues surfaced, but there wasn't anything driving splitting the code up.
Sam Newman discusses on his book how to find the proper seams to split services (DDD aggregates being the most common answer) and it seems to people are making rather arbitrary decisions.
In general, I fully concur: there are very few services that actually warrant a microservice architecture.
Many large companies have millions of LOC behind their microservices. Your average startup probably doesn't.
I predict the same will happen to the 'superstar', it just isn't old enough yet (and at least it was built with some badly needed domain knowledge).
https://twitter.com/kelseyhightower/status/94025989833123840...
If you had ten teams of seven and they managed two services each... it's easier to see how that architecture could actually help.
Same as if you have a two person team building a web-app and they go for client-server architecture rather than a basic full stack web framework. If you're both working the frontend and backend at the same time, save yourself the extra ops effort.
I'm not interested in figuring out exactly what the right marketing term for it is, but I've had good experiences with teams of 6-10 engineers owning something like 2-5 services with a larger ecosystem of dozens to hundreds of services. Of course, I've been working at very large companies with extremely high traffic for several years now, so my experience is skewed in that direction.
If I had three engineers on my team I'd be unlikely to end up with more than a small handful of services. Half the benefit of splitting up your services has to do with keeping your independent teams actually independent -- if it's just one team then that isn't a problem in the first place.
We have a small team and have a handful of services. We also have a fully automated CI/CD pipeline. It's worked really well. I doubt any would be considered 'micro', but instead they are designed around functional areas like authentication or backend processing.
Yeah, who knows what micro means, but that's exactly how I like to split up services. If at some point a service gets too large, split it up. Hundreds of services out the gate is a gross premature optimization. (And like most premature optimizations, ends up costing much more time both in development and in maintenance.)
I've actually had experiences with seemingly this same problem at a previous startup. Once we started spinning off individual repositories for small pieces of business logic stuff started to go downhill as the logistics of communicating and sharing one another's code became more and more complex.
The real solution is microservices in monorepos.
I have seen this phenomena of thinking that a large system, broken down into tiny parts, is somehow easier to manage time and time again over 30 years of development. In every case, the one thing the central thinkers fail to realize is that complexity is like conservation of energy - it can be transformed, but it cannot be destroyed.
Also, when it comes to large teams I have seen one thing work when it comes to sharing a resource(s) critical to a larger system - shared pain. If the central/reusable code/service breaks everyone's stuff, then everyone forms a team to immediately address the problem before continuing on. The solution is almost never "find a way to let the other teams continue while something important is on fire." It seems like a major motivation for a microservices architecture seeks to avoid the pain - which perhaps is not the best reason to use microservices.
I like the idea of unseen, but indispensable, complexity. For instance, the human brain is probably the most complex thing in the world, but the interface is fairly simple :)
We have architects here that dictate the design of the system but IMO they have not done the simplest implementation of anything. We have Kafka to provide ways of making each service eventually consistent so we delete something out of our domain service, but it requires absolutely huge amounts of code in various different other services to delete things in each place listening for events. Every feature is split across N different services which means N times more work + N times more difficult to debug + N times more difficult to deploy.
The system has been designed with buzzwords in mind - Go and GRPC have been a disaster in terms of how quickly people have developed software (as has concourse - so many man hours wasted trying to run our own CI infrastructure it's unreal), loads of small services that are individually difficult to deploy and configure (and come with scary defaults like shared secret keys for auth - use a dev JWT on prod for example). The difficulty in dealing with debugging the system - there simply aren't the tools to understand what is going wrong or why - you have to build dashboards yourself and make your application resilient to services not existing.
Never ever try to build Microservices before you know what your customers really want - we've spent the last 6 months building a really buggy CRUD app that doesn't even have C and D fully yet. Love your Monolith.
In my experience, you won't know whether you need microservices until you're on at least v2.0 of your application. By then, you have a better understanding of what your real problems are.
- a strong API between components
- network calls
The former is a very good idea that should be implemented widely in most code bases, especially as they mature, using techniques like modules and interfaces.
The latter is incredibly powerful in some cases but comes at a huge cost in system complexity, performance and comprehensibility. It should be used sparingly.
Small services are easier to develop with several teams, in my opinion. Each team knows what to input, and output. They can do whatever in between as long as these two contracts are respected.
But the overlooking of all these moving pieces changing at different paces is tough. And the smaller the services get, the harder it will become.
Not a size fits all, clearly!
There are some pretty big asterisks next to running "nanoservices". Mostly how expensive they actually are to run at large scale and the weird caveats that can happen due to them not always being up.
And I wouldn't advocate for "always-up nanoservices".
The basic answer to both "nanoservices" and microservices is do what you think is right but don't go too far. There are good reasons to make a nanoservice and good reasons not to, same with microservices.
As soon as you have shared library, you now have coordinated deployments. And that is just not fun and will cause problems.
The trick here is that this does mean you will duplicate things in different spots. But that duplication is there for a reason. It is literally done in two places. When you update the service, you have to do it in a backwards compatible way. And then you can follow with updates to the callers. This makes it obvious you will have a split fleet at some point, but it also means you can easily control it.
If your build and test process doesn't actually exercise your deployment pipeline across versions, then it's not testing anything. I don't think shared libraries are a problem at all - they're probably a good idea - the problem is when they're used as an excuse to not worry about testing your upgrade and rollback scenarios.
I'd pan that criticism a bit wider too: it's not just about having a testing process either - it's about making sure your devs are able to easily use it and watch it work as part of their regular cycle.
This is in fact something I'm about to start working on at my new job for a new project - pushing the development of each service down so someone can write `make test` and not only run tests, but see if what they've done can upgrade between the currently deployed version.
What did each service own, if they all shared code? Make those ownership lines as crisp as you can. And unless you have x teams, consider not having too many more than x services.
The more I read about the problems people have with microservices, the more I'm convinced they've never read about flow-based programming.
If you have two different teams collaborating or you expand beyond what a single box can do, create services. But if you can express things reasonably as a single service, why make things more complicated and error prone?
Whether they were wanting to factor into another service with a defined and mostly static API or a common library with a defined and mostly static API, it was a failure to factor common code into an amorphous blob that gates the release of all the other services. Instead, they've de-modularized the code and called that a success.
If you're drawing hard lines between services and having them talk to one another, having one of these models for at least parts of that makes a lot of sense: * filter system of the flow-based nature * REST API that does a transformation and returns transformed data * message broker with producer/consumer model where the consumer of one queue does a transform and puts the data into another queue * a full actor model * a full flow-based model
The biggest challenge is making the shared library forward and backward compatible with itself for at least a few releases in either direction, because not everyone will redeploy at the exact same moment.
If you can't solve that problem everything gets painful. Doing that right was the second hardest part of that job (meetings were the hardest). The job title (securing the data interchange) came in third place.
Now those micro services shouldn’t always be out of process modules that communicate over HTTP/queues, etc. A microservice can just as easily be separately compiled modules within a monolithic solution with different namespaces, public versus private classes, and communicate with each other in process.
Then if you see that you need to share a “service” across teams/projects, or a module needs to be separately, deployed, scaled, it’s quite easy to separate out the service into versioned packages or a separate out of process service.
I've endured a lot of suffering at the hands of the microservices fan club. It's good to see reason finally prevail over rhetoric.
It would have been nice if people had written articles like this 2 years ago but unfortunately, people with such good reasoning abilities would probably not have been able to find work back then.
Software development rhetoric is like religion. If you're not on board you will be burned at the stake.
So many times during technical discussions, I had to keep my mouth shut in the name of self-preservation.
This sounds like an issue with being able to articulate why something is or isn't going to net the expected benefits or being able to foresee unexpected risks. Keeping silent is better than throwing out silly hyperbole risks, but not bring up real risks because "they don't want to hear it" is completely bogus. Any solid engineer will bite at another potential risk to ensure they don't find themselves engineered into a corner 65% through a project. Your comment also makes it out like the notion the article is making, monolith over microserves, is gospel for every situation; that in no condition would it ever make sense to use microserves and that only naive zealots would espouse the wisdom (dogma) to use them. You can use any piece of technology poorly, that doesn't mean the core concept is flawed, just that your problem space is different than what that software is trying to solve. Consider using HDFS as a primary data store in place of MySQL where it doesn't make sense and you might cry the wisdom of wishing someone had told you HDFS is terrible and to just use the tried and true MySQL of olden days.
The issue is not articulation of ideas; the issue is that when all the books, all the articles and all people believe that something is true, there is no amount of articulation which will be able to convince them otherwise.
You have to wait for the hype to go away before even considering bringing up the argument.
This didn't challenge any "common wisdom". Common wisdom didn't tell them to do this. Cargo cult engineering is not common wisdom.
> You push data consistency concerns out of the database and between service boundaries.
At work, we have close to ~50 services(no one calls them microservices), but they do not suffer from this brittleness. We segregate our services based on languages. So, all C services go under coco/ , all Java services go under jumanji/ , all go services go under goat/ , all JS services go under js/. This means, everytime you touch something under a repo, it affects everyone. You are forced to use existing code or improve it, or you risk breaking code for everyone else. What does this solve? This solves the fundamental problem a lot of leetcode/hackerrank monkeys miss, programming is a Social activity it is not a go into a cave and come out with a perfect solution in a month activity. More interaction among developers means Engineers are forced to account for trade offs. Software Engineering in its entirety is all about trade offs, unlike theoretical Comp Science.
Anyway, this helps because as Engineers we must respect and account for other Engineers decisions. This methods helps tremendously to do this. No one complains, everyone who wants 1000 more microservices usually turns out to be a code monkey entangled in new fad, or who doesn't want to work with other Engineers.
You want to use rust? There is a repo named fe2O3/, go on. Accountability and responsibility is on your shoulders now.
If you think about it, an Engineer is tied to his tools, why not segregate repos at language level instead of some arbitrary boundary no one knows about in a dynamic ecosystem?
The irony being that anything approaching SOA (or microservices) requires exactly the same amount of communication. More likely they require more since it's almost certain that such a decision introduced chaos.
A lot of software is used to control or impose someone's will on others organizationally.
Currently our JIRA kanban board is crippled because someone just had to take a simple system and impose a workflow that can't be deviated from.
Well-written APIs are SLAs on steroids.
The nice thing about a SOA architecture is that it makes it harder to do this kind of cowboy programming. Yes maybe some quick business wins are harder than in a monolith, but I think (at least in our org) the cleaner architecture pays off by letting us move quicker on a different class of product projects since there is less technical debt.
You cannot possibly have every destination be a separate repo and then have the development lifecycle of your shared code be so active that it ultimately puts at risk the architecture of your entire organization.
What makes shared code so perfect is having stability such that you extract your variant code into your non-shared code. Shared code should evolve at a much slower pace than your non-shared code or you risk this very outcome.
Microservices are not dead, nor are they the solution to everything. We need better architects.
An architect is just a another man with an opinion.
The certificate means nothing, the training and instruction is priceless. Add to that things like TOGAF and a deep understanding of the current state of existing architectures and you'll understand what I originally meant, but failed to explain.
UPDATE: misspelled "meant"
We need science and theory, and skilled architects who can apply it to the problem at hand.
I'll take a TLA+ specification over a diagram any day.
Systems design in real life is an "artistic science": there are always known limitations (and some expected unknown ones) that rule out the theoretic optimal design for good reasons. The problem is that limitations and compromises are often not disclosed, and that many programmers and architects are too inexperienced to really grok the implicit meaning behind specific design decisions.
So we struggle along, with bloggers, researchers and FOSS contributors halfheartedly collaborating to make point improvements as solutions are discovered. Stodgy enterprises suffer, big tech companies make decisions with global impact, startups are left wondering WTF to do, and the rest of us largely don't care. Why? Because nit picking doesn't solve business problems [almost ever. I'd argue this point in cases of things like SCADA systems and other mission critical control systems.].
Nobody argues about which algorithm is better for sorted data sets: linear search or binary search. Theory already establishes one is faster than the other, but no theory establishes which architecture is better than the other.
TLA+ is one such system that includes a language for writing models and a model checker to verify them for you. In the context of microservices you would write a model of your services and the checker would help you to verify that certain properties of your model will hold for all possible executions of the model. Properties people seem to be interested in are consistency and transaction isolation. You can develop a model of your proposed microservices architecture and work out the errors in your design before you even write a lick of code.
Or if you already have a microservice system you could write a model of it and find if there are flaws in its design causing those annoying error reports.
Amazon wrote a paper about how they use it within the AWS team [0]. Highly worth the read. And if any of this sounds interesting I suggest checking out Hillel Wayne's course he's building [1].
[0] https://lamport.azurewebsites.net/tla/formal-methods-amazon.... [1] https://learntla.com/
Everyone else, including doctors, who call themselves scientists are just trying to float on the cachet physicists earned with their astonishingly good predictions. Properly speaking, they are phenomenologists . Please note that I'm not saying what they do isn't of great societal and intellectual value! The study and categorization of phenomena is certainly a noble enterprise. But none of them can make predictions good to 9 decimal places.
*might be billionths or quadrillionths by now in QED, but doesn't change my point.
Things like algorithmic complexity are well understood and formalized but design patterns are not a science nor has the concepts ever been formalized..
There is no theory or formalized system that says monolithic is better than micro or vice versa, it's all opinion. That's why its' called design.
For your latter point, it’s a matter of elegance, which is a pretty way of saying cognitively manageable. Think of epicycles vs Newtonian mechanics as an analogy. With enough epicycles you can compute the same result, but Newton’s approach is still a clear scientific advance.
Your point about design is well taken. Any given design is analogous to a theory. So we should aim for the simplest and most cognitively manageable design that satisfies our needs. That’s not literally formalized, but it’s a well established principle with an excellent record.
That being said, practical medicine is much less scientific than many think. There is a lot of master/apprentice learning going on, just as in software engineering.
In the world of math you don't need empirical data to verify a point. It's all logic derived from a small set of axioms.
From an architectural perspective, there is absolutely no difference between a micro service and a library. The only real difference is in the dispatch mechanism.
The problems around configuration management are the same problems we've had as programmers for decades. It's just that the people who are keen on micro services are usually not old enough to have experienced the pain in a different context.
Should shared libraries be in different repos, or should you put everything in a single repo? How do you deal with versioning? What happens if one app wants version 1 of the library and another app wants version 2? How do you deal with backwards compatibility of the API? Do you make a whole new library when you decide the old API is incompatible with the new vision? Blah, blah, blah, blah. None of this is a new problem.
I can't remember which version of Windows it was (maybe 7?) where they were seriously delayed mainly because they had so many programs using different versions of libraries. Integrating it at the end was apparently complete hell. Since they wanted to have separation of responsibility in their groups, each group was just pounding away implementing the features that they needed, but not integrating as they went -- because that would mean lots of cross team communication. The exact same thing is likely in a large organisation with a ton of micro services.
There isn't just one way to solve the problem. Mono repos and monoliths help in certain ways and cause problems in other ways. There are other techniques as well (should we implement ld.so for micro services? :-) ) But as you mention, the real answer is that the solution requires humans, not technology.
I’ll stick to linking glibc, friend architect. Performance is also architecture.
I always tell people if you can't write and maintain a library then don't do microservices.
It's so silly that we are wasting so much time rebuilding existing things, poorly.
Working exactly along those lines this week on multiple services exposed over REST APIs, I was wondering if tools exist to check compatibility between them. Said differently,
- I have a `swagger.yaml` for my service A managing chipmunks and it says endpoint `/chipmunks` supports an 'color' query parameter.
- In service B, I have a `handleToServiceA` that encapsulates calling A. Then say I write `const chipmunks = handleToServiceA.getChipmunks({'colour': 'blue'})`.
Are there (whatever the ecosystem) tools that would read serviceA's swagger.yaml, detect my error in serviceB (color -> colour) and report the issue at compile time rather than run time?
One step better is autogenerated client libraries where a new version is created every time a new version of swagger.yaml is deployed. However, I don't know an open source project that does this.
Could you share methods that worked well for you? Asking because those question come up a lot and there never seems to be any conclusion.
Architecturally speaking, there are some similarities between micro services and libraries because they’re both forms of modularization and usually have an API, but there are some stark differences beyond the “dispatch mechanism”.
The main difference is that a service’s deployment lifecycle is completely up to the service admin. Microservices are like websites - they can continuously evolve (within their API’s contract) without asking permission from consumers. This is their main superpower and why they’re a way of scaling a development organization without slowing it down too much.
In the case of a shared library, it’s completely up to the host admin as to when to upgrade. In the case of a static library, it’s up to the consuming software to determine when to upgrade. A service can upgrade when it feels like it.
Issues of API backwards compatibility, forwards compatibility, extensibility, self descriptiveness, versioning, etc. are old issues but usually have different answers when upgrades are truly happening all the time and not just in theory. It tends towards much fewer hard versions and more evolutionary backwards compatibility.
IOW, microservices aren’t a cure all, but they do encourage a set of behaviors. Many articles detracting from them seem to have not wanted those behaviors in their org in the first place.
Shared libraries can be an implementation of micro services. On a platform (e.g. Android) it can be that a shared library is updated and then all consumers are forced to update.
Services, whether SOA, web services, messaging services, or network services etc, as in SOA, describe a client/server architecture with an API. Usually these APIs are designed by the principles of Domain Driven Design, where different teams map to bounded contexts that have their own published API.
Microservices are a form of SOA where each API also runs in its own process and thus has independent deployment lifecycle. Many in the SOA world advocated for this 10-15 years ago (and often ignored), and now the industry has gone back and coined a term for this practice.
In an organization that does microservices properly at scale, like Amazon, you have teams that build and run and upgrade their service autonomously from others. Read the Steve Yegge rant about Amazon and Platforms to understand this. It allows tremendous parallelisation of effort and allows for thousands of deploys to production daily without breakage. This is hard to pull off though, and a new initiative doesn’t generally won’t more than a small handful of services.
If it’s a shared library, then call it a shared library. It has a completely different lifecycle from a microservice. In your Android example, shared libraries by definition are controlled by whomever can dictate OS updates, not the app developer.
Now, people are pondering things like: Ok, management of elasticsearch has performance issues and it's a general pain if the elasticsearch documents change. So let's try to move the schema of elasticsearch documents into a strictly semver'd artifact. And let's move management of our search indexes into a service depending on that artifact, so the schema changes in a controlled way. And ops and the search team can scale and optimize the service as they need to minimize search outages.
Creating smaller services based on problems is a good thing. Creating smaller services because of ... reasons... not so much.
Conway’s Law in action, or in reverse https://en.wikipedia.org/wiki/Conway's_law
e.g When we added Go, we had to present what it brings? A good testing framework, light weight(subjective), goroutines(which fit our use case), in built benchmark support, mature support of third party packages etc etc.
Again, YMMV benchmark for your use case.
That said, Java doesn't have go routines, doesn't provide low level access to memory, benchmarking framework is easy to work with in Go. Go gives access to sync Pool which when use correctly gives extremely good gains.
One of our measly dual core VMs, can do more than 40k QPS with predictable latency at 99 percentile. With JAVA oom panics were common and latency percentile was all over the place. Sure, you can perhaps tweak the Java code, but it was far easier to do with Go.
Meanwhile our Go services almost never had a working set larger than 256MB and most didn't even need that. We could schedule on average 16 Go services for each JVM service.
One of our core services is written in C, it never crashes, it sustains QPS no other service can in the entire ecosystem, why change it to something new unless there is a solid reasoning to it?
I kid you not, one of the companies we integrated with gave us 2 64 Core VMs in their private cloud to run our service. We run 3 instances of our process on each VM along with Redis taking about 60GB of memory in total, with 59.996 MB of memory taken by Redis. And it was written in Go. Apparently, we miscommunicated that we will be using a GC language for this service. They assumed Java and based on their experience allocated 2 64 core VMs with 256GB memory each :D.
Just like any other software, interfaces live in separate repos e.g protoplasm/ for proto-buf definitions. avrobber/ for avro and so on so forth.
So, any change to protoplasm/ triggers automated tests on all other services irrespective of boundaries.
What I mean this this. Any IDL interfaces live at their functional boundaries. e.g all proto definitions shared internally by java services live under jumanji/proto/, all proto definitions shared internally bu Go service live under goat/proto/ and so on so forth.
But anything which is shared in a public manner e.g any proto definitions between Java and C live outside either of these repos in top level common_proto/ repo. The build system is configured to take care of pulling the correct proto version from common_proto repo.
We use maven for this, but nothing much we can do about that. Build systems are ugly and we have to live with maven in near future.
EDIT: Oh that said, stop using required field in protobufs if you are stuck using proto2 like us, whoever decided to add required field to proto2 grammar made a mess. Things change and deprecate repeatedly, gazillions of required fields are annoying and recipe for disaster.
Full disclosure: INTP here.
It just seems like it just creates an us vs. them ideology.
Aww, no clever name or the JS services?
Buts it's especially terrible because it was designed by a guy without much language expertise in a handful of weeks.
The core isn't too bad; I'd argue that the main reason why it's hard to work with is because you don't have tight control over your execution environment, mostly since the language wasn't really standardized until pretty late. It also solves a much different problem than most languages since it has to optimize for better UX (not crashing the whole program on exceptions, for example).
Yeah, it's a bit slow. It's faster than Javascript, though. The best benefit is flexible number size, so it's good for some types of computation.
What improvements in package management would you want? At least it has some form of management, unlike C/C++...
Mypy coverage, even in the stdlib, is awful. When it comes to thirdparty, mostly nonexistent. Mypy feels so young - I love the team, love their work, but I still run into cases where inference fails when it shouldn't, where error messages are extremely unhelpful, etc.
Slow.
Plenty of 'wat' like mistakes: https://stackoverflow.com/questions/3270680/how-does-python-...
Exceptions everywhere, for control flow even - iterators are implemented with exceptions.
Absolutely AOT unoptimizable. Pypy's cool, never got it working for my use case. In theory a JIT could help.
Speaking of calling out to C... you think you're writing in a memory safe language, but actually, you're writing in a memory safe language that's probably been hollowed out and replaced with a fast C implementation. But it's actually worse
https://hackernoon.com/python-sandbox-escape-via-a-memory-co...
Exploiting C code loaded by Python is like exploiting C code from the 1990s.
No parallelism. Multiprocessing? Good luck with that - pay the cost of pickling, pay the cost of an additional interpreter, pay the cost of debugging hell.
An ecosystem split in two, and don't let anyone tell you otherwise - a few hundred top packages moving over after many, many years, is a sad state for a language that was known for having an absurdly large ecosystem.
I could really just go on and on and on, but at some point it just feels mean.
Huge respect for the project and the team but Python has made mistakes (as all languages do). They were understandable mistakes, but they were mistakes. It's fine for some things, but there's plenty wrong with it, just like there's a ton wrong with javascript. But javascript gets probably 1000x the flack.
Yeah it is, choose 3. Python 2 is reaching EOL.
> Mypy coverage, even in the stdlib, is awful. When it comes to thirdparty, mostly nonexistent. Mypy feels so young - I love the team, love their work, but I still run into cases where inference fails when it shouldn't, where error messages are extremely helpful, etc.
Yeah, the typing stuff is a pretty big disappointment. The ergonomics are pretty terrible (probably because they wanted to push as far as they could without introducing more syntax support for typing). Mypy isn't just young, but it's buggy and its codebase was a sloppy mess last I checked. There's no support for recursive types (you can't define a JSON type, for example). And it absolutely falls over in the face of common libraries, like SQLAlchemy, which are too dynamic for it.
> Absolutely AOT unoptimizable. Pypy's cool, never got it working for my use case. In theory a JIT could help.
Yeah, Pypy is the best hope for Python's performance. They're making great progress, but I also couldn't get it working in our Python 3 codebase (Numpy and Pandas installation issues).
> No parallelism. Multiprocessing? Good luck with that - pay the cost of pickling, pay the cost of an additional interpreter, pay the cost of debugging hell.
Yeah, this is a real pain point. Some die-hard Python folks say otherwise, but there's really no good parallelism option for lots of workloads. Pickling is just too expensive.
> An ecosystem split in two, and don't let anyone tell you otherwise - a few hundred top packages moving over after many, many years, is a sad state for a language that was known for having an absurdly large ecosystem.
This hasn't been a problem for me for years. Most things that have seen active development in the last 5 years have good Python 3 support. The only time I've run into a Python-2 only utility, it was 7 years stale. Well, except for Centos's `yum`.
These are all reasons I like Go, by the way. Super fast, great tooling, and fairly stable (except for the package management story). I do wish there was a lightweight scripting language with a great VM and real parallelism--something like JS without the inheritance, OO baggage, etc; just objects and arrays and functions running on a JIT VM like V8 but with parallelism as a first-class citizen a la BEAM. And preferably optionally typed from the start, in a way that the runtime could leverage for optimization purposes.
Easy to say. Tell that to Google and Dropbox - Guido has worked for both and they're on 2. Tell that to the companies that can't afford the creator of the language.
3 is not the easy choice.
Anyways, maybe try nim? I've heard good things.
Plenty of the environment still has not moved over even outside of companies.
I don't know the state of its coroutines though as I've never used them.
People hack the `continue` statement into it all the time, it wouldn't be terribly surprising.
Some things are more pleasant in Python--everything is sync by default, there's less churn, the standard library gets me a lot farther, it's less permissive (no 'undefined is not a function' nor '!= vs !=='), etc. Both languages serve similar niches well, but the feeling I get is that JS is quite a bit faster and nicer for IO-heavy workloads (JS just has a more mature async story than Python) while Python is generally more intuitive and perhaps better for general application development. But for most things, one isn't dramatically better than the other, and they're both a good deal more disappointing than Go. :p
This is horrible at scale.
> This is a classic case of not understanding micro services and trying to fit a problem around a tool.
That much I agree with. TFA even acknowledges that, in the conclusion. Not in so many words, but they basically admit they did it wrong.
>This is horrible at scale.
I just want to reiterate this. In the early 2000s I worked in the online platform group at EA. The list of things done poorly there was long, but picture:
* 40+ engineers
* Monorepo with hundreds of thousands of classes; all code deployed to all servers.
* Hundreds of different services running across thousands of servers.
* Communication based on Java serialization, so all code had to be deployed to all servers at the same time.
* Deployments (and thus downtime) sometimes lasted up to a hour. Worldwide audience; it was always in the middle of someone's day.
* Rational Clearcase for version control. It took nearly an hour to sync to tip.
Pretty much every morning you'd come in, spend an hour syncing, find that someone broke the build, hunt them down, and resync for another hour. Generally speaking the first few hours of every day were wasted for 40 engineers.
This was a very poor platform.
Sometimes I wonder how it's going there these days.
When they rolled out clear case in my team at IBM, I quit a few weeks later.
This can work fine at scale, Google does it with however many tens of thousands of engineers they have these days. Having everything in a monorepo doesn't solve the communication problem, it doesn't prevent solving it either.
This sentence needs to be repeated for everything on the 'wrong end' of the hype curve.
This is a classic case of not understanding NoSQL and trying to fit a problem around a tool.
This is a classic case of not understanding OOP and trying to fit a problem around a tool.
This is a classic case of not understanding dynamic typing and trying to fit a problem around a tool.
This is a classic case of generally understanding
Turing completeness and eventually, inadvertently
implementing a half of Common Lisp in C anyway.
Something like that?arguably they aren't close to each other in tiers of the stack, and don't really overlap. The sorts of libraries you might choose to use to do something could differ, (e.g. json processing - in the front end, you'd pick usability and security over a lower level more optimized transcoder for the api).
That said, I agree with your core premise- not understanding, and chasing a nail with a hammer, but geekily named repos per language just sounds like someone who doesn't understand how something git works...
Please?
Seriously though constraining dev's to use a monolithic repository in the name of discouraging cowboys is like tying people's legs together to ensure they walk in an aligned direction. Sure it works but seriously unnecessary pain. Your company should invest in integration tests
In a micro service architecture, the request bounces through different service layers, json serialization, network transfers.
In a monthlith normal application, especially if its on one machine, the request touches main memory and then CPU cache.
Latency Numbers Every Programmer Should Know, I especially think of the Send 1K bytes over 1 Gbps network vs cache latency of microservices/monolith. https://gist.github.com/jboner/2841832
Ie instead of micro services 10000 ns for a network transfer you could fetch within 10ns from CPU cache in a monolith, that is 1000 times more efficient.
Are we beyond the peak of inflated expectations on Micro services and towards the Plateau of productivity? https://en.wikipedia.org/wiki/Hype_cycle
If you get your DevOps ducks in order, a lot of the issues you have with standing up and maintaining individual microservices should be manageable. I certainly understand the pain of dependency management, but that's also a part of good architecture.
I'm willing to hear more of these stories though. We can always learn about edge cases or even new paradigms that come out of current thinking.
If you don't understand these concepts and how much work must be don't to correctly implement microservice architecture - you SHOULD STAY with monoliths, they are much easier.
Does that mean that because of all the work required to properly set those up for a monolith, we better stay with microservices?
> if a bug is introduced in one destination that causes the service to crash, the service will crash for all destinations.
This sounds dangerous. If I was a destination provider, and returned some garbage, could that take system down?
Also, reading up on centrifuge, instead of breaking up by source, destination; would it be more spatially conservant to create queues by response type, which are finite. The entire time I was wondering why bad requests are put back into the same queue they came from. Shouldn't they be isolated? You could then treat those isolated requests as one unit, so in a destination outage, your not backlogged on your actual compute resources for good requests.
Nonetheless, seems like just another day in an engineers workday :)
This ^^^. Everyone has an opinion on the Internet... frontend, backend, no-end. It's probably not going to align perfectly with your needs (are you Google/Facebook?) . There are a lot of tools & specs out there. If your blindly following someone else's opinion without truly understanding your users, product, & team you'll probably accrue debt.
Also, devs often forget the team. If your choice of tech cuts the team's velocity in half then it probably wasn't a good choice... even if it had some other technical benefits.
software engineering walks in circles. NoSQL people seem to be growing up and learning about consistency and transactionality. I'm waiting for somebody to discover threads and shared variables back as a revolutionary way to improve performance and greatly simplify the implementation of a system of "actors".
Joking aside, i think greatest reason for Segment's microservices failure was that they didn't use Kubernetes.
I say this as someone who worked in an environment widely castigated for its "monolithic" nature: It sounds like trying to have "modularity" hurt you, because you didn't actually have it in the first place. When you stopped pretending, the pain went away.
My company is probably typical in that we implicitly use microservices because we have consume dozens of microservices provided by other companies.
We only have a handful of services we maintain, each with dedicated engineers.
And what becomes a service? You want to answer two questions:
1. Does it have a concrete business justification? 2. Does it have clear functional requirements?
If it has a business justification, it will get people assigned to keeping it running. If it has clear functional requirements, it will make sense to the people working on it which service does what.
That's still pretty vague, so you want to look at who has done it well, and why it worked for them.
Companies like AWS have been extremely successful in using microservices because every last one service has a business rationale, it's either directly making money (EC2, S3, SQS) or it supports the needs of customers who are using a service that makes money (VPC, Cloudformation, all their internal auth, billing, provisioning, security and such).
The caveat there is that the big B2B service providers are not a great model for companies that have a lot of business logic.
And shared library shouldn't contain business logic...
What's the point then? They become libraries then, not services.
edit: also it seems they're missing a solid release cycle and manager. For every release before going to production there's a TON of testing that needs to be conducted, they're mentioned over 40 improvements this year, that's two releases a week. It's not possible that you're properly testing each of those releases. Had they done a single, well tested release with multiple fixes every 1-2 months, the burden would be significantly reduced.
Originally you couldn't trigger Lambda functions from SQS (seemingly the most obvious integration). You could use Kinesis but small print says Lambda concurrency is restricted to the Kinesis streams which gets very expensive.
Visibility/monitoring into most microservices is not good (Iron.io is quite nice but any concurrency is really expensive). I don't like the workflow for deployment and testing either.
So I shifted to single EC2 instance with my application and Beanstalkd w/ my own configurable workers. Way cheaper, easier to manage, normal programming workflow, etc.
For some use-cases Lambda and other services are really nice and efficient but there's usually a lot of hidden limitations so be sure to spend a lot of time evaluating before committing. You often spend way more time fighting the microservice drawbacks than the benefits are worth.
Furthermore, if one uses modules (as one should), one can arbitrarily and somewhat trivially run those modules either in-process (compiled in) or out-of-process (via REST, gRPC, Cap’n Proto, or another RPC system), e.g., in a separate service/microservice/whatever you want to call it. This gives you a best of both worlds approach where code can be arbitrarily run in a monolith or a separate service as-needed. This changes the thought dynamics from a rigid "monolith vs. microservice" decision to a more fluid process where things can be rather easily changed on a whim. When modularity is the goal, then services become something of a secondary concern.
Microservices are used as something of a sledgehammer to force modularity and performance in languages that lack proper modularity and/or are innately slow or otherwise inefficient, while suffering orchestration costs and the performance penalties of copying data across multiple processes and networks as well as making it harder to derive a single-source of truth in some cases.
Probably a good approach for a typical webappp looking to improve performance would be to first port core logic to a modern, fast, compiled language with modules, then evaluate the performance from there, and then determine if any modules should be split out into separate processes or services.
Like NoSQL, microservices can be (but not always are) a case of the cure being worse than the disease; however, they can also be useful in certain situations or architectures. Like anything in engineering, there are tradeoffs and it depends on your situation.
- All teams will henceforth expose their data and functionality through service interfaces.
- Teams must communicate with each other through these interfaces.
- There will be no other form of inter-process communication allowed: no direct linking, no direct reads of another team’s data store, no shared-memory model, no back-doors whatsoever. The only communication allowed is via service interface calls over the network.
- It doesn’t matter what technology they use.
- All service interfaces, without exception, must be designed from the ground up to be externalizable. That is to say, the team must plan and design to be able to expose the interface to developers in the outside world. No exceptions.
- Anyone who doesn’t do this will be fired. Thank you; have a nice day!
still kinda works.It seems like Segment didn’t really understand this at all, and instead decided to have seemingly arbitrary and rediculous service boundaries that had no relationship to the real world.
See also people who create poor abstractions in their code and other sources of technical debt.
This isn’t an article about how microservices architecture is somehow bad, it’s an article about how bad Segment’s engineering team is. You can make the same argument for anything. At the end of the day some architecture or process or programming language or tech can’t replace real thinking about your problem and correct application.
> In early 2017 we reached a tipping point with a core piece of Segment’s product. It seemed as if we were falling from the microservices tree, hitting every branch on the way down. Instead of enabling us to move faster, the small team found themselves mired in exploding complexity. Essential benefits of this architecture became burdens. As our velocity plummeted, our defect rate exploded.
How do you address something like that coherently? I sort of agree with the OP, it kinda sounds like they just didn't know what they were doing. Or maybe that quote wasn't really the meat of their complaint?
A single team should never be maintaining 140 different microservices. That's not a failing of the microservices architecture, that is a failing of massively abusing the microservices architecture. Arguably, a small team should never be managing more than 10 services (totally arbitrary number, but it seems like the upper limit to what a human can focus on).
It sounds like Segment very much prematurely optimized. Microservices is probably not a great approach below 20 developers, and becomes necessary past 100, at least in my experience.
You can quibble with our implementations, but at the end of the day, the proof of the engineering is in the working. Our system works even better than before and our customers appreciate it!
That's totally unnecessary and reflects poorly on you for having said so.
It doesn't! Your tech stack has next to no impact on the success of your product.
Making something useful is really all that matters. Is the service running? Great, that's all the end user cares about. The back end could be a bunch of code copy-pasted off old form posts from 1998 and string together with scotch tape as long as it works.
The end user doesn't care about what language you use. They don't care if you use redis, mysql, postgres, or mssql 2000. They don't care if the HTML or CSS is messy. They won't judge you if you don't use SASS or LESS or whatever is popular today.
The most successful side project of mine runs on a $20/m VPS and has been re-written 3 times and the code has been pretty terrible every time because every re-write was done in a new language I learnt over a weekend and never touched again. The first was some horrible PHP, then some less horrible PHP, then GO, and recently someone re-write it (properly) in elixir. The end user doesn't know, the end user doesn't care. The site gets about 250K uniques a month and the API handles about 350 million requests a month. The crappy GO I wrote in a weekend worked just the same for the end user as the fancy elixir.
Anyways, sorry for the rant. I've just seen so much wasted effort go on projects that never launch because people are too concerned about stuff that really doesn't matter in the long run.
If your stuff runs like crap because it found a market fit and you have people signing up and your crappy code can't handle it, then that is a FANTASTIC problem you now get to solve.
Wow. This is the level of discourse we've reached here now.
Why the ad hominem? If you have an actual point to make on technical merit, by all means, make it. Were you passed on in a job interview at Segment? Why the hostility? (honest question)
If the data doesn't need to exist in the same database, put it in a new service with a separate database (or at least completely independent tables).
Ideally there is no shared code between services (and there shouldn't need to be, because each service owns completely disparate data), so the only coupling is the API definitions.
Each service is free to use whatever internal architecture they see fit as long as they honor the API definitions they provide to their dependent services.
In the case outlined in the article, the fact that each microservice had similar enough concerns to all use some common libraries and the same database makes me doubt that this should ever have been built as microservices to begin with.
> However, a new problem began to arise. Testing and deploying changes to these shared libraries impacted all of our destinations. It began to require considerable time and effort to maintain. Making changes to improve our libraries, knowing we’d have to test and deploy dozens of services, was a risky proposition. When pressed for time, engineers would only include the updated versions of these libraries on a single destination’s codebase.
When pressed for time, engineers would only include the updated versions of these libraries on a single destination’s codebase.
I think I see the problem and it wasn't with microservices.
* splitting everything up into separate repos and services seems like a pretty radical move. You could start with separate queues that are handled by a single service, potentially with many instances, or you could try to break out a few parts that change often or have a very high load
* A big chunk of the complexity seems to be in transforming one message format to another. That is something that should be very easy to write tests for. So you need a CI that tests all the services when you change the base libraries, and then it's pretty easy to find out if a change to a base library is backwards compatible or not. And for libraries that are shared between many services, you should mostly stick to backwards compatible changes.
+
-js
-java
-monolith.jar {guava, netty, etc }
-svc1
-svc2
-python
-go
For biz code, I've seen this kind of lib-ifying architecture provide a nice microservices workflow. The unifying heuristic is: Any dependency goes into the lib project. Anything unique to the service that requires no dep should usually go into the service project. The nice thing about this is it modular with fast compiles and preserves optionality. Since everyone is using the same deps, the code can live in the svc project or the lib project. A bit of namespacing convention makes it's trivial to shuttle the code between the two projects to wherever it's most natural to have it.Well, I'm not surprised. Microservices are made possible by advances in automated testing, CI, and deployment. You should have those things anyway, even if you have a monolith -- but to go to microservices without them is a pretty bad decision.
But that's an "all in" environment. You are Elixir/Erlang/OTP, or you are out. That is, understandably, not an option for many use cases.
At least they finally got to the right conclusion, they were on the wrong path to begin with.
People seem to be blaming Microservices when they weren’t even close to understanding what they were doing and why. I’d be much more interested in an article about issues faced with Microservices where they actually tried to slice their functionality based on their domain.
Had they taken the time to explore the root causes of these problems and how to best approach them, they probably wouldn‘t have taken the microservices approach to begin with.
I see this far too often. Instead of looking for solutions to problems, people look for problems to try the new hot solution they read about. Happened with ML/AI, NoSQL, Microservices, Blockchain, …
A shared library works up until a certain point. If every service uses a shared library then you are already getting to a world of a monorepo. A monorepo for different services should work fine if the overall architecture is feasible.
In my professional experience, where I've run into numerous examples of good and bad examples of both, microservices tend to win because it's a lot easier to unwind the badly implemented microservices compared to the badly implemented monolithic service.
Of course, this is just another small set of anecdata.
Sounds like they just redrew their service boundary, from Integration APIs to Business Function.
Centrifuge sounds like a new service deals with connecting to integration APIs, so they've replaced 140+ services with one.
Another service they've spun up is Traffic Recorder, and its responsibility is to eliminate the need for http requests when testing integrations.
Feels like the biggest change is going from so many repos to a monorepo.
The title and opening paragraphs gave me the impression they felt they were moving away from microservices, but maybe I didn't those bits correctly.
1. Tests hitting 3rd party APIs are flakey & slow 2. The job queuing mechanism can cause all jobs to be slowed by a single 3rd party API outage/slowdown
Eventually they arrived at:
1. Replay responses to speed up HTTP based tests 2. Create a smarter queuing mechanism in house
I'm not sure what microservices has to do with any of this. Anyway, kudos to them for having the courage openness to share their learning from mistakes!
... 140 tightly coupled services wasn't the problem then?
Isn't the recommended practice to treat third-party services and libraries like a black box? That's what I do, and that's how I figured what their solution was (roughly) before I read it. Felt a bit proud of myself.
Sure, but at the same time the team learned some valuable lessons, gained experience and hopefully matured. I applaud Alexandra for opening up and sharing that with us.
To use words that are mine, microservices are a hack on Conway's law. One team should be responsible for 2-3 microservices and should have a lot of autonomy.
It would have saved them a lot of man-years of time.
I would be curious as to what the focus and scope of these shared libraries were such that they required frequent updates with cascading side-effects requiring everything that leverages them to have to be updated.
I swear more people writing microservices need to read about flow-based programming and perhaps the actor model.
https://www.sandimetz.com/blog/2016/1/20/the-wrong-abstracti...
This seems like a weird reason to adopt microservices. Can't you isolate tests within the same repo using folders?
The real reason people should investigate Erlang/Elixir.
But does perl5 even need hype? Soon it'll age into retro-chic status, ala lisp. :)
businesses/companies at early stages of growth make decisions that allow them to be competitive and profitable, "the holy grail of application architecture" might not apply yet or for the first 5 years, or ever.
Monolithic Scylla DB built on top of venerable Seastar framework is very good example of under-rock-living competitive advantages.
I can't imagine you have applied Conway's law.
I think there's some serious confusion between FaaS and microservice architecture.
? There is nothing new about microservices. They've been a hot topic for far longer than Segment has been an idea.
Then I was tasked with building a scalable platform around a message broker system.
Then the switch clicked. I get it now.
Introducing components to match on various forms of value seems odd. Ideally standardise on those names; but if that's not possible (and you're using microservices), why not have a service to correct the name, rather than a library deployed to every endpoint? Then you call that service with data containing `first_name` or `givenname` and it returns that message with the standardised form. Or have a "global key" service; where you send the source system name and value, and have that translated for the destination via a lookup, allowing any field names or defined values to be translated by a generic reusable component:
Entity | System Name | System Value | GlobalKey
-------------------------------------------------------------------------------------------
Boolean | HR | TRUE | 52582622-4322-445b-bb7a-8ca118d0ca2b
Boolean | HR | FALSE | 688c2298-6b99-4c31-a356-1ab9b1caacde
Boolean | Finance | 1 | 52582622-4322-445b-bb7a-8ca118d0ca2b
Boolean | Finance | 0 | 688c2298-6b99-4c31-a356-1ab9b1caacde
Boolean | Sales | Yes | 52582622-4322-445b-bb7a-8ca118d0ca2b
Boolean | Sales | No | 688c2298-6b99-4c31-a356-1ab9b1caacde
FieldName | HR | GivenName | 3de31cff-e4bb-4d4a-a819-fd96c7c5032e
FieldName | HR | Surname | e799b891-1eeb-4598-85a5-a9534d3a3a4c
FieldName | Finance | First_Name | 3de31cff-e4bb-4d4a-a819-fd96c7c5032e
FieldName | Finance | Last_Name | e799b891-1eeb-4598-85a5-a9534d3a3a4c
FieldName | Sales | FirstName | 3de31cff-e4bb-4d4a-a819-fd96c7c5032e
FieldName | Sales | Surname | e799b891-1eeb-4598-85a5-a9534d3a3a4c
As for "handcrafted XML"... the alternative looks like handcrafted code... For translation a language like XSLT which was designed for translating data from one format to another seems like a good choice. You can also use this to get around your translation issue by defining a central "universal" format; so you can have XSLTs to translate messages from source systems to the universal format, then other XSLTs to translate from universal to the destination; so that adding or removing a system (/service) only impacts that system / you don't need to rewrite every point-to-point interaction of that service with another (OK, not point to point since we have microservices; but if they're not adding value by abstracting you away from point to point then essentially you've just got complex point-to-point interactions rather than simple point-to-point).Microservices work when they are separate, independent, single-concern systems that coordinate using APIs. People often go overboard in splitting apps up into small pieces, even when those pieces logically belong to a single system. Start with figuring out the subsystems, then considering whether they are worth splitting in the first place.
It's worth pointing out that microservices don't mean separate repos or even codebase separation. What matters the most is encapsulation. Monoliths grow horrible because they end up being balls of spaghetti, and forcing modularization at the service level is a way to avoid such messes by reducing the individual parts manageable sizes, allowing a part to be replaced without being concerned about its tendrils having grown through the whole system.
For me, the biggest value of microservices is composition, of thinking abou modules as off-the-shelf components that you use as parts to build something bigger. Using a complete enough set of microservices, I can build a frontend or client that has zero app-specific backend code. For example, if I have a generic data layer (think Firebase), a user database layer with OAuth/OIDC, and a way to store images, then I can build Instagram from scratch with no backend development at all. That's very powerful.
But once I need some specialized, app-specific stuff ("business rules"), such as rating of photos, commenting, moderation, etc., then those probably wouldn't be microservices! The concerns are unified there for the most part, and disentangling them would just lead to annoying fragmentation. A single use-case specific service ("monolith") that would exist at the center of it all.
On the other hand, composition is mostly useful if your pieces are going to be reusable. If I intend to build more than one Instagram, or maybe a Facebook (which also needs data storage, and logins, and photos, etc.), then the individual pieces would be reusable and could just be shared between the apps. But if I'm just building Instagram for 5 years and I'm not building a series of apps for different use cases, reusability has zero importance, and I might as well just move everything into a single monorepo and forget about making anything general-purpose. (Each piece should be general-purpose enough, but they usually don't need to be so generic that you could open-source it for everyone.)
I never liked the word "microservice", and I think we'd be better off if we called them, say, modules or subsystems.
> Eventually, all of them were using different versions of these shared libraries. We could’ve built tools to automate rolling out changes, but at this point, not only was developer productivity suffering but we began to encounter other issues with the microservice architecture.
In the extreme case where you have 150 microservices and 3 devs, I think that spending a few days to build a tool that auto-updates your common deps and re-runs your tests would be a good investment. Or you could pay someone else to do this with a service like https://www.dependencies.io/. Handling common code is one of the known pain points in microservices, so it's worth tackling head-on. (Last I saw Netflix handles this by the rule "no services can share code unless it's by an open source library", which encourages common code to be thoughtfully packaged and released.)
> The additional problem is that each service had a distinct load pattern. Some services would handle a handful of events per day while others handled thousands of events per second. For destinations that handled a small number of events, an operator would have to manually scale the service up to meet demand whenever there was an unexpected spike in load.
I can see this being tricky to tune, and am hesitant to opine without knowing the details, but if you can fix the problem by bundling all the services into a single monolith (i.e. aggregating all load into N nodes), then you should also be able to fix the problem by using a cluster scheduler like k8s with equivalently sized nodes. As long as you don't have bursts that are an integer factor of your baseline system load, both approaches should work equivalently (to the first order). It sounds like they were running individual instance(s) per microservice, which isn't a good fit for very bursty services.
And as a bonus, with a cluster scheduler you get a number of primitives to do resource reservation, which you don't get for free if you're merging all of your services back inside a single monolith. This means the problem of back-pressure from a single misbehaving endpoint -- which was one of the reasons they moved to microservices in the first place -- will probably come up in some form down the road.
> Recall that the original motivation for separating each destination codebase into its own repo was to isolate test failures. However, it turned out this was a false advantage. Tests that made HTTP requests were still failing with some frequency. With destinations separated into their own repos, there was little motivation to clean up failing tests. This poor hygiene led to a constant source of frustrating technical debt.
I don't have anything to say here except... don't do this? If your tests are failing you should be fixing your tests (or removing them if they aren't adding value), not adding new integrations. Consistently-failing tests are a big warning sign that your CI/CD process is not in good shape, and a healthy CI/CD process is a strict precursor to doing microservices successfully.
> The outbound HTTP requests to destination endpoints during the test run was the primary cause of failing tests. Unrelated issues like expired credentials shouldn’t fail tests.
Your UTs shouldn't be hitting your external dependencies; the correct solution here is the one that they eventually landed on, i.e. to either record/replay real HTTP requests, or to mock out the HTTP responses manually. You still need real integration and smoke tests in a production-like environment to make sure that you've not missed a change in the remote API schema. IME this is one of the biggest challenges of working with external APIs, and I don't envy the task of maintaining 150 integrations, however this issue seems unrelated to microservices.
Thanks, but no thanks.
Key sentence. Using the wrong tool for a small team, blaming the tool.
I was a product manager on a team of really rockstar developers. They all earn at least $200k a year.
Instead of demanding more ambitious projects, you could keep about 99% of them happy by just letting them use the new framework of the week to build their next web app. Their excitement when they were green-lighted to use react was mind-boggling. Building another dumb website with the new framework, yay.
Mindlessly applying microservice architecture is the same issue at heart.
React (and others like it, e.g. Vue) really is a huge win and solves a ton of pain points common to front end development. It still has its own pain points, but it's hard to overstate how much of an improvement it is over other older approaches like jQuery.
The parent didn't understand why developers would be happy that React got the greenlight. My response was that React (and Vue) really is a much better tool for building interactive websites than many older tools like jQuery.
> Building another dumb website with the new framework, yay.
He's just astonished how excited they are although what they build is "another dumb website" in his opinion. Many developers love the "how?" and don't care about the "why?". I don't judge it, but it can explain some of the excitement. Angular, React and Vue are great for development. But Ember, Backbone and others were also good back in their times. What he essentially says is "We use specific technology so they get excited although the problems we solve are boring af"
Developers have to use the new and shiny to keep themselves marketable and be ready to jump ship at the first opportunity or out of necessity.
My salary has gone up by $45K - $50K in the last 4 years by changing jobs three times. I’m not complaining.
But on the companys’ side, it was completely illogical not to pay me market rates. They still had to pay my replacement market rates and they loss institutional knowledge when I left.
Why is it that HR will approve a req for a new developer at market rates but have strict limits on what they can pay current employees.
On other hand you still have things broken in react not to mention few years with uncertain patent situation.
https://twitter.com/sveltejs/status/999704064937156611 sums this up pretty well.
Programming world is in a constant flux, now you have lit-html that gives you react in 5kb without vdom overhead. And in next few years there will be other interesting projects to look at.
Uh, lit-html seems like just another templating library. Marginally better than something like Soy. Calling it "React" is pretty reductive.
Because doing something new is much more fun that doing something effective. And most software developers don't really get into this job out of doing things effectively - we're get into it out of the sheer fun of it, altough we would never admit it to our bosses and often even to ourselves.
Trouble is, "fun" for the enthusiast / self-perceived-non-workdrone does align with "productive". =)
My initial reaction is that it should be the opposite: people who are enthusiasts would want to learn something new because they enjoy using and learning technology, whereas a "workdrone" would focus on being productive because they are focused on work.
I'm not saying you're wrong or that there's anything wrong with either mentality, I just don't find your statement self explanatory. Please do elaborate =]
Granted, it's absolutely true they're often used in places they probably shouldn't be and can easily result in JS bloat. When working on the frontend it's tempting to use them for everything because of how comparatively fast/easy development is. Ultimately the tradeoffs should be considered on a case by case basis.
- Next year, the new kid on the block will be the most amazing thing ever created (tm) and everything else will be obsolete (knockoutjs, angular just to name a few)
- If you pack everything now and leave the code untouched, in a couple of months there's a big chance half your tools will need upgrading (angular?) When you upgrade, half the things won't work and you will spend a ton of time fixing things that were perfectly working before, just to get the thing compiling again
-If, down the line, someone else has to take the work where we left, not having the exact development machines set up, they're going to have it much harder to make simple fixes.For example, in six months of development, a simple datepicker javascript library changed the way where it declares the locale config, and we needed to change every declaration of that function to make it work everywhere again. When a framework deprecates a function that you use while you're in development, you fix it. If someone has to pick the project two years later, they are f*
With "old" stuff this doesn't happen. Sure, if you pick php 3 code and try to run it in PHP 7 you'll find problems, but nothing like with webpack/yarn/whatever is doing the packing work this month
Nowadays is all Vue and react, a couple years ago it was Ruby, next year will be something else. In the end, whatever works for you will be good enough, but if you have to think in advance and try to get some future proof code, half the things that are 'what needs to be used right now', will be unmanageable legacy code a year or two down the line (Just my two cents)
It's unglamorous, monotonous and boring, but it's wildly important to the continued success of the project. Who can blame engineers for not being interested in trudging through the boring middle?
Interestingly sometimes efforts to keep engineers from leaving, like splitting an application into 100's of micro services, can actually undermine the entire business. Kind of fascinating if you think about it.
A person in charge may be so afraid to lose an employee they fail to see that the very efforts to keep them may derail the success of the business.
I don't know if adhering to these principles will make microservices better.
Then just release a package and reference it whenever there is a change to the target.
Should work well in either a monolith or micro service.