Modern Applications at AWS
allthingsdistributed.com
allthingsdistributed.com
Unfortunately I see a lot of mid and senior devs try to propose these architectures during interviews and when asked questions on integration testing, local environment testing, enforcing contracts, breaking changes, logical refactoring a lot of it falls apart.
There's a lot of pressure for those guys to deliver these sort of highly complex systems and all the DevOps that surround it when they enter a new company or project. Rarely is the project scope or team size considered and huge amounts of time are wasted implementing the bones since they skip the whole monolith stage.
Precious few companies are at the size where they can keep an entire Dev team working on a shopping cart component in perpetuity. For most it's just something that gets worked on for a sprint or two.
AWS made a fundamental assumption in this that monoliths and big balls of mud are the same thing. Monoliths can and should be architected with internal logical separation of concerns and loose coupling. Domain driven design helps achieve that. The other assumption that microservices is a fix to this isn't, because the same poor design can be written but now with a network in between everything.
I think this advice is worth a read, but it's not law. The most pragmatic approach to microservices is to start with a monolith and shave off the intensive parts to separate services as you need to.
I've untangled more than a couple small eng teams (~ < 20 devs) that were trying to do microservices using this recipe on aws and just horrendously failing at it.
A well-architeched monolith will really go a long way - generally perhaps easily up to 100 or more devs before thinking about microservices.
- If it isn't something you could plausibly imagine using as a 3rd party service (emailer/sms/push messages, authentication/authorization, payments, image processing, et c.) probably don't make it a microservice.
- If you don't have at least two applications at least in development that need the same service/functionality, probably don't make a microservice.
- If the rest of your app will completely fall over if this service fails, probably don't make it a microservice.
- Do write anything that resembles the above, but that you don't actually need to make a microservice yet, as a library with totally decoupled deps from the rest of your program.
I've seen that attempted several times but I've never seen it succeed even after several years of shaving things off here and there.
The problem with this approach is the same problem that causes all technical debt -- cleaning up technical debt takes time away from developing new features, which is the company's highest priority. Adding more engineers to the team to help tackle debt won't help because management will see that as an opportunity to get more new features.
Now that the dust has settled, it's become clear that the real advantage of microservices is to decouple teams.
Shared codebases become harder to work with as the number of engineers increases, and that hurts developer productivity.
Don't build a monolithic service for the whole company -- have every team manage their own services, on their own development and release schedule, and their own production scale. Their dependencies on other teams should be based on APIs, not code sharing. Whether an individual team chooses a monolith or microservices doesn't really matter.
Monoliths shared across dozens of people are operational nightmares because there’s not a strong sense of operational ownership, so you often end up with a bunch of process and bureaucratic overhead to compensate.
Your advice to start with a monolith is too simplistic. If you know how your service will scale and your engineering teams will grow, then you have enough information a priori to avoid the waste of attempting to split apart a monolith later.
Then again, I was chuckling in the a similar way a few years ago when not using mongodb et al made me not "web scale".
When I worked at AWS, it used to take us months to change the position of a button in the UI. Things that our competitors did in hours and days, we did in months and years.
I'd say that as a general heuristic, if you want "to increase agility and innovation speed" do the opposite of what Amazon does.
For people who don't want to go to Twitter since they are here, Tweet is about that AWS changed the limit of ElastiCache names from 20 characters to 40 and author says about it: "If you can make a change like this to your SaaS without any fuss, you can probably outcompete AWS on innovation pace and agility. For AWS, this is a significant accomplishment worthy of the “2019 What’s New” page"
It will be interesting to see if that happens soon then. I know IBM and the many other Openshift partners are chasing this. Will be interesting to see if they can deliver.
My own view - do not underestimate the depth of services and innovation AWS and GCP can bring to bear. I'm just not hearing folks go - oh, this provider let me use a longer name for my elasticache cluster. Instead its, why shouldn't we just use AWS or GCP.
The decision making processes are influenced by different cultures: Debian aims for a very stable system with no breakages at least within a major version, Windows aims for close to endless backwards compatibility (hence the bloat), while Apple designs its systems with more of a "move fast" attitude (hence the quick proliferation of new features).
Not being smarty pants here; It truly boggles my mind how fast they move.
It took a VERY long time for ELB(and thus autoscale/cloudformation) to get the ability to drain connections. This was a show stopper for most serious players on anything customer facing but unless you were, or knew somebody who was, spending close to 500k/mo you would never be able to get a sense of timeframe out of AWS for addressing this.
ALB launched without the ability to route requests based on sub-domain even though this was the only direct upgrade path for most people who had been building SOAs on top of ELB and Route53.. Took way too long to get that added in and of course with no public acknowledgement of time frame.
Now that I think about it.. AWS is NOT very agile at all because they barely include a small fraction of their stakeholders in their product design and iteration process. You'd have better luck trying to squeeze water out of a moon rock than getting AWS to comment on roadmap for a service.
Hell, I'm still waiting for AWS to offer an upgrade path to PostgreSQL 10 from October 2017.
So if AWS's attitude lead to its console, I think they're doing great.
That’s a pretty vague and useless heuristic. Since you worked at AWS, maybe you could answer why AWS moved slowly on releasing features for existing services? An Amazon-y “five whys” analysis would probably yield a lot more specific and useful heuristics.
From what I understand, the culture in a large company varies from team to team. Was it your experience that all the teams were that slow?
Are you claiming the whole of AWS is that way or that since your team was that way we should disregard any advice about AWS?
If I invert all the advice from the post, would that be better as a development practice? A large monolith, backed by an off the shelf database where there is a central security team looking after every line of code?
Customers come first in AWS culture. With a massive user base, they need to care about not breaking existing customer applications/integrations.
If AWS did not evolve its software processes, they would have already fallen apart by now. The teams move reasonably well for the size of a big company.
As a customer of AWS, it never seemed that way.
I ran an enterprise-backed startup. We built way faster than our enterprise backers when it was our own name. If something carried an association or backing of the enterprise, it took at least 10 times longer. They had a reputation to uphold and needed to ensure everything reflected that.
I’ve noticed that software quality never seems to enter into conversations about scaling. Do code hygenie, writing clean code, and good documentation not matter when you’re moving fast and growing like crazy? I’ve noticed the approach seems to be to focus on operations rather than the way the code is written. Use CI, monitoring, organize teams around services, etc. From my friends in the bay however, I hear the code itself is mostly written and band aided over and over again with refactoring and rewriting seen as “wasting time” when you could be writing more band aids or improving operations.
Does the band aid and if it’s not broke don’t fix it and do exactly what you need to fix it approaches really yield the same or better results as approaching software as a craftsmen? How do you convince an AWS engineer that these things are important when what’s in prod is working and a rewrite just seems like pot stirring? Or a refactor isn’t critical work to add more features?
Moving to autoscaling is really more of a cultural problem than a technical one. If teams are used to being able to reserve excess capacity they will be wary of autoscaling -- what if they try to scale but there is not capacity to do so because other teams used it up? People need to trust that the system will give them the capacity they need or they will continue to overprovision.
Don’t take the path of cloud based development if you are aiming to simplify.