Testing Microservices the sane way
medium.com
medium.com
This is honestly not that hard to get working, if you have Docker and good tools. For simple cases, "docker-compose" will work well enough. For complex cases, my employer open-sourced a tool for exactly this: http://cage.faraday.io/ This extends docker-compose with the idea of multiple pods, staging/production environments, etc.
For testing, we have plenty of unit tests, and we've been experimenting with pact https://github.com/pact-foundation/pact-js, which turns microservice consumer mocks into provider contracts.
But at the end of the day, it appears to be absolutely necessary to have end-to-end tests that run in a staging environment whenever a service is updated. Otherwise, something inevitably falls through the cracks. It's a nice idea to think that you can develop each microservice in isolation, but somebody ultimately has to watch site-wide quality.
There's also all the stuff that's tricky to automate - say you're using AWS Lambdas, step functions, SQS queues. Translating that across different developer envs is a challenging problem.
I'm very much on the monolith-first approach. There are cases for microservices for specific business functions when you have few engineers, but in general it's far more of a scaling layer for engineer count. Single-app workflows are typically much more productive compared to microservice workflows up until a certain point.
I'll admit, I'm not a fan of testing in production, especially when you're doing that testing against paying customers. I've just seen it go poorly too often, with the common result of lost customers. When you're B2C, it's bad since you're now having to acquire more customers not to grow, but to remain even; a death-knell for VC-backed startups. When you're B2B, the loss of customers signals a coming winter for the business. I've been through a couple of layoffs due entirely to lost customers.
Remember that your product is most likely a fungible asset - your service easily replaceable by a competitor - and if you piss off your customer base by allowing out blatant bugs, you will lose customers. The sheen of newness has worn off tech companies, and customers are not going to be as forgiving of their time being wasted as they once were.
Looking at her list of production tests, her list generally falls into 3 categories:
* instrumentation of code for observability * staged rollout of new software to observe issues on a subset of traffic/hosts against the rest of Prod * simulating failures in Prod to test against possible random failure scenarios
Of those, the first should have zero impact on your customers, the second should be making your customer experience better by reducing the blast radius of a bad deployment, and the third should be making your service more resilient to failures over time for your customers. Would you rather find out that you can't survive an availability zone outage when you can easily return that AZ back to healthy in a minute, or when you have to wait for recovery for an hour? Honestly I think its disrespectful to your customers to gamble on what might happen in the future rather than asserting that your service can handle the wide array of known failures that can happen in this world.
WRT the three points you bring up:
1) Instrumentation is a part of any good deploy, regardless of where, when, or why. Instrumentation is not a replacement for testing.
2) Staged rollouts are a good thing, but they don't prove the new object being rolled out is production ready. They can cordon off major failures, yes, but see my comment above about missing steps when it comes to major failures.
3) Not all companies can afford the Netflix model of having multiple fully redundant data centers at all times. I'd even hazard a guess that most companies can't really afford to triple (or more) their infrastructure costs in addition to the higher development and maintenance costs. Ultimately that's one thing that testing is good at: ensuring your disaster recovery plan works without honking off customers when you discover gaps in that plan.
None of these three strategies will replace pre-customer testing. At best full adherence to the policies (in the absence of end-to-end testing) will only limit the amount of damage major bugs can do. The question to me is: why are we OK with just limiting the damage of otherwise findable bugs?
http://www.draconianoverlord.com/2017/08/23/futility-of-cros...
My preferred solution, that I haven't had a chance to actually flush out at scale, so disclaimer/YMMV, is for all services to ship with stubs:
http://www.draconianoverlord.com/2013/04/13/services-should-...
Where the stub is an in-memory version of the service that its authors maintain (not the client), so you can achieve the proverbial "deploy all systems on your local machine", but since they're stubs, they're extremely quick to boot/reset/etc., and also with allowances (again since they are stubs) to let you set the per-test input data of any stub your system talks to.
I believe this would work best with homogenous/noun/REST-based services, e.g. all entities in your corporation have a strict/unified CRUD API, so then "integration" tests (e.g. the proposed stub/no-actual-wire call tests) can define their input data in terms of entities and be fairly oblivious about which stubs/systems those entities actually live in.
I was recently reading about Hypermedia HAL APIs[1] (TLDR: add metadata to API responses, helps w/ discoverability etc) which could conceivably play a role in solving this kind of problem.
You're doing it wrong. If these services are really so deeply entangled that you can't change and test them one at a time, they shouldn't be independent services. Merge them, or otherwise rethink your service boundaries.
In my apps I use fakes in dev and test mode, which makes development very fast and easy. A few tests run against the actual UAT environment, but these are skipped unless a command line flag is passed.
Mostly it's inspired from the article "Mocks and explicit contracts": http://blog.plataformatec.com.br/2015/10/mocks-and-explicit-...
1. The tests themselves are housed in a separate repository, so you can't update the tests alongside your service. This means every change to a service has to be backwards compatible. Hello, multi-step rollouts.
2. Our environments must be highly configurable, so that every permutation of versioned services can be integration-tested. This is forcing us to adopt an unnecessarily complex container orchestrator.
3. Service owners are not excited about setting up and contributing to yet another project. We end up with a lot of out-of-date integration tests. Lots of noise, if you ask me.
I think a combination of contract testing (e.g. using Swagger, Pact) with monitoring, canary deployments, and automatic rollbacks would be easier to maintain and just as effective at catching bugs.
It's interesting to see how things like Golang may well have evolved out of this problem. Calculating a golang app's dependencies is as trivial as a grep or two, so it's easy to know what tests to run when libraries change.
In my company, we had a legacy PHP system that seemed to me excessively layered and cut too thinly. In reality, it allowed to completely replace that system with Python and Java piecemeal, without ever stopping it.
Fitting together is an aspect that becomes easier when you decompose your app properly microservices or not. When a component has clear responsibilities, it's easy(-ier) to understand, modify, and replace. When a component is reasonably small and isolated, it's easy(-ier) to analyze and understand, even if it's written in an ancient language using dust-covered unmaintained libraries.
Most companies don't have specific needs. They think they do, but 90%+ of the companies out there are either building run of the mill CRUD apps or something barely more technically difficult.
Even teams that are building "run of the mill CRUD apps" need to be expert on the specific context and needs of their run of the mill CRUD app. If it truly is generic then it doesn't need to be built (and thus doesn't need to be tested or discussed on HN).
Also anecdotally, in my career, while I've seen lots of CRUD apps they haven't been the focus, the challenge, or the money maker for any of the businesses I've worked with/for. So I'll buy that 90% of companies don't need this advice because 90% (or more) of companies don't do software development at all. But for the ones that do, saying that 90% of them don't have specific needs rings hallow.
Do you have anything to back up the link you're claiming between not being technically difficult and not having specific needs? Interfacing with several dozen external services probably doesn't fit what you'd call "technically difficult", but making sure they don't break would count as a specific need.
What is your definition of "technically difficult", anyway? Is it something that links in the Tensorflow library? Something with separate transactional vs reporting databases that have to stay mostly in sync? Something that has to stay easy to update when regulations change? Something like I mentioned above, with multiple-dozen external integrations and contractual or legal penalties if they break? Is it things like the auto-level feature that hobby quadcopters have?
I have spent a good chunk of my life explaining companies why they don't need any special hand rolled platform - leaving hundreds of thousands of dollars on the table.
I guess parroting buzzwords is always more popular in this industry.
For situations where the true overall complexity is manageable, I think that making sure that the environment can be entirely replicated and bootstrapped easily is actually very good discipline for keeping complexity under control.
Edit: English