Staging is dead: The rise of preview environments
withcoherence.com
withcoherence.com
They way I've set things up at my current company is this:
When a developer creates a pull request an ephemeral (preview) environment is generated.
When that code is merged it automatically gets deployed first to our staging environment and then immediately to production.
Staging gets a sanitized copy of production's data every night so both the code and the data are close mirrors of production and that's what developers (and those preview environments) integrate with when they're calling services other than their own.
If you were to get rid of staging would you stand up your entire backend every time you made a preview environment? Does that include replicating databases or would you just use seed data that may or may not resemble what's in production?
It seems like the solution they're pitching works great if you have a small team and only a few services/sites.
What I've seen done is that the staging environment is skipped, but instead concentric releases are performed on the customers. So released code goes first to a small group of users, and then widens from there if everything looks good.
If you really want to invest in extensive testing, you can do both! Hell, _and_ you can have a load testing environment as well.
Agree that it would be harder to implement with customers sharing a schema, but you can still have different release pods
And like you, we have both "preview environments" and staging and for the exact reason you've highlight. I'm rather skeptical of the company behind this product now since they don't seem to understand the difference and really overselling something that's quite common.
Before: Your developers have a shared Linux box that serves the preview environment. Your ruby app has a tiny loader in front of it that loads the app code from a specific directory based on subdomain.
After: You need to route subdomains to different kubernetes containers, and handle deploying new containers when the developer opens a PR (because it's much more work than just copying files onto a server), and you need to handle destroying old containers (because it's much more work than just deleting files on a server).
Not to mention, usually kubernetes is managed by "the ops" who don't know anything about the app. Much more difficult to interface with them when it's not "oh one server, and the developers figure it out, and if they royally fuck it up we'll rescue them".
Routing domains to a container shouldn’t be significantly different to doing it in production.
This isn't it, and I think it's pretty much an impossible problem to solve without an entirely new engineering culture around testing.
How do you duplicate the behavior of the 10 direct data dependencies of your app and the constellation of transitive microservice dependencies? Wingman, Galactus, etc. At a company with 1000+ microservices?
And how do you populate it with actual data so you can log in and perform actions? Every single microservice probably needs test data, and it would take a lot of engineering forethought to populate the entire graph.
It gets gnarly. I've been on teams at the center of it all - user accounts and login, representation of the core business entities that the rest of the company is built upon. We couldn't even solve the problem for ourselves. There were so many permutations of account states, creating a dedicated API to make test data deviates from a first class API. So how do you coordinate with downstream services to skip IDV (matching SSN), compromised password checks, etc.
Perhaps you only bring online a subset of services and fake out the rest. At what level can you fake stuff out? And how do you ensure that the test data works with whatever subset you don't fake? If a single service responds in an unexpected way, it could break or corrupt the state you're relying upon to test.
Maybe "preview" works for frontend. Backend, not so much. This is such a massive problem, and I'm only scratching the surface with my description of the kinds of issues you have to solve.
At a minimum, the product owner should be able to try out the new feature in the UI without having to deal with a staging environment.
Yeah this is the surprising thing to me - I've worked with otherwise bright people but whenever we've introduced services/microservices, it seems the default reaction is to just make a "distributed monolith" with circular dependencies and tight coupling.
I'm only half joking that 90% of my "architect" role is just point out we shouldn't have circular dependencies.
…and with added network calls!
> Microservices. grug wonder why big brain take hardest problem, factoring system correctly, and introduce network call too. seem very confusing to grug
(Source: https://grugbrain.dev/)
I'll try to simplify what I mean by that but something akin to having Services A, B, and C.
Each of A, B, and C need their own definitions to define how they need to be spun up as well as a definition on how to populate their data. That data population, say through RDS, needs to be spin up RDS instances and make them ready and available whenever someone creates a corresponding ephemeral version of a service. Toss the database afterwards.
In the above case, assuming you needed a full graph, you would spin up the three services as well as attach each service to its corresponding database. This approach works well when the number of services remains small.
Once that number starts growing, it may not be cost effective to spin everything up. Say service N is needed, but you really only need to read from it. In this case, you might have a dedicated staging environment for N and create a dependency graph that says "in my ephemeral environment, create new instances of A, B, and C with their databases, but also allow network requests to make their way to N."
That might be over simplifying, but I've been working in this space for the past few years and have been thinking about this exact issue recently
I think what made it quite easy for us is that we maintain a staging environment. Staging is constantly tested by QA simulating our workflows so non-production data is being constantly generated. When we spin up our preview environments, we use snapshots from the staging systems to seed the preview environments.
Your point on data though really gets to the crux of the problem. Deploying the services is trivial these days with k8s, Cloud Formation, Terraform, Docker, etc. but the data part is hard.
I'd love to learn more about the challenges you have with data. My company essentially provides data-anonymization-as-a-service and we integrate with ephemeral environments, but we could probably improve the integrations.
For those, like me, who had to think for a moment where they had heard that reference: https://www.youtube.com/watch?v=y8OnoxKotPQ
All of our codebases have some level of automated testing that runs when the PR is created. That could be unit tests of a single function or Playwright tests which exercise an UI flow. What those tests are dependent on the type of software and the team building it.
The preview environment links that are generated when the PR is open are given to QA, Design, and Product to take a look at. Not every one of those roles review every PR, it really depends on the task. But if their feedback is required, they'll give it at this stage.
Probably worth mentioning at this point that we have separate repos for each web application and service. It's service oriented but not microservices, services at our company encapsulate a relatively large domain and it's similar with web applications, right now each web application has its own subdomain.
So PR gets merged, the pipeline will deploy that code to staging. At this point the service owners have the option to define a set of integration tests, not every repo has them but it is an option to run them. If those tests fail we do a rollback. If those tests pass (or if they didn't have any defined) then a second deployment to production is triggered.
All of our deploys are blue-green, so no downtime except for the occasional blocking database migration in which we have to schedule downtime. User facing features on the web applications are all required to be released first behind a feature flag. For those there will be several releases behind a feature flag, then we'll flip the flag first in staging. We'll have Product and QA do a complete run through and make sure they're okay with everything, then the "real" release happen when we flip the feature flag in production (no deployment needed).
Every night we clone the production database, sanitize it (swapping out things like names, phone numbers, etc.) then point staging to that clone. Then we run an end to end regression test that goes through all of the happy path flows for all of our apps.
Quite the importance over a constant wave of changing code on development in certain projects.
Staging seems like a more code locked stable area for all previews that can be polished via quality practices and keen eyes.
Now it seems a little fleeting to me as well to have the shock of previews enter into this instead.
Let the paint dry adequately before taking it outside.
The typical workflow is to anonymize production data, snapshot it, and make it available to preview and testing environments as replicas. It's pretty fast.
Do you have any pointers?
The obvious problem here is that this approach is far too expensive for any org that isn’t a tiny startup whose production system fits on a half dozen hosts. Consider a column store DB that’s configured in a multi region manner. Is every engineer bringing that up in their preview environment? If they aren’t, it isn’t faithful to a critical production performance constraint. If they are, the company is probably paying twice the cloud costs of their competitors at least, and they don’t get the capacity planning benefits of a fixed staging deployment.
So once you've validated everything, you maybe stand up the prod-scale stack for less than an hour to run your at-scale tests, then tear it down again. Definitely not doubling your cost because you're not keeping anything around, and you're only using it when needed. Most places in my experience actually don't do this, and instead have a copy stood up somewhere just wasting money as you've described.
Your points are all correct of course, but then that wouldn’t be a preview environment, would it? It would be shared database state in a database staging environment.
Which, for what it’s worth, is also how I have seen this problem managed when it came up.
As for costs, maybe true, but it could also take a lot longer than an hour to bring up a large deployment. And if every engineer is doing this for one hour a week, and you have a few dozen engineers, there’s a 2x cost increase with a tougher capacity planning problem since your load is now tied to your hiring plans.
Everywhere I've worked either has spin up environments, a developer issued test environment or various lower testing environments where we can deploy individual features for testing and QA.
But more to the point everywhere I've worked has also combined these lower level testing environments with some kind of pre-production environment which features more data (typically a clone of prod). On this environment often a further layer of tests are ran and new features are manually reviewed one last time before being deployed to production.
Where does a "staging" environment fit in this? Unless I'm miss understanding the article seems to suggest code goes, dev -> staging -> prod, but aside from one developer role I had in the early 00s it's always been, dev -> test env -> pre-prod -> prod.
Is the article suggesting you can just do away with pre-prod environments for "preview" environments? And why would you even want to do that? As I understand it one of the features of a pre-prod environment is that is persists and isn't a clean slate each time. Any mess that can accumulate in prod can accumulate in a pre-prod environment.
I guess I'm not following what this article is advocating for.
Sounds like staging to me...
Which it absolutely should. If the idea is to go from your local environment with test data, to some staging or pre-prod environment in the cloud with (potentially) prod data. Then that cloud environment (regardless if it's called staging or pre-prod, or if it houses prod data or not), should be an exact replication of your production environment. Down to every detail.
Arguably you should need compliance on both staging and prod, but you do the anonymisation to reduce the risk of exposure from (less tested) code in staging.
I’ve always viewed staging as the environment that gets prod data (maybe anonymised), but has ideally no exposure to actually affecting prod, rather than a completely fake environment.
> dev -> test env -> pre-prod -> prod
You're just using a different word and adding one additional environment. You swapped out the word staging and replaced it with pre-prod.
A preview environment is an ephemeral environment for testing a single piece of unreleased work. Framing it according to the terms you've used, the article is advocating for replacing "dev -> test env -> pre-prod -> production" with "dev -> test env -> production".
In that case what they're advocating for seems like a really bad idea. One nice thing about having staging / pre-prod environment is that it can replicate prod architecture, content and config (for the most part anyway) which allows for things like performance checks to be ran after code has been tested for functionality in lower environments. Something I've also seen is that prod websites typically load various analytic and marketing scripts which might not be included on dev environments. Even if those scripts are not maintained by the dev team, you probably still want to be checking any code being deployed isn't going to break an analytic script because it's expecting links to have a certain class, etc.
The usual staging and/or preproduction environments would be the preview of their own branches.
It sounds nice. The only problems I can see are:
1) The cost, because each one of those instances could have its own queues, databases, connections with third party APIs, etc.
2) Seeding the database with a meaningful amount of data. I often see preproduction database built by testers month by month after a minimal seeding. Each instance here has to be seeded with all data plus what's required to demo the new feature. That must be saved before destroying the database, reused for further branches, merged with the data for other features.
Depending on the application architecture, you can spin up database containers and seed them with replicas of the anonymized data.
it's the place where more biz-facing people can tell you the color's off by like half a shade just before a big release :)
We've been experimenting with a pretty reliable way to do this in-house, though. The key is that a single Kubernetes namespace per preview app iteration contains everything your preview app needs. We have a Github action, triggered by a "preview" label on a PR and any subsequent commits, that:
- builds and pushes a docker image with a dedicated tag
- within the specific preview namespace named after the PR ID, spins up and seeds with test data (in our case, a sanitized subset of production) a dedicated database statefulset
- runs the Helm chart that we use for production, but against this specific preview namespace, with this specific image tag, and overriding our normal ingress domain with one specific to this preview app
Then when the label is removed, we drop the namespace and reclaim all resources. "Rebasing" off later data, or if reviving a stale PR, is as simple as removing and re-adding the label.
The biggest issue we had with preview/review environments was resources not being cleaned up properly after destroying the environment (along with the associated costs), but that was more an indictment of our IaC code than anything else. It also didn't eliminate staged environments for things like the non-standard DB that we were using. As long as the changes on the DB side were net-new, it wasn't really an issue, though.
The technical implementation part can be interesting to readers and potential engineers who will use the platform and is worthwhile to have available, but this content is purely meant to cast a wide net.
We still have a sandbox instance, and that is hooked up to the rest of the stack.
It works pretty well if you need to iterate on something quickly that you are unable to do with a test.
Going from static staging envs to ephemeral envs is a lot of work for non-startups with traditional release processes.
Data is an example of a challenging hurdle. It is (somewhat) straightforward to Terraform production and modularize it to make it repeatable. But what do you do when your most current tables have customer PII or other sensitive data in them and migrations are done manually during release? Now you need to audit the entire database for fields where PII might exist so that automation can be written to dump those databases and sanitize that data.
Then there's scavenging the envs. Ephemeral envs can cost A LOT of money. Many places also do not apply as rigorous reporting on resourcing in those envs as they do in production. So what happens when you get a new CEO who wants to cut infra cost by an insanely high figure and you start with the preview envs, but you have no idea which are for active PRs, which were created out of band, which are prod-in-all-but-name, etc?
It's definitely worth the work, as I believe that ephemeral envs lead to safer, faster releases, but getting leadership to buy in and invest capital is often a requirement.
IMO the REAL challenge is creating quality, ephemeral LOCAL environments. So many apps need a crap ton of infra to stand up for no other reason than "we use a crap ton of infra to build and test our app". Like, if I NEED to create a VPC to run an app as a Lambda function because the author tests by shipping directly to Lambda from their IDE or whatever, that's a waste of money and productivity. No reason why we can't run that in Docker locally and mock external dependencies as needed.
There are plenty of automated and technology driven ways to assert the quality of a system, which are far more reliable, faster, widespread (across more of the system) than preview envs. A lot of testing involves creating chains of logic to make resultant assertions. If A works, and B works when A works, then B works.
And things like compile time checkers (eg types), linters, and TLA+ so I've heard, are examples of things that are better than humans, rapid once setup, and in my experience give massive ROI.
I had a manager the other day say every PR needed a manual "smoke" test before release. I challenged him why would I spend 15-60 minutes manual testing one PR when I could spend that time writing tests that check _every PR_ ? He essentially told me to just fall in line and do as I was told. (No rational rebuttal)
I made the mistake of bringing preview environments into a small company and quickly discovered I was pushing responsibility for reconciling different streams of work onto the non-technical people.
Traditional staging environments that sit as the stage before production are problematic because they become a living thing that has to be maintained, but there’s a middle ground between preview environments and traditional staging environments: one staging environment that is ephemeral specifically for the next release, with the ability to spin up more if a situation necessitates it (e.g: a long term project that needs to be tested while other streams of work continue).
Having one gate between is probably okay (test/demo), but more means there's more chance of breaking things when trying to merge and reason with ever broadening code bases. Especially with smaller teams.
Yes, this all means production may break, but it also means the lesson is learned, devs get more careful, and testing evolves more rapidly out of necessity or pain.
Its great to see next-gen infrastructure tools support this, both from the article and also from infrastructure products like planetscale and railway.
In addition to this you can’t always test against prod data, especially when that data is classified or otherwise sensitive. Or you simply respect the privacy of your users. (Also you might not want to run staging code against prod for obvious reasons.)
Netlify and Vercel made preview environments a standard part of the workflow for all frontend developers. my usual soundbite on this stuff is we need to treat our envs like cattle, not pets, to make up and discard as casually as a git branch.
i think there are a number of startups (eg Railway, Brev) are all feeling their way towards this future. my one wish is that someone make a credible industry standard for environment branching that everyone can build toward, including 3rd party vendors (eg Planetscale/Neon) and open source multi-service frameworks (eg Temporal, Airbyte where i work).
You can take my localhost from my cold dead hands :p
I could get behind the death of staging tho.
Not every product is the same, and there's no reason for a single deployment approach to the problem. These different types of deployments are just tools to be picked and chosen from to deliver the best solution for a given problem. This article made Coherence look not particularly savvy, and I'm not sure I'd trust them as an expert partner after reading this. If anyone from Coherence is reading this, you can talk to experts in deployments by talking about how complementary preview environments can help solve problems in staging while not handwaving over the necessity of pre-production environments like you did in the Traditional staging bottlenecks section.
The points of traditional dev and qa environments are to lower the risk of unproven code, but the shared nature often causes more problems than it solves. Ad-hoc environments provides the opportunities to better utilize infrastructure, better test changes in isolation, and have one granular purpose per environment.
Another point is that this doesn't mean shared-nothing. You don't have to stand up expensive infrastructure every time if that infrastructure should be indistinguishable between environments. You can build smarts into how it's done such that multiple databases go onto the same DB server, or you reuse instances until something meaningfully changes in an isolated env you need (e.g. by looking for migration version differences). All of this is much more tenable now that infrastructure is so much more easily programmable and declarative.
And all of this works at multiple scales. Large horizontally-scalable databases don't have to be deployed at production scale on your machine. Do that in a load test env that is a last stop before prod, if you need to (e.g. if you don't trust how well you scale with # shards). Add in options from tools/concepts like teleport, and you can potentially extend this sharing to local dev environments for when things no longer fit on a dev laptop.
My experience despite having 7ish years with k8s is that people on k8s still have their dev, qa, prod clusters and dogmatically stick with that despite the fact that there are definitely better (and now simple) ways, and you really get better economics if you can binpack in the same set of machines and prioritize work appropriate to its classification.
How about mobile apps?
How about 3rd party integrations? And integration with legacy systems?
How about systems where there's non-trivial data that needs to be set up?
...there you go.
To first address the idea that preview environments are trivial to build or are already solved, generally we think that’s not true except in a “rsync=dropbox” kind of way. On top of what we’ve already built at Coherence, which represents a considerable investment of resources (especially when compared to what’s available to software teams <30 heads), there’s still so much more to do before we are even close to 10x better than what teams do today with PaaS or Infra-as-code.
Staging and other static pre-prod still has a place in the SDLC, the rise of previews doesn’t mean that these have no place. If your app has static integrations with 3rd parties, special data that needs to be manually managed, or integration with other workflows/teams/customers that require manual integration - staging/pre-prod has a huge place. In our case, Coherence supports static environments just as easily and fully as it supports previews (as do many other devops automation tools).
To the point above about tons of work left to do, one of the biggest pieces is helping with data generation/automation for new environments. Integration with data anonymoization tools to make moving data from prod possible, data generation/seeding tools that can introspect a schema and generate realistic data, and a library of snapshotted datasets from other envs are all things that some teams do, but which should be available to everyone. Features like load testing and security scanning for ephemeral environments are adjacent areas of opportunity.
There is absolutely no question that sprawling microservices make automating previews more complex, as do bigger teams. And the concerns around cloud cost are 100% valid. On cost, Coherence automatically does things like use spot instances and pause inactive environments. In practice, cost is not the primary issue with usability. Complex serverless systems often have the most trouble with dependencies across 100’s of functions and cloud-provided resources, and to us that only means that these kinds of systems will benefit the most as preview automation scales to encompass them (it still doesn’t do a good enough job today). That said, for a lot of real-world teams, there are cheap productivity benefits on the table if they’d invest in better developer experience.
Alongside reproducible cloud-hosted dev envs, automated pipelines, and ephemeral environments, feature flags and canary releases are also relevant best practices for release management. Coherence doesn’t really help with either of those at this point, but great tools exist for lots of language/framework ecosystems already.
We’re still a new product/company and we love hearing from users - feel free to ping hn@withcoherence.com with any questions/comments!
Do not walk; run away from that trashcan bonfire.
As opposed to the last place I worked, where there were 2-3 month+ long branches alive in various stages of features/testing/bugfixes. THAT was painful. Give me trunk to production with a small team over that mess any day.
> Give me trunk to production
Okay, but three times a day?
Three test builds off the trunk in a day, I can swallow.