Dear Staging: We’re Done
devops.com
devops.com
If we have something fail in production, the first question we ask ourselves is: Why didn't we catch this in staging? 90% of the time, the answer is: "We didn't test if properly but thought we did."
Our rate of patching for defect is rare.
As far as the cost, our staging is run in hardware retired from production, with some upgrades (mostly RAM and SSDs). I'm not sure we've spent the authors $100K over 2 decades, let alone in a single year. But YMMV.
And are you really going to run your entire staging environment in spot instances? What happens on peak days?
https://www.ibm.com/blogs/cloud-computing/2019/03/05/20-perc...
(My recollection was they loved everything about AWS except the price tag, although creating a personal staging environment for me to debug their issue in would have only made a marginal difference to that total cost.)
That's not true at all. It is a common misconception that feature flags have no risk.
Common approach is to just plaster the codebase with "if feature_flag(x)", so you complicate dozens of files/modules/functions, make them harder to reason about and debug. And cleaning it up is commonly a mess.
Another approach is to duplicate code which frequently leaves one branch untested.
FWIW, I do like feature flags, but as a solution for a different problem: long-lived feature branches (forks) of the master. And even then the complexity they introduce is sometimes a burden, but overall it's worth it over alternative approaches.
And sure, you can start implementing stuff iteratively, in small chunks and perusing feature flags to hide bits of unfinished user flow, but the benefits you are getting there are from the former (iterative, small changes), rather than from feature flags. It is still no substitution for testing in a prod-like environment, at the very least that you are not breaking any of the integration places which are not covered by tests (and integration and system tests are slow, which always puts a boundary on how many you can have).
But isn't this case what staging originally tried to solve? "this big feature isn't ready for production but we need to regression test in a setup as similar to production as possible" ?
For the last 15 years I've been at teams using a "staging", it was simply a single production-like environment (well, my first project was exactly like production on slightly less beefy hardware with sanitized production DB import nightly) that gets the "trunk" (or "main") deployed to automatically (on my current team, it can be used to run non-production code as part of the deployment process which you bail out of).
I am pretty sure there is no one true definition of staging, so it's a spectrum rather than an exact way to do things.
Feature flags is about multiple branching. Staging doesn't really give you that in full, just by hacking the environment. It's then that you end up with environments that don't properly reflect future production and start missing stuff.
The author sounds like someone familiar with testing, that needs to learn about CI/CD. Sure, for trivial front end stuff, you can often get away with doing it in production. But without synchronization between environments, you end up in trouble and take on more risk, just for the luxury of treating your paying customers as beta-testers.
It's a valid opinion depending on what one is working on, but I bet it doesn't apply for most outside unimportant side-projects.
Implementing a flag takes fewer lines of code than the feature, and so while the odds are very high that the bug will be within a feature, there is a non zero chance that the bug is in the flag.
A bug might break a new feature. It might break the old behavior. It may turn the system into a latch (you can launch it darkly, turn it on, and then not be able to turn it off).
And then you can break the system while retiring the flag and making the new path permanent.
I have to admit that I often find it funny when some talk about unit tests being the end of all. No manual QA needed, no staging needed, tests passed and everything is good to roll to production.
Unit tests are important, no question but it’s still a piece of code which could have bugs or not enough scenarios covered.
https://dougseven.com/2014/04/17/knightmare-a-devops-caution...
A practice you can adopt which can prevent this: suppose you use JIRA (the practice works the same for any such system, I'm just using JIRA as an example), and for every new feature you must have a JIRA issue, e.g. FEAT-1234. Now you have the rule that the feature flag name must embed both (1) the year and quarter the feature was merged and (2) the JIRA issue. So as part of FEAT-1234 I create feature flag "2020q3-FEAT-1234-split-big-trades". And then you have a custom field in JIRA called "feature flag name", and you record "2020q3-FEAT-1234-split-big-trades" in that field for FEAT-1234. And then you need to have this as a standard check in processes such as code review and release management. It might even be useful to have a report which compares JIRA to the feature flag system to check this rule is being followed. If they'd adopted this practice, that particular bug would likely have been avoided. Adding a new feature, you would not be allowed to reuse an old feature flag from years ago, for the year and quarter and the JIRA ticket would all be different.
From services that suddenly require new ports to be open through VPNs (development happens without intercontinental VPN’s of course) — to commandline options not correctly propagating to our hardened variant of configs — to even latency or bandwidth issues and timeouts due to assumptions.
There’s a difficult balance to strike when it comes to ease of development and not letting issues (security or issues with distributed systems) propagate to the customer.
Staging is also the place where you can safely let the chaos monkey reign supreme without fear of losing customer data! It’s the last place to test the “fit” of your system.
Staging has a place, and if puff pieces like this cause the company as a whole to downplay funding for it then I’m going to be mighty displeased.
This sounds like more of a problem with the author's staging environment than a general issue with the concept of staging environments.
The idea is staging and production should mirror eachother as closely as possible. If your staging environment is wildly different, of course it will be less useful. But it's also your responsibility to maintain a useful-enough staging environment. Why not just say, "it's too much effort for me to maintain a separate, identical environment" rather than dismiss the concept entirely?
The longer you are here, the more weird shit you see, and the definition of “know” gets a lot murkier.
You aren’t going to “know” if it works in production, even if you put it into production. Which is something that only makes sense about the fourth time you encounter a piece of shipping code that is so fundamentally rotten that you are not sure how it ever came up with the right answer.
You don’t know if the whole car is working. You only know it’s moving the direction you wanted it to move and it hasn’t exploded, made horrible sounds, or smells. So it’s mostly working.
Every tool and process should be working to make you less afraid of trying things (like deploying to prod). Otherwise they’re just helping you build a bureaucracy instead of software.
There will always be known differences between staging and production, e.g. service names, networking.
The point about staging is that it lets you run risk-free tests on a set of things, in a way that give you confidence about the things that are the same as production. The fact that a set of things works in staging is never an absolute guarantee that it will work in production.
If you are in the position of needing to ship releases rapidly, and you can't run meaningful tests in the available time, then no, it's not useful.
Staging is supposed to be architecturally identical to prod, that doesn't mean it should be actually identical. In practice, this means that if your prod has 10 API servers that are load-balanced behind HAProxy or something, you can probably get away with only have 1-2 API servers in your stag environment, as long as they're still behind HAProxy. If your prod PostgreSQL primary has 64GB of RAM, in stag you can probably get away with 4GB (or less). The key here is to have the same architecture, but on smaller scale (because you're not actually trying to service your customers).
From there, if you have a staging env that costs "$100k", but you don't "infra as code" (with Ansible, K8s, Chef, Puppet, Docker, etc), you're doing it wrong. The difference between deploying to stag and deploying to prod should probably be a one-line change.
I think if you have more than one instance of a component in production, you really need to have more than one instance in staging. You can get a lot of issues with more than one node (e.g. session cookie stickiness in load balancers, clustering framework issues, replication issues) which you won't get with only one. I've seen before people get burnt by production issues which they can't reproduce in staging because it only has a single node for everything.
I don't think the article does a convincing job of arguing that releasing behind feature flags would be much easier to maintain than a staging deployment, though.
For making this consideration, it'd be worth seeing why the costs can't be lowered, why the effectiveness of staging can't be increased, or what cases might be difficult to handle with feature flags.
As others have said, staging is useful for certain types of changes that may be high-risk and not time-sensitive. For other changes that are time-sensitive and are low risk (like changing some incorrect/offensive copy), being able to bypass staging is also useful. I don't mean to disrespect the author but it feels like the article was designed to take an extreme stance for views. It even reads like clickbait. I bet this article leaves a bunch of people believing that using an intermediate testing environment can never be useful.
If you spend 1M on production, you might have a staging environment 1/10th of that size. Say 10 machines in a cluster instead of 100. But some technologies also have minimum implementation cost requirements per tier of usage. Like if you have a multiple data centers in production, you will need to test that in staging. You can’t just have 1 machine in a cluster, so your staging environment needs at least 3 to make a quorum. Therefore, to have 5 data centers in staging, you now need 15 machines... which is more than the base of 1/10th of production.
Issues like that can balloon staging costs if you try to properly implement it.
my preproduction uses a three machines setup so it can catch clustering problems, but the code itself can run just as well on a single machine interacting with a local single instance db, and devs run it like that.
dropdb $PGDATABASE || true
createdb -T template0 $PGDATABASE
pg_dump $DUMP_DATABASE | psql -q $PGDATABASEBe it an infra change or application. It should always be like:
Staging -> Canary -> Production
If you invest time in making the above pipeline automated it will cause less outages.
P.S we use feature flags but it won't solve many other challenges around infra and application changes.
Some examples: - Changing from Classic ELBs to ALBs (We care about latencies so ALBs served 1% traffic in the start) - Zero downtime single mysql master upgrade. (ProxySql) - New k8s controllers. - New metrics adapters.
Staging should be like a playground to test all edge cases before you go to canary and then to prod.
> Staging -> Canary -> Production
Ironically, the author says exactly this when she presents on it. Why she wrote this article in this fashion is beyond me.
Personally, I have worked in an environment where staging was a pain, full of cobwebs and issues and data mismatches. I sympathize because I know how painful it can be.
But if you respect your end-users, you cannot delegate the duties of a staging environment to them.
Than blaming staging env as a whole, I would blame archaic devops practices for SaaS.
Because realistically, something will go wrong sooner or later and it will cost the company (and its customers) a lot more money and reputation if it happens in the user-facing environment.
I still have concerns with deploying production upgrades behind flags due to the inherent issues surrounding things like database schema mismatches, terraform/infra upgrades that are not compatible bidirectionally etc. Has anyone worked on ways to manage these with feature flags?
If instead of converging to versioning, you want full-on branching that never merge, this sounds like very chaotic design. You'd then maybe want to design everything around feature isolation (see: UNIX philosophy), but comes with own caveats around new requirements and cross-cutting concerns.
Whether this is worth it, depends a lot on the particulars of your apps.
And where do you test if your feature flags work correctly with production data and environment if not staging? Yes yes, you might have tested in on your development environment, but if testing on staging doesn’t work because the environments are too different, testing on localdev or CI is even worse.
Everything I've worked on in the last few years would have been in far, far better shape if they had just started with a no-schema data management system like Bigtable and just adapted their backend code around its limitations.
Imagining you are safe because you are smoke-checking DML statements in staging is just begging for a week-long outage.
Heck, I don't need to imagine, I have seen it time and time again.
After you’re onboard with this philosophy then I’m ready to sell you the fact that you don’t need joins, indexes, or transactions. If a giant complicated service like gmail doesn’t need any of those things then very likely that your thing doesn’t need them either.
second: even if you don't have an explicit schema in the database, you still have to have a way to know which data goes in which column. just because it isn't written in a definition language doesn't mean the concept doesn't exist and doesn't require to be maintained as the application domain evolves.
The last time I tried a project with a no-schema DB, it was a huge PITA. It sucks having a function fail on one single document because it had a number instead of a string. I love flexibility, but it is very valuable to have the right constraints in the right places. I quickly learned that a schema is just that.
By "typed language" I suppose you mean "statically typed language". This isn't relevant because A) JSON types exist and B) I've never had a problem with unexpected types from user data in a dynamic language, not even in PHP.
> Otherwise, how is this not caught when it enters your system?
It should be caught, and it's very important that it be caught all of the time and not have any issues due to bugs or version mismatches. That's why the schema should live in the DB.
YMMV I guess.
Thats not to say that a staging env is bad, they can be very useful, if setup properly. It also might be the case that you really can't test in prod. (stuff that provides safety)
I am very keen on feature flags/magic routers. They have the advantage that you can test the inputs of the system with _real_ data. But that can also be a drawback.
I suspect this DevOps is working in loneliness, without enough communication with devs. Devs ought to be educated, it's part of DevOps job.
Literally their solution was "just make a product that costs $0 and test it against our production service".
To me, that was a very red flag.
You lovingly hand-provisioned your infra rather than writing infra as code. This makes it non-trivial to set up a stag environment, so you just don't.
with staging: keeping almost identical environment with production, syncing data if necessary to test load. can be considered as mostly Ops related work.
without staging: of course feature flags should be used when necessary, but if there is no staging environment we should create feature flag for almost every pull request, which in turn create a lot of legacy which should be deleted later. add nested feature flags into this and now you have problem of seeing which conditions are executed when multiple teams are working in parallel
We had some issues with Akamai WAF blocking things in PROD which only got caught when pushing to PROD so we added staging to Akamai now its closer to 1:1!
For example oracle e-business suite runs completely different with 900GB of prod data vs 100gb of staging data.
ASP.Net applications are slower talking to databases with 1600gb of MSSQL data in prod vs 40gb in staging vs 3gb in test etc.
DR performance different again because its scaled down version of PROD but has the same amount of data and traffic loads when it takes full production load.
Blog writer is taking the piss.
Not everything can be feature flagged. Not all tests can be automated. Not all features can be safely tested in production. You can't always roll-back code.
Scenarios where I've found staging environments to be invaluable:
- Features that make wide-scale changes to the application
- Changes that require user acceptance testing. Management/Customers might ask for Feature X, you have designers mock up what it'll look like, you might even build prototypes of it... but until they see it and play with it, they might not fully realise that what they actually asked for isn't what they actually intended.
- Integrating changes from multiple developers and seeing how they play together
- Testing that your feature flagging actually works
- Testing how the system responds to dumping large amounts of data/requests at it - a half-day outage in production might easily exceed the cost of a whole year of running a staging environment.
- Testing how the system behaves after data migrations/changes (things where there's no roll-back without losing data)
- Testing integrations with external systems.
- Detecting poor performance, memory leaks, deadlocks, and other stuff that isn't possible on a dev's machine
Testing code on your machine is fine as a smoke test, but your laptop probably has multiple times the capacity of a single node on production, for a single user. So your change that results in 1GB/sec of full table scans on the database might not even be noticed because hey, you have an NVME drive. Your tests also run for maybe a few minutes, but that staging instance might be up for days or hours, and deals with hundreds
I've picked up numerous issues because someone (I'm including myself here) committed some code that resulted in Staging going down or behaving poorly.
I've even overheard conversations that were along the lines of "Yeah, this feature isn't working right at the moment in testing, but that's because staging is running slow today It's fine, we can push this out to production and it'll be better there." if it wern't for me overhearing that, it'd have resulted in a multi-GB video file being sent for every visitor to the site, costing a fortune in CDN charges and still running like crap because it was a 30 second video.