As you might expect, it wasn't really the fault of feature flags but rather indicative of a process problem. We'd either have to prod the team or person in charge of it, verify that work was actively pushing things forward to ship, and so on.
At various times we'd have, for lack of a better phrase, the "No (wo)man", who would come in and say "no" to a lot of things. One of the best ways to achieve this is to see that nothing has happened on a feature flag for awhile and then send a pull request to remove the feature flag (and the feature entirely). This got people out of the woodwork who said waiiittttt a minute let me just finish that up, or if no one vehemently disagreed with the pull then you could actually just remove it entirely. But it did take some explicit reflection on whether the flags were defensible to remain flags.
The only thing you have to remember is: Its either going to production, or its getting dropped.
There is an expiry date built into the feature flags framework (90 days), and either they need to be explicitly cleaned up and deprecated, or extended with a valid reason attached.
Also, we code review every commit. If any diff is coding around/through a feature flag, the first question is always: "Can you clean up that experiment first?"
* Always be developing against the current running features. No brainer.
* Design things so that they integrate feature flags, not work around them. This usually means pushing feature flag determination to more generic/common code.
* Separate backend/frontend changes into separate feature flags when possible. Turn on backend changes early and often to better measure your feature's impact.
* Give individual features their own flag, but also have a global flag that manages the entire experience. This makes it easier to manage your gradual dial up as well as shut off problematic features that would otherwise mess up the launch.
* Be diligent about removing feature flags once they're turned on. Schedule it into sprint time, reward teams that remove them, make it a management mandate, whatever. Just get rid of them once they're no longer needed.
* Invest in monitoring around your services that (ideally) can correlate failures with features. you should turn on features over the course of a few hours/days to mitigate customer impact in the event of failures and gain data about performance at 50/50.
I think the answer to your specific question of "testing every combination" is that you can't, easily. But by keeping the number of feature flags that are inactive low (< 150 is very liberal) for a given service, having everyone develop against the current running features + dev overrides, and using gradual dial up with integrated monitoring to catch poor interactions when the impact is small, you'll have mitigated a lot of your concerns.Flags were there because of two main reasons:
* Our app was deployed more frequently than a service we depended on (e.g. every commit for app, nightly for service in staging, every 2 weeks for service in prod, etc)
* We wanted to have it rolled out to a subset of users for testing
We did not run into exponential trees of combinations because we rarely had two different flags interacting. Maybe it was a happy accident of trying to make work parallelizable or maybe because our feature flags never lasted more than a month or so.
The code was intentionally dumb and the flags were stored directly in the source code (not in a database table or another config file or something). Simple, stupid calls to `FeatureFlipper.isFooReportEnabled()`. We did not test this class because each method was a simple boolean check that the current user appeared in a list or ENV != prod.
We stubbed out `FeatureFlipper` when using it throughout the app. Stub the feature to be enabled, check behavior. Repeat for disabled.
Most of the features were simply hidden at the view level. For users, if there is no button, it doesnt exist. We didn't particularly care if the user would "guess" the url of a feature flagged page -- not worth the effort.
Doing a "full rollout" of the feature was a non-event. Just delete the method on FeatureFlipper and go fix the compile errors :)
Why do you have so many feature flags at a single time? Do you have a team of 20-40 people or more? Are your features taking months to create? Are you trying to create a feature flag for every new line of code before deploying it?
A lot of times code can be deployed without using a feature flag, because it's small enough, maybe you are overlooking those cases?
>Also please re-read my post carefully. The details are there
Clearly there is something wrong with your project that you aren't mentioning. Plenty of other projects do flags without problems. Your team has issues.