This problem is amplified further because flags can come in so many different shapes and sizes:
- Rollout: binary on/off, percentage rollout, beta group rollout, single user rollout, etc.
- Control plane: eg. emergency shut off valve for the caching layer, switch that controls routing, etc.
- Multi-valued or numeric: ie. not just true/false, typically combined with A/B testing
Conflating these can lead to extremely messy code. Each of these cases deserves its own special type of treatment.
Rollout flags should have a self-destruct or mechanism that strongly encourages engineers to remove them as soon as they're no longer used.
Control plane stuff shouldn't really be mixed with feature flagging at all. It would be dangerous to remove them, and regularly vetting the states should be a necessary requirement for any type of disaster recovery assurance testing.
I really appreciate your call to represent these states as an enumeration. Capturing the flag states at the start of a request and then determining the proper bucket helps to make the problem more concrete. Engineers can then better reason about the explosion of states and the impact to control flow.
Ultimately, I think flagging is one of those "hard problems" that comes with the job territory. You don't want to throw the technique out entirely along with the bath water, because it has plenty of justifiable benefits. There are best practices and proper hygienic steps we can take to make the use of flags easier on ourselves.