Basically: Wrap every new feature in an Experiment Toggle. Deploy to production (once ready for 'non sandbox' testing). Turn on for test users, test. If good, start the rollout to experimental users. If after weeks the numbers worked out, turn it on for everyone...
...but if at any point unexpectedly bad things happened, just flip the feature toggle switch. Most of the time if "bad things" were happening in a relatively new feature, this would stop the panic and let analysis of the root cause of the problem happen without the pressure of everything being on fire.