For example, suppose you did a controlled study and found that the new autocomplete UI was presenting users with only "yes" options, X% of the time. Is there a threshold beyond which you would unlaunch or adjust the feature?
What if the new UI was actually changing user behaviour, causing them to answer "yes" more often than they otherwise would have? Is there a delta (positive or negative) beyond which you would unlaunch or adjust the feature? If you had to adjust it, how would you know you were adjusting it correctly?
What sort of telemetry would you need to put in the product so that you really would know if this was happening?
I'm curious if your team thought about these questions before launching, and what ideas came up.