A relevant Twitter thread:
https://twitter.com/NeuroStats/status/1192679554306887681At the risk of projecting, this has the hallmark of bad experimental design. The best experiments are designed to determine which of many theories better account for what we observe.
(When I write "you" or "your" below, I don't mean YOU specifically, but anyone designing the kind of experiment you describe.)
One model of gravity says the postition/time curve of a ball dropped from a height should look like X. Another model of gravty says it should look like Y.
You drop many balls, plot their position/time, and see which of the two models' curves match what you observe. The goal isn't to get the curve; the goal is to decide which model is a better picture of our universe. If the plotted curve looks kinda-sorta like X but NOTHING like Y, you've at least learned that Y is not a good model.
What models/theories of customer behavior were your experiments designed to distinguish between? My guess is "none" because someone thinking about the problem scientifically would start with a single experiment whose results are maximally dispositive and go from there. They wouldn't spend a bunch of time up-front designing 12 distinct experiments.
So it wasn't really an experiment in the scientific sense, but rather a kind of random optimization exercise: do 12 somewhat-less-than-random things and see which, if any, improve the metrics we care about.
Random observations aren't bad, but you'd do them when you're trying to build a model, not when you're trying to determine to what extent a model corresponds with reality.
For example, are there any dimensions along which the 12 variants ARE distinguishable from one another? That might point the way to learning something interesting and actionable about your customers.
Did the team treat the random algorithm as the control? Well, if you believe some of your customers are engaged by novelty then maybe random is maximally novel (or at least equivalently novel), and so it's not really a control.
What about negative experiments, i.e., recommendations your current model would predict have a NEGATIVE impact on your KPIs? If those experiments DON't produce a negative impact then you've learned that some combination of the following is the case:
1. The current customer model is inaccurate
2. The model is accurate but the KPIs don't measure what you believe they do (test validity)
3. The KPIs measure what you believe they do but the instrumentation is broken
Some examples of NEGATIVE experiments:
What if you always recommend a video that consists of nothing but 90 minutes of static?
What if you always recommend the video a user just watched?
What if you recommend the Nth prior video a user watched, creating a recommendation cycle?
Imagine if THOSE experiments didn't impact the KPIs, either. In that universe, you'd expect the outcome you observed with your 12 ML experiments.
In fact, after observering 12 distinct ML models give indistingiushable results, I'd be seriously wondering if my analytics infrastructure was broken and/or whether KPIs measured what we thought they did.