Assuming the usual sqrt(n+1) error on a count, as n is low. Combining the uncertainties from 83/66 [0] gives 1.258 +- 0.231 (1sd). It's only just likely that it is an actual improvement, but there is about a 30% chance that the orignal was better.
[0] (((1+66^0.5)/66)^2+((1+83^0.5)/83)^2)^0.5x83.0/66
EDIT
If we also include the errors on the total counts and use the fraction 83.0/66/6362x6392, then the error is:
(((1+66^0.5)/66)^2+((1+83^0.5)/83)^2+((1+6362^0.5)/6362)^2+((1+6392^0.5)/6392)^2)^0.5x83.0/66/6362x6392
which shows an improvement of 1.264+-0.234. Stastically nout.
using Distributions
b_old = Beta(66+1, 6392-66+1)
b_yel = Beta(83+1, 6362-83+1)
N = 1000000
# Sample from both distributions, count the fraction of samples that are better
sum(rand(b_old, N) .> rand(b_yel, N)) / N
This yields 7.7% chance that the old one is better than "yellow". It's fascinating to see how we can get such different answers to a simple question.I did an A/B test on an older framework which didn't automate statistical significance at all, but the website was getting more than 2000-3000 orders per day, so after a single week we had enough data to determine that sales had increased by 36% (reduced the page's load time by almost half, changed the checkout to use Ajax, and a few other small changes.) without the need to quantify things. In fact, at the time, I didn't even know what "statistical significance" was... not that I know too much more about statistics now than I did then.
Anyway, in theory all of the exact p values, etc. matter, but in practice, the bottom line is all that matters, because the p values can change in a moment based on something I haven't factored in. That's where, at least for the time being, intuition still plays a great part in being actually correct, which is why we still have people with repeated successes.
So: probably better, but should have run it longer.
I wanted Optimizely to say it was 100% significant for a full week of it running before I ended the test, but the chart was interesting to me, because the conversion rate difference between the two remained the same for the entire test, rather than there being a specific period where "yellow" excelled.
That's not how statistical significance works...
https://www.optimizely.com/contact/
They'll need to reeducate their statisticians right away!
Years ago my team's statistician did a competitive review of various AB test apps, and reported various ways in which the UIs make statistically invalid statements to the user.
Optimizely is in a rough spot. People don't like having to think through experimental design, and they are really, really bad at reasoning about p-values. To try to fix the people part, they came out with the sequential stopping rule stuff (their "stats engine), but they never really published much justifying it. The other alternative would be to move the experiments into a Bayesian framework, but that has a lot of it's own problems. When they acquired Synference, that was one of the likely directions to take (along with offering bandits), but that didn't work out and those guys have since left.
Not that I'm trying to defend Optimizely (I'm not a huge fan, but for other reasons...).
I can't vouch for the quality either, but they did publish something about it[0] - that at least looks quite scientific. Happy to read any critique of course.
[0] http://pages.optimizely.com/rs/optimizely/images/stats_engin...
They are still having people make very fundamentally flawed assumptions about the data, which results in incorrect conclusions, and they are still not presenting the results in a way that people correctly interpret them. That being said, those are really hard to solve, and models that would try to correct for them would likely require a lot more data and be overly conservative for more people.
What are your reasons for disliking Optimizely?
Disclaimer: I do stats work at VWO, an Optimizely competitor.
(Also if you want to read our tech paper, here it is: https://cdn2.hubspot.net/hubfs/310840/VWO_SmartStats_technic... This describes our Bayesian approach, which we believe to be less likely to be wrongly interpreted by non-statisticians.)
Should have everything you would ever want to know about the method.
I agree with you that the problem of inference and interpretation between A/B data, algorithms, and the people who make decisions from them is a hard one and worth working on.
That said, I do think the two sources of error our stats engine addresses - repeatedly checking results, and cherry picking from many metrics and variations - did make progress in having folks correctly interpret A/B Tests. This did result in more conservative results, but the benefit was that the variations that do become significant are more trustworthy. I think this was absolutely the right tradeoff to make for our customers, and trustworthyness is a pretty important aspiration for stats/ML/data science in general.
Of course I did write the thing, so I'm not very impartial.
My issue with Optimizely are mainly how they essentially ditched us as (paying) customers. We were admittedly small-fish, but we were paying and were willing to pay more, but they switched to Enterprise-vs-Free without any middle grounds. Enterprise was way too expensive for us. Free didn't include essential features, so we were just stuck.
I ended up writing an open-source javascript A/B test client[0] (and recently also an AWS-lambda backend[1]), but it still has a way to go...
Optimizely has a bandit based 'traffic auto-allocation' feature in production on select enterprise plans [1]; bandits are excellent in a wide range of situations, and have many advantages, but like anything, have design parameters and there are some caveats you have to be aware of to make sure you are using them effectively.
On Frequentist and Bayesian: Optimizely's stats engine combines elements of both Frequentist and Bayesian statistics. They have a blog that tries to touch on this issue [2] But this is subtle stuff - and there are a lot of trade-offs, and different perspectives; look at the Bayesian/frequentist debate which has been going on for decades among statisticians.
But, FWIW, I definitely saw Optimizely as an organisation make a big investment to produce a stats engine which had the right trade-offs for how their customers were trying to test; and I think the end result was way more suitable than 'traditional' statistics were.
[1] https://help.optimizely.com/hc/en-us/articles/200040115-Traf... "Traffic Auto-allocation automatically adjusts your traffic allocation over time to maximize the number of conversions for your primary goal. [...] To learn more about how algorithms like this work, you might want to read about a popular statistics problem called the “multi-armed bandit.”"
[2] https://blog.optimizely.com/2015/03/04/bayesian-vs-frequenti... "Yet as we developed a statistical model that would more accurately match how Optimizely’s customers use their experiment results to make decisions (Stats Engine), it became clear that the best solution would need to blend elements of both Frequentist and Bayesian methods to deliver both the reliability of Frequentist statistics and the speed and agility of Bayesian ones."
I didn't realize that the auto-allocation ever shipped, but I'm glad it finally did. Hopefully there was work done to empirically show that they solved a lot of the issues around time to convert and other messy parts of the data that killed earlier efforts, but I think everyone who knew about those was gone before you joined :)
There are very subtle issues with both frequentist and bayesian stats, which makes combining them sounds insane to me.
What are you up to these days?
Being smug and condescending really backfires when you don't know what you're talking about.
> Being smug and condescending really backfires when you don't know what you're talking about.
How's that working out for you?