During this time there were a total of 67 new
subscriptions. Of these 58% (39) came from the
new design and 42% (28) came from the old design.
Looks like the new one is a clear winner.
Is it? This seems a small population to settle on a clear winner.Using R's prop test, I get a p value of 0.22. (Type "prop.test(39,67)" to calculate it).
I think this means that in a world where it makes no difference which design is used, you would get a result as significant as this 22% of the time.
An alternative is the Adjusted Wald method. You can try it online here:
Which gives some confidence intervals which also range from "could be better" to "could be worse". Even when you reduce the confidence level from the typical 95% to 90%.
a quick check with an A/B testing calculator
even says that this result has significance
(~90% likely)
Which calculator was that?