Show HN: Open-source A/B testing framework
github.com
github.com
[0] https://github.com/Alephbet/alephbet
Georgi at Analytics Toolkit definitely knows his stuff. We're taking a Bayesian approach instead, which I know he isn't the biggest fan of, but I think it is much easier to understand. Itamar Faran, the author of our stats engine, has a great article that goes into a lot more detail if you're interested: https://towardsdatascience.com/why-you-should-switch-to-baye...
Quite true, sadly. FWIW, we keep maintaining Alephbet, even though honestly I have no clue who's actively using it besides us :) the codebase is simple enough that it doesn't require a lot of work luckily.
Re stats: I did implement a Bayesian dashboard of sorts with Alephbet, but I'm not sure it prevents the peeking problem on its own. It requires some discipline when planning the tests to decide ahead when to look at results. Disclaimer: my stats chops are virtually non-existent, but that's what I learned over the years. Georgi's platform really helps structure this process of planning and when to stop the experiment (either when successful or when it's failed).
Another small (but in my experience important) thing that sets Alephbet apart from other A/B testing platforms: ad blockers. Mixpanel, GA, Amplitude etc frequently and trivially get blocked by ad blockers on the client. For client-side A/B tests this can reduce the data quality (even though typically A/B tests are not privacy invasive). Alephbet's Lamed[0] backend allows you to create a custom AWS url that's far less likely to get blocked. The "data quality" with Alephbet is higher in my experience than the data we see, e.g. in Amplitude.
This looks like a great tool!
Does your system store the stats, or does it trust the stats to be stored in, eg. GA, and then just allow you to analyse them?
Is it appropriate to send email alerts when "significance" is reached? Without adhering to minimum sample sizes calculated in advance won't this result in a bunch of Type 1 errors?
Are the changes to the pages made client side or server side? I think clientside but I'm not sure. If so are they sync or asynchronous?
Thanks!
1. We don't store any raw user data. We pull things like mean and standard deviation from data sources, run the statistics, and store the result.
2. We use a Bayesian statistics engine which is much more immune to peeking problems and Type I errors than frequentist approaches.
3. Tests can be run either client or server side. For client side, we recommend bundling the SDK with your app (webpack, etc). We really care about performance so never want to add additional http requests or script tags of any kind if at all possible.
1. How do you get around needing session level data instead of aggregate data when working with non parametric KPIs? GA in particular is notorious for sampling data.
2. True, but you can't get away from the fact that a split test only run for a day or two isn't going to give you trustworthy results. It's things like this that abstract away the statistical reality for lay users that cause poor decisions to be made under the guise of being "data driven". I think as testers, and you as a provider of a testing system, have a duty not to lead businesses to believe that they are making statistically sound choices when they may not be.
2. We have a minimum sample size threshold before we run any statistics on the data. To your point, we don't want to say something is "significant" if it's 5 conversions vs 1. This is one area we're looking to improve with better heuristics. We can't completely take the human out of the loop, but we can help give them all the info they need to make the best decision. On that front, we do show Bayesian expected loss (risk) and credible intervals in addition to just the "chance to beat control".
Can you use the system to analyse results of tests it didn't run? ie. If I run tests using some SAAS that only supports frequentist stats could I use your system as a bayesian analysis backend?
Otherwise, this looks great.
They seem to allow filtering drilling down by various categories, which would make statistical significance even more of a concern.
I would also suggest supporting an alternative to MongoDB. Postgres using jsonb is a great option.
I try to always use and support open source components, as open source provides much less business risk. Since MongoDB isn't itself open source, I would be hesitant to adopt it or a product that depends on it. Mongo also has a bad reputation...
I would definitely evaluate and likely use your product if it did not depend on MongoDB.
What's the bad reputation of MongoDB that you're concerned about?
Also, you seem to have a really strong bias against it - can you explain?
Really? I have used it pretty extensively and like it... I don't do a lot of complex manipulations though, it might be a pain for some use cases.
> Why would MongoDB Community not be an ok choice?
MongoDB community is SSPL licensed, which is not Open Source. While I don't intend to offer a MongoDB hosting service, I want the option to fork the code and create (or pay some one else to fork the code and create) a hosting service for me to use. This is important because MongoDB Inc's business may not always align well with my business and my needs. (or they may just decide that they don't want to do business with me, maybe they go out of business or their business focus shifts or political pressures come to bear.) The option to create a viable community fork is critical to ensuring that the software remains viably usable. The business risk of relying on proprietary software is great. The more reliant you are on it, the bigger the risk.
> What's the bad reputation of MongoDB that you're concerned about?
Mongo has a long history with Jepsen test failures. See http://jepsen.io/analyses/mongodb-4.2.6 and the linked articles from that page. In addition, I have heard many confirmations of issues from folks who have used it in production.
> Also, you seem to have a really strong bias against it - can you explain?
I think I have explained my position above. I don't have any interest in Mongo or any of its competitors. I don't personally know anyone involved with it or any of its competitors (Though I have naturally had professional contact with some.) My strong preference, as previously stated, is for Open Source software. This preference applies broadly to all software, but especially to infrastructure software, and is by no means specific to MongoDB.
You should offload that onto DBT or some other data modeling tools.
I'd love to see an open source standard way to define metrics, but haven't found anything yet.
Any plan to support Matomo (formerly Piwik) analytics as a data source?
Growth Book is for hypothesis testing. It let's you define and run a specific controlled experiment and then you can analyze the results and make a decision.