Google Optimize now free for everyone
blog.google
blog.google
Chrome extension: https://chrome.google.com/webstore/detail/recipe-filter/ahlc...
Source: https://github.com/sean-public/RecipeFilter
Video demo & explanation: https://youtu.be/3Xq1p10f3v4
That seems specious. By that reasoning we could predict long form articles would triumph over short ones, but that doesn't bear out.
It seems just as plausible that there's a population of readers that do like those terrible rambling stories and tend to be more loyal to a site if they do, vs a large population who will just click on whatever recipes show up at the top of the search results with a reasonable sounding recipe name.
Recipe sites also rip each other off all the time, and it would be difficult to tell where the authoritative source of a recipe was. The personal story part is harder to rip off without being caught as a scraper (and may then boost your SEO prospects).
Your other point are good though.
I highly doubt this is true. Scrolling to the bottom of a page is a very low cost to pay for a recipe.
Yeah I don't either, I just wanna know how to make the damn cookie.
Had Henry Ford used ML he would've invented a faster horse.
Behavioral analytics on user interaction can help you fix a bad design but shouldn't matter that much.
That said, it certainly can be.
(not necessarily true, but it has happened before with YouTube)
We have a few thousand visitors a month and are starting to convert, but my guess is A/B testing language and buttons would be premature optimization for us. Just curious at what point that's no longer the case.
The challenge is gaining statistically significant data. I think it is easier for an early stage customer to talk to their customers versus go through the time of a split test.
Suppose you have 10k monthly sessions with a 0.5% conversion rate (50 conversions). How many more customers would you need in order to prioritize running a test? If 55 conversions in a given month means you crush important KPIs, then that's probably worth testing -- you just need a 10% lift.*
Also keep in mind that running A/B tests (1 control, 1 treatment) is suboptimal. That tests, "does this beat what I have now?" The more important question is "what is my best option?".
OTOH, if other things like messaging and product are stable, you can test a smaller traffic site by leaving it running longer.
My rough estimate is 100 conversion events in the time the test runs. So if I have 100 conversion events in 1 month, it may make sense to run a 2-3 option + 1 control test for 1 month.
(You can also test much larger things than buttons. For startups, I like to suggest trying out positioning or value statements and seeing how visitors respond!)
* however, it'll take a long time for you to reach statistical confidence for a 10% lift in rate, with only 50 conversion events across all tests.
The idea is that to get results, you need some combination of a lot of data, or a big impact from the changes your testing. If you just change the color of the signup button, it probably won't have a major impact on the conversion rate, so you'll need a lot more data to reach a conclusion. But if you test a completely new landing page, it might have a better chance of being meaningfully different (better or worse, who knows until you test?) and so you wouldn't need as many visitors to get a result.
There are a few variables to consider
- What goal would you like to AB test? Conversion rate is an end-of-funnel goal that needs a lot of traffic, you can use upper funnel goals like product views, add to bag etc to get quicker conclusions (not as accurate but often a good approximation)
- The stats engine/AB testing tool you are using. More simple tools might conclude quicker but in my experience they can be so inaccurate they are counter productive. Usually a long time to conclude = reliable results. I've never used Google Optimize so I'm not sure where it stands.
- How many people are being exposed to the AB test, for example is it all web traffic or just mobile?
- How much of an affect the AB test has on behavior. A button color/text change will normally take long to conclude than a feature that's really helping your users.
- How confident do you want to be before reaching a conclusion? I'd recommend looking for 95% confidence in uplift before concluding an AB test.
In an A/B test, the probability of selecting an "arm" (a treatment to show to the visitor) is equally distributed. Google Optimizer differs by adjusting the traffic distribution, sending more traffic to better performing options, what's usually called MAB approach.
The difference is that MABs will more quickly converge on the "winning" variation, but are more likely to get stuck in a local minimum. EG, if an "worse" option performs better right off the bat, the MAB might send most traffic to it. This would take a while for the algorithm to "recover".
The major advantages of the MAB approach are minimizing opportunity costs and the ability to capture seasonality because you can continuously run tests. Traditional split testing runs for a while, gets a result, and moves on. With a MAB, you can assign an "explore" budget that keeps tests running in the background to capture the seasonal/periodic change in conversions.
That comes with a cost: every visitor who doesn't see your best page is lost revenue.
I'd suggest thinking about the following BEFORE YOU RUN A SINGLE A/B TEST:
1) Key Metrics: Define these. They are the general, "I don't care what your experiment is about, these numbers are important." Every experiment you run should automatically track these metrics. You should also give the ability to define custom metrics, since an experiment that changes some random button color probably wants to look at how many people clicked the button, which is almost definitely NOT a key metric.
2) Logging Infrastructure: Make sure that you have a easy-to-use, reliable data pipeline set up for logging and processing events. Bad logging == bad experiment results. Also consider streaming vs batch processing for updating experiment results.
3) Population Management: How do your experiments segment users? Are variants calculated in realtime? Batched with some SLA for lag? Are they sticky?
4) Mutual Exclusion: People running experiments often want "their" users excluded from other experiments.
5) Guardrails: Do your experiments automatically shut off if there is a catastrophic decline in one or more key metrics? What safety measures do you have around determining if an experiment is safe/valid? How do you handle cleaning up data when there's a problem? What sorts of actions invalidate an experiement's existing results? Does your entire site break if your A/B Testing service is down for whatever reason?
6) Cleanup/Ownership: Experiments don't run forever (at least they shouldn't!). Cleaning up old features, populations, etc can a pain, especially when the people that wrote the stuff originally no longer work at the company. Make cleanup mandatory and as easy as possible.
There's a lot more, but I'm tired now. A/B testing is complex. There are lots of resources out there, though. Look for white papers on the subject, they're surprisingly approachable. Example from Microsoft: https://exp-platform.com/Documents/2017-08%20KDDMetricInterp...
[1] https://cloud.google.com/ml-engine/docs/tensorflow/using-hyp...
(I'm not complaining–it's interesting and I'll give it a try. But it's not brand new)
Also: the headline has "now" in it. It's not the end of the world, but I've seen correction notes on NYT etc articles for far smaller matters. It's just a general principle that true information > false information.