"Network Experimentation at Scale" from Facebook describes how difficult this problem is. Most A/B test frameworks don't reach this level of sophistication. It does make some sense to just ship things if you don't have time to build out something like that. (disclosure: I worked at Twitter long ago)
In my experience, they mostly just catch bugs. Stuff like "hmm, our much better looking signup flow underperforms... oh, the form is broken on Safari." That kind of thing.
This can be the case for A/B testing. Sure, you can increase ad clicks by 30% ... if you trick the user into clicking it through a carefully timed layout jump.
I think GP's argumentation may go in this direction. I'd probably not say A/B testing is the problem itself, it is a tool after all, but I could imagine it's sometimes not used very well.
Another point: Spotify's core flow changes so much (feels like almost daily) that I've lost all confidence in using it.
a double blind test can help you determine the most effective way to cause pain to a monkey, but it will never answer the question of whether you should be doing so.
Science is just observation and experimentation.
Science doesn't dictate how you do the above. Now, someone would find it impossible to reproduce your findings, but - that would just suggest bad science
For example, if your metric is "time spent interacting on the platform", then a testing of a rollout of a feature ends up with longer page load times, so users spend more time there because they're waiting for pages to load would increase that metric, and management decides it's a good idea.
That's not enough. If you don't include some sense of both 'systematic' and 'rigorous' (and yes, these terms are slippery), you aren't doing science.
Until you can objectively measure how "systematic" and "rigorous" an experiment is, your definition of science only applies to you.
We haven't been doing science for very long. The primary difference is the desire and effort to add rigor and systematic thinking. The difference in efficacy is hard to understate.
Again, this comes back to your definition. Many would disagree in that science has existed for, at least, as long as recorded human history because an actual tenable definition of science is something along the lines of "the endeavor to build knowledge by experimenting and observing the results". The rigor of the experiment is part of the quality of the science, not whether it's science itself.
No one is arguing that rigor isn't important to good science. It is important because rigor lends to reliable and valid results. What we ultimately want is results that are reproducible and can be used to predict. If you observe bad results, it's because you did a bad experiment, thus bad science, not that you didn't do science at all.
As an analogy: if I took notes during a meeting that no one can understand or use, that doesn't mean I didn't take notes. It just means that I took bad notes.
Not sensibly. Or at least; let's avoid bogging down on the semantics. We started doing something quantifiable different recently, which has had a massive impact on our world. It is quite sensible to ask "what changed?" and try to understand it. If you want to give it a different name from "science", ok, but that's mostly likely to confuse people. If you want to claim such a shift didn't happen, you've got a hard row to hoe.
Did it? The nature of Tesla and his other current businesses buffer that a bit even if it has been his approach, and it seems to have gotten him thrown out as CEO at X.com twice; among the things going on Twitter seems to be Musk trying to relitigate his failure at X.com without other investors being in a position to kick him out, but he seems to be piling up existential threats without resolving them.
I'm not sure what you mean by this. Many quality studies are A/B tests. A/B just refers to the two IV states you're testing, which you're then observing a DV - sales, engagement, errors, etc.
A/B tests can be double blinded (don't tell the error monitoring people which results are from a trial), and have high number of samples, far beyond even most pharmaceutical trials.
They can also be really crappy, changing too many variables at once, etc. But they are certainly "real science".
EDIT: an example, Drug vs placebo - is an A/B test.
Science is more “what’s true if humans didn’t exist.”
Marketing is more “what widget generates more revenue?”
I put this earler in the phrase "reflection completeness": https://sdrinf.com/reflection-completeness ie there are things which stops working when people know about it.
In particular with A/B testing, this means that the initial A/B test is intermingled from at least 3 effects: specifically it measures how the naive population's behavior changes as a function of new functionality being made available. This is heavily, heavily time-dependent; specifically there's a "novelty effect" (early data collection will not be representative to long-term usage patterns); and there's "reflection effect" (once the outcome of the test is widely known, people can change their behavior based on that). Controlling for the first is difficult, but possible; controlling for the second, beyond just "keeping everything secret", is significantly more so, as the timelines for that might be years in length.
I strongly suspect GP was pointing at this timeline factor, and specifically that market engineering, as currently, generally, widely practiced, is grounded on the immediately available signal of "does it increases sales in 2 weeks of A/B test running". Which, given novelty effects, is heavily biased towards "yes"; and these people aren't incentivized (nor have the time/energy) to measure _very_ long-term effects beyond novelty, and reflection period.
An A/B test just refers to observing how a dependent variable changes when an independent variable is in two different states, State A and State B.
Drug vs placebo - is an A/B test.
To borrow the Lindy effect; whether someone likes the jacket in color A or B is of such short lived value it’s a huge waste of the resources that went into the pipeline needed to come to the conclusion.
Here’s an A/B test; rethink logistics to increase customization of outputs or continue to create design jobs who define what’s trendy and acceptable?
In the context of what we're talking about, you can A/B test more than marketing, you're can test variables like UI/UX.
Yes clothes fall in and out of fashion, but changing the placement, color, size of the "add to cart" button isn't something that's going to be changing frequently.
Another example might be adding a "trending" tab the top navigation of a page or whether the "what's trending" vs "what you like" provide more engagement as the default page.
Youtube recently tested randomly lowering people's video resolution to see who changed it back to gauge the importance of the resolution to their customers.
I disagree that I am “caught up” on anything.
I have a preference that’s been refined over time. Not a psychological error in perception.
I wish they start gauging how frustrating such tests are, particularly for the test group. I've been cursing at YouTube many times over the past weeks because of this very issue - and now I learn it's not even a bug, but an A/B test.
For example with changing the font size of a button:
Your null hypothesis is there is no difference in the number of clicks. Your alternative hypothesis is that there is an increase in number of clicks.
Your IV is the button font size. Your DV is the number of button clicks over a set period of time.
You randomly sample 50% of the population to State A (same button size) You put the other group into State B (increased button size)
You observe the number of clicks of the button.
You analyze this data, and can determine the statistical significance between your null and alternative hypothesis.
Sociologists will frequently not have as good access to such a large participant pool under near ideal experimental conditions with such good ways to observe behavior. And the stuff you have to keep in mind when running experiments is not terribly complex. A bit of statistics, a few things you absolutely have to get right, that’s it.
Obviously there are reasons why AB tests are often not run rigorously (statistical illiteracy and pressure to get things done quick as well as to produce tangible results as often as possible – all three of which might lead you to run underpowered experiments with too few participants and to stop testing early which will lead to too many false positives). However, stopping to do experiments (and instead just releasing new stuff and observing the reaction) isn’t really an improvement that leads to better outcomes compared to that.
I mean, I'm sure the parameters to the math are proprietary. But the basic math seems simple enough.
Trying to tease out the pieces that aren't coupled to Twitter's User class is probably more effort than it's worth
I guess the comment I read implied Twitter had an amazing A/B test suite, as opposed to a tightly specialized A/B test suite.
To some extent, you grow your company and codebase around your A/B testing system, not the other way around, because it has to slot into so many places: deployment, testing, front-end, back-end, analytics, monitoring, etc. Almost no two companies share the same stacks across all of those dimensions, and an A/B testing system that doesn't hook in tightly to every one of those systems is not complete. This is why I always get a bit scared when people think they can just use one of those "drop-in" services and be done with it: yeah, great, you can now fiddle your JavaScript from some third-party website, but you're going to have a big project ahead of you passing group assignments out to all of the different systems that will need to know about them, and almost guaranteed some part of your stack will not have a library available from the service you picked so you're going to be writing your own REST wrapper, and dammit they don't document that API very well and it seems to be responding differently since the last update, and man it'd be a lot easier if I could just pipe results straight from my own service into Big query rather than running a daily user export, and damn, their dashboard doesn't let me set exclusion criteria on my own metrics, I have to send activation events now, and etc, etc. By the end you've basically built your own A/B test system from scratch, you've just paid someone else to do the "int myGroup = Random.next()" call, which is the easiest part to build.
A case could be made that A/B testing is insufficiently rigorous given specific goals, resources, limitations, context, etc. But that case isn't being made here.
Thanks though. I've narrowed my original comment to more accurately represent the scope I am referring to.
[0]:https://philosophy.stackexchange.com/questions/10178/is-math...
Where A/B studies may go wrong in my view is a few other elements:
- A/B studies have difficulty in determining differences based on multiple interacting characteristics. In fairness, so does empirical science, and the principle of "holding all else constant" is a frequent assumption of scientific processes.
- A/B studies face an inherent self-selection / exclusion bias: the participants in this round of A/B testing are those who've not been driven off the project/product from past experiments and design changes. Given that many Web 2.0 companies eventually dance with pushing people right up to the border of tolerance, it's quite possible that A/B testing has a long-term effect of pushing those participants whose tolerance has been exceeded out of the study population entirely. I don't know how large a factor this is, though loud / rage quitters are certainly a prominent (if not necessarily large) cohort. Whether or not they're also influential, or perhaps more importantly when they become influential is another question. Again, this is a fairly common problem with any social experiment, including natural social experiments, see various forms of brain-drain and social flight.
- A/B testing tends to focus on short term changes and behaviours, which may mask longer-term outcomes. This has some overlap with the above, but also with subjects' general response to change. See the classic case of this in the Hawthorne Effect (<https://en.wikipedia.org/wiki/Hawthorne_effect>).
The upshot is that A/B testing can be valid and useful, but that experimental design, particularly in the case of social and psychological experiments, where subject feed back into the study and its methodology itself, is exceptionally thorny.
I'd compare this to how evaporation rate increases with temperature, as more particles find themselves with enough energy to escape the liquid.
From my personal experience, even if I can tolerate a lot of UX abuse, each such "optimization" lowers my threshold of switching to a competitor. Software in general, and SaaS specifically, resists commoditization, but every now and then an actual alternative to a product/service I'm using shows up - and whether or not I switch (and when) is correlated with how much I resent the incumbent for their UX "improvements".
I'd add one bullet point to your list:
- Unlike regular scientific experimentation, A/B testing is a methodology primarily spread in business circles using regular hype channels. That is, the average practitioner is not qualified to execute it correctly, which is one of the reasons I see A/B testing more as tools to launder arbitrary decisions. Because consequences of doing it wrong are typically not immediately apparent or obvious, both companies and customers suffer (and a vast space for fraudsters is created).
I'm in a charitable mood, so I'm not passing judgement on people for not having PhD-level understanding of statistics - just pointing out that, to the degree much larger than in sciences (even soft ones, which suffer some of the same structural problems), there's little pressure to do such tests correctly (and there's lot of ways to make money or status by doing them without regards for correctness).
From what I hear, a common way of executing A/B test badly and getting bullshit results, is by terminating the test early when it shows the relevant metrics improving for the test group - vs. running it longer if no big improvements are observed (or the metrics start getting worse for the test group). This biases the experiment towards giving false positives. This problem was big enough that there was a debacle around Optimizely few years ago, whose UI was accused to promote this early termination of tests. The cynical take I'm still somewhat partial to is that it wasn't an accident (if not done on purpose, then possibly... a result of an A/B test!) - false positives make the (statistically naive) users feel they're getting more value from Optimizely than they actually are.
There's a reason that the technical term for "A/B testing in SaaS products" is gaslighting.
O hai werd uv yeer!