A/B testing isn't a new thing - I'm sure the inventor of the wheel experimented with different shapes, and the buyer of the hexagonal wheel probably didn't have the best user experience.
Multiply that by the number of people in the world and the number of products people use, and A/B testing is really up there as possibly one of the most beneficial ideas ever.
I really don't understand those who claim it should be banned - I see no way that testing two different versions of a website with people who desire to use that website can bring sufficient harm to outweigh those massive benefits.
I was in the "B" group, and felt so humiliated at how many times I tried to reset my password to get into Facebook.
* https://www.forbes.com/sites/kashmirhill/2014/06/28/facebook...
It's a great tool, and its impact all depends on how you use it.
The major issue with A/B testing in the workplace is it causes confusion and slows people down when you change things. Which makes these tests really expensive even if they are seemingly easy to preform. So, I would call it useful but flawed.
If you're building a tool to make life easier for the user, something that gives them a better experience is your optimal outcome. This seems like a scenario where A/B can produce a good outcome.
The challenge is when you throw in an ad-based revenue model, and the A/B testing is then optimized for the opposite (eyeball-hours, linear metres scrolled per session, ad spots passed, ads clicked) - engagement-based business models end up (I'd argue) A/B optimizing for the opposite of what their users want, to get them to spend longer doing a task they could have done quicker.
The funny thing is - the ad-based revenue model is not the only possible variant. Last time I’ve checked Facebook’s profits per user were $7 per quarter, that is $28 a year. At the same time I am paying LiveJournal $25 a year for the ad-free version. Just taking my money looks like a much better model in many respects:
- less overhead: a lot of people doing these studies how to force me to look at something I do not want to look at will be free to do something more useful to the society;
- streamlined relationship between me and my publisher: in this model there is no advertiser who can say “I do not like these texts, no revenue for you”.
That’s why I prefer to pay for some Substack authors, like Matt Taibbi and Glen Greenwald, than to try to fish their texts for free amid some sea of “clever” advertising (hey AI testers, I bought this thing already, what’s the point of forcing it on me again and again?).
I kinda wish that Brave model (my money distributed between sites I visited) got more traction. It looks much more healthy.
I bet by at least 4 orders of magnitude and likely 5 or 6.
Perhaps you should just agree that, "not all a/b testing is the same".
Quote - "Speaking from work experience: A/B testing doesn't look into the nuances of usability or productivity, it looks at easy-to-quantify metrics like conversion rates and money spent"
There's also a neat sleight of hand here. Your inventor of the wheel surely tested multiple variants to optimize for the utility of his invention to the user. The A/B testing that's problematic is about optimizing taking advantage of the user. That doesn't lead to better experience, but the opposite. This is what's increasingly popular, and this is what people complain about or want to see banned.
Related: attention economy is predicated on bad user experience, because it makes money from friction.
The primary goal of A/B testing is to see what's more profitable.
If that happens to result in better UI that's a side effect.
In fact, it could result in less usability (relevant to this conversation, it probably resulted in the frustrating "algorithm-based" timeline at FB/Twitter/etc).
Well, I think I'll provide a disagreeing opinion. :)
I assume this opinion probably comes from your past experiences, and I believe it is true in many cases. Since I'm not American and have never worked in an American corporate environment, I can't say what is true over there... but my experience in EU and Canada with A/A/B, A/B/C and typical A/B testing (as well as building such testing tools for others) was not like that.
For example, when building tutorials for users, profitability is far from being the primary objective. Same goes for building documentation, programming languages, open-source software, internal tooling and other such things.
Of course, I get that in the end, profitability is the primary goal of the company (with some exceptions). But I maintain that not all A/B tests have profitability as their primary goal, which makes the previous statement an incorrect generalization IMO.
A/B testing lead to the development of effective "dark patterns" in UI that trick users into doing things they don't want or don't understand, and then making it difficult to undo.
Totally the same as A/B testing button placement. Totally.
Consider yourself that inventor, would you A/B test hexagonal vs round?
As for the other aspects they sound like great targets for testing within different use-cases, but I'm not sure why that'd be an A/B test as we think of them now.
Worse for whom? I feel like a lot of the A/B testing results in more revenue, a more addictive app, and less user satisfaction, because they're not testing for anything beneficial to the user, because at least with FB, you're not the customer, their advertisers are the customer.
For example, I have used A/B testing to see find ways to help users get a task done with fewer clicks, saving them time.
I have not experienced a product become better for the user as a result of involuntary A/B testing in my entire adult life.
Producers and consumer have both an adversarial relationship and a mutually beneficial relationship, and the distinction between these two is essentially the split between voluntary A/B tests and involuntary ones. In the adversarial component, the producer is trying to figure out how to extract more money from the consumer, without improving the product. Alternatively, (and equivalently), how to make the product cheaper, but also worse, in a way that yhe customer doesnt notice (with their wallet). A proactive version of the "market for lemons".
For instance, if you A/B test your cancellation process to minimize the number of people who cancel their subscriptions, you will almost certainly do something that makes you some additional money, and is also unambiguously evil.
Any A/B testing that is mutual benefit to consumers and producers can be done with consent, by volunteers. And the miniscule amount of scientific rigor you would lose by doing so is not worth the tremendous sacrifice we have seen in quality of consumables in the past 2 decades (probably longer, but i do not have the personal experience to go longer)
You might be compelled to describe involuntary A/B testing as a strategy for maximizing evil subject to the constraint that it be legal, but it often dips its toes into seeing what is illegal but still profitable, and is capable of fundamentally undermining our legal system and even our political system.
The technology has grown more powerful. The addition of computers that can optimize essentially arbitrary objective functions has serious existential implications for humanity.
A blanket ban on the practice, incurring the total dissolution of any corporate entity found guilty of the practice of involuntary A/B testing, would be a start.
If you did, how would you know?
When craigslist added the map that shows you where all of the people are offering the thing you are interested in. That was a very good change, but thats pretty far from how craigslist operates.
When dominos stopped serving hot glue on cardboard, its pretty easy to see how that didnt come about by furtive A/B testing. They were pretty confident people would like the new pizza more than the old pizza. So they told them about it. Boy did that work for dominos.
That actually speaks more generally to my point. If you're making a change that you think people will like, you tell them about it, because even if it turns out that they dont like it more, the fact that they thought they would and you did it generates quite a lot of good will for them.
Do we live in the same universe? As far as I can tell, software keeps trending worse. Usability is terrible, options and settings keep getting moved around and hidden, software is less responsive than it used to be...
But the majority of people choose to use a modern computer, presumably because they find it overall more useful than their old MS DOS computer and software.
Sure - there are gripes, but they must be outweighed by something pretty big for 99.99% of people to choose a new computer over a 30 year old one.
Imagine thinking that seriously.
To me the latter view is the one that’s hard to take seriously.
Within a university, research with human subjects is required to pass an ethical review before it is allowed to proceed. Given the scale and impact of the research conducted by Facebook on its users, it is entirely reasonable to hold them to the same standard.
So we agree that ab testing is good for optimizing toward a numerical objective. You then seem to think that either:
A) There are simply no numerical objectives that correlate with good ux or
B) Every software company ever is optimizing toward perverse incentives by which they take more money from their users while making their products worse
It’s probably B that you believe, and this is such a myopic and paternalistic view. There are a couple cases where it’s a problem, eg cancellation flows. But this problem is orthogonal to AB testing (try cancelling your newspaper subscription in 1994). AB testing is mostly just trying to improve the rate at which people sign up or buy something, and in this case, your objection hinges on the hidden premise that people are idiots.
> Within a university, research with human subjects is required to pass an ethical review before it is allowed to proceed. Given the scale and impact of the research conducted by Facebook on its users, it is entirely reasonable to hold them to the same standard.
One man’s ponens is another’s tollens. See: https://slatestarcodex.com/2017/08/29/my-irb-nightmare/
My qualms are not so much with the method as the morals that guide it. It's agnostic but when operationalized in a faulty moral framework can definitely lead to bad results.
obviously "A/B testing" is not "hiring psychologists", because one is in the category of experimental methods the other one is about human resources
yet there are obvious connections. A/B testing is used to increase "conversion", which is profitability. which is the same fuckin' thing as addictiveness in case of a site where you pay with your eyeballs
Absolutely, go bananas A/B testing different colors for a “Sign Up” button or testing different pricing models.
But let’s not go bananas optimizing algorithms that are damaging to users mental health at a massive scale.
That is untrue.
A/B testing is not the same thing as seeing how bad your users' mental state becomes if you muck with what they read.
A/B testing is a tool. It can be used for good or evil. That was Facebook's choice how to use it.
They did it to make sure they could keep people's attention, despite the psychological damage.
Read the whistle blower report, witness the evolution from seeing content from your friends posted in less addictive chronological feed to addictive content your friends like in the internet sorted by addictive news. Hell, the site started as a PHP hack to creep on pretty women. They sold a bunch of data to foreign adversaries. For years, they let people sell ads to Nazis. They don’t give the people faced with the psychologically brutal jobs of moderation get benefits. They have been a platform for genocide and government surveillance.
I might be missing some examples of them missing some money to do the right thing, but nothing comes to mind.
In research, these types of experiments typically require consent..
(This is not an attack on you or your otherwise valid point. Just a reminder that people should be mindful of their ethical obligations to get informed consent and not cause harm to others with their experiments).