Dear Software Makers
blog.jim-nielsen.com
blog.jim-nielsen.com
----------------
It's going to be awkward if you share a youtube link with somebody and what they see is significantly different from what you saw, perhaps even to the point of them replying, "Why on Earth did you send this to me? Are you on crack?"
More importantly, this will almost inevitably lead to content creators being given the ability to not just randomly A/B test versions of a video, but produce different versions of the same video that are shown to users based on their data.
e.g. Shania Twain used to produce different versions of her albums with different instrumentation based on which section of the music store they'd be sold in. There was a Country version for the Country section and a Rock version for the Rock section. She's still bootin' around today, so she could produce different versions of music videos targeting users based on whether Google thinks they like Rock or Country more. This would be relatively harmless, although one might be surprised by the version that appears on a friend's phone.
Musical taste isn't what really divides people these days. What might content creators do if they could show different videos to people based on their political views? This might be good for their business, but it undermines objective reality. People would be shown different "facts" based on their beliefs. This is precisely the opposite of what needs to happen to reduce political polarization and bring people closer together. A common reality is necessary for society to function.
Yep, and A/B testing as experienced by uninformed, unaware end-users is a dark pattern.
It undermines the perception of (and trust in) continuity which is necessary to make effective use of a tool. The best way I can describe it to the skeptical is: imagine the dials on your car's dashboard rearrange themselves occasionally overnight, and on some commutes to work you suddenly can't work the radio or the AC while moving at ≥35mph. Of course, since the widespread use of touchscreens, that example became very literal.
So the car manufacturer has figured out the "optimal" arrangement of dials and buttons on their dashboard for their preferred levels of user engagement. Great. How many of those users now associate their car's brand with inconsistency? "I can't trust the damn buttons to be in the same place the next time I drive."
There are small A/B tests which absolutely make sense. You often see marketing sites making small tweaks to banners and copy. It's not that one has worse UX or even that one is objectively worse, just that different users have different preferences and it's often difficult to know exactly what will work best.
Similarly you can be very confident of something, but A/B testing it still reduces risk. Any significant change should probably always be rolled out to a small fraction of the user base first in case you accidentally change something for the worse.
I agree if you're talking about some BS experiment where a company uses A/B testing as an alternative to putting the hours into product design and user research.
Is it unethical to do phased rollouts (where a small percentage get the new version) as a way to do safe deploys? If the issue is that two users making requests at the same time might see different things, then this would also be unethical? Yet, these sorts of phases rollouts is the best way to release something safely. When I worked at a large CDN with 50,000 servers around the world, we ALWAYS did phased releases, to make sure we didn't take down everything all at once, and to make sure we caught any performance regressions right away.
Is your issue that the user might be getting a version that won't stick around? That seems always the case, whether you do A/B or not. You might rollback if there is an issue, and you will certainly roll forward at some point, meaning users will get a new version at some point.
Would it be an issue if the A/B test was temporal? Like all users got one version today, and a different version tomorrow?
I guess I am just confused by this statement:
> If you don't have enough confidence in your changes to make them carte blanche, then don't make them at all.
This seems contrary to so many other complaints we see all the time, that companies push changes out without taking into account what users actually want. So, do we want companies that push out changes with no user feedback because they are confident that they know what users want, or do we want companies that get feedback from users on whether new changes are helping or hurting.
> or do we want companies that get feedback from users on whether new changes are helping or hurting.
You don't get that feedback. The feedback you get is whether some telemetry KPI goes up or down. That's not the same as actual utility for the user.
I constantly experiment on users in my work. It’s all around extracting the most money you possibly can. Meanwhile we have mountains of UX interview material where people tell us exactly what’s wrong with our site, and we don’t implement any of it lol.
Profits are up though! In a big way! And our users continue to hate us more and more.
(I have this weird feeling that you're not upset with blue/green deployments...)
> Users don't want their shit changing all the time.
Eliminating A/B testing won't solve this problem. Even without A/B testing, they make updates, etc.
You might as well just say "Updating an online service without asking the user first is unethical."
Oh, and how much are you paying for that service...?
Look I get it feels weird but in practice most A/B tests are stuff like “does this copy change if ppl use this feature”.
The reasons they don’t is the same reason RCTs for new drugs don’t tell patients either. You end up with selection bias.
> If you don't have enough confidence in your changes to make them carte blanche, then don't make them at all
This is a bit hyperbolic, empirics is something that should be used more by decision makers not fit their own sake but for people who don’t understand why they are making them, especially in government (although It’s harder because finding cases where it’s appropriate is hard).
An A/B test isn’t just about what’s better, it’s about understanding all other things being equal how does one change to X affect Y.
A lot of ppl have bad takes on what makes a product better, and they would otherwise have a greater say in the product design.
Some product managers are just really stupid and are there due to nepotism so it’s an external equaliser and allowing the thoughtful ones to have more of a say.
> Users don't want their shit changing all the time.
Yep that’s why you don’t ask them.
When these products operate in competitive markets and not some winner takes all market these are often about improving users experience.
If this was something more high stakes like a medial trial I’d get, but for stuff like filling out a document or watching a piece of media. The stakes for most SASS app are really low.
But it's not that simple. We definitely need to "pave the bare spots" (desire paths); It's just that we need to start off, at what we sincerely believe to be an optimal place, knowing that it isn't, in fact, optimal.
this seems wildly naive
Truth is: “number go up” is the only valid strategy for growing because that’s what favours the platforms the most.
Unfortunate. Depressing.
If almost everything is different, it’s hard to learn for next time what exactly what led to a change in which ever dependent variable your observating.
It seems easy enough to say "keep the youtube we love" but how do you think this variant came about in the first place? I can assure you there have been numerous A/B tests that have led to the current feature set. And even if its a local rather than global maximum, at least there are measurable qualities by which it is preferable. Also - do you believe that everyone who says this is harkening back to the same historical reality? This is quickly approaching "make youtube great again" territory - when exactly was it great again? and why? This statement is easy to agree with and hard to prove.
It also seems easy to handle links to different variants; just supply a URL parameter. This is a non-issue.
If someone out there has a non-tea leaf divination style alternative to measuring things as a way to determine success, I'm all ears, but I don't believe in fortune tellers, and people who make software shouldn't either.
Just sounds rubbish both for viewers (look how terrible titles and thumbnails are after a/b tests) and uploaders alike
intro (2x 15 second clips pulled from random places in all the other sections)
section 1 (a/b/c)
section 2 (a/b/c/omit)
section 3 (short/long)
section 4 (a/b/c)
It would be like a choose your own adventure video, without the choosing, or the adventure.
> Enough A/B testing will turn any website into a porn site
> The funny thing about scoring systems is they are kind of little dictators. They tell you what you’re supposed to want and value. And that’s the weird thing. Scoring systems are little definitions of success and failure. I think one of the biggest differences is that, in games, those definitions are temporary and playful and under your control. And if you don’t like it, you can throw it away and you never have to play again. And in institutions, they’re authoritarian. […] After a period of time, [metrics] seem to drain what’s genuinely valuable from the system because they point people at something that’s very easily and mechanically checkable and measurable.
[0] https://99percentinvisible.org/episode/673-the-score/transcr...
I don't think I am the audience for YouTube any more.
YouTube is converging towards where all the social media sites are:
* shorts
* photo/text posts
* longer videos too - typically 12 minutes approx
* AI videos - some of which are fine but I want to be able to filter them out
* their algorithm/feed is very bad at letting me explore my interests, when I choose to - instead it feeds me stuff that leaves me unsatisified
But I came here years ago to watch TV made by people NOT bound to 12 minutes. I watch pretty much nothing else at night on the couch except YouTube but I am coming to realise it's no longer what I want.
That "original YouTube" seems to be gone. Nothing has replaced it.
Youtube certainly took a lot of unpopular decisions lately, but being able to stream any hi-res video ever posted in an instant and for "free" is a miracle. I will be happy with youtube for as long as content creators I care about are happy and I can get their fresh content via browser, yt-dlp, or other means.