Don't build growth teams
conversionxl.com
conversionxl.com
As a crude example if you have a community site like Reddit you could harm the community by over-optimizing your conversion funnel for Reddit Gold, while still substantially growing the Reddit Gold business. For a time.
He touches on this by admitting that the super-optimized KISSMetrics homepage was not the right strategy but I would have loved read more.
Essentially, it's become really easy to get leads for cheap and incredibly expensive to get them to convert. The question becomes, which master do you serve?
I think marketing leaders need to delicately balance the KPIs with the vision of the company or product. In High Output Management, Andy Grove talked about negative indicators for KPIs as well (forgive me, don't have book within reach). I don't think any time gets spent on that in my network. Granted, it is hard to do with small teams and you might say it's the wrong thing to focus on for small teams. Eventually, you reach the point where the numbers look too good to be true and they usually are--meaning you won't convert these people because they're not ready to buy what you're selling. They're just ready to get the free info you're offering them in two clicks of their time.
I appreciate the candidness about failure even if it wasn't detailed. Can't wait for more similar stories.
That is: if you want a car to go faster you may improve the engine, or focus to make the car lighter. If you hyper-optimize the engine the smallest variation in the fuel will affect performance. If you hyper-optimize the weight you better watch out on your own diet :) otherwise the car won't be as fast as intended.
All in all it often goes back to following the Pareto principle. If you need to hyper-optimize you landing pages you probably need to focus elsewhere. You need good landing pages, not perfect ones.
so instead don't necessary want perfectly optimized, we want good enough, because inside that good enough is where adaptability to changes comes from.
At 95% certainty, you have 19 people saying “yes” and 1 person saying “no.”
At 99% certainty, you have 99 people saying “yes” and 1 person saying “no.”
It feels like a difference of four people when, in reality, it’s a difference of 80. That’s a much bigger difference than we expect."I'm a data scientist, I feel like this is a really wildly confusing way to put this.
95 & 99 are very 'close together'
But if you look at 5 & 1 - the first number is 5 times the second. Huge relative difference
Now, imagine a grid, where the unit on the x-axis is "No", vs "Yes" on the y-axis.
Place both studies in terms of their respective No/Yes relationship.
Study A: 1:19, 2:38, 3:47, ...
Study B: 1:99, 2:198, 3:297, ...
Can you picture the difference in their slopes?
- 95% is 1/20
- 99% is 1/100
This makes it intuitive that 95% has 5 times more people saying "no" than 99%. The wording made sense, but IMO using percentage adds more complexity than necessary.
A brand had two models: one that blocked 98% of the wind and another that blocked 99%. To a casual observer, it wouldn’t be immediately obvious why the second one was significantly more expensive. After all, it’s just a 1% difference!
(Of course, the answer is obvious: the second one was twice as effective at blocking the wind.)
To be more explicit, if there are 100 units of wind and one blocks 98 units, the other blocks 99 units. So one lets twice as much air in than the other, but I wouldn't consider that saying that one was 'twice as effective'. The relative improvement is (99-98)/98 ~ 1%.
If windbreaker A let's in 2 units of wind out of 100 and windbreaker B lets in 1 unit of wind out of 100, that means windbreaker B is twice as effective at blocking wind, because it lets in half as much wind as windbreaker A.
For others interested, basically at 95% confidence level, there's a 5% chance due to random events. At 99%, the randomness is limited to 1%, a 5X improvement, i.e. a huge 500% improvement not a tiny 4% improvement.
At 95% certainty, you have 95 people saying "yes" and 5 people saying "no." At 99% certainty, you have 99 people saying "yes" and 1 person saying "no."
While that is 5x improvement in uncertainity, it's also only 4 more people, a 4.21% improvement in overall "yes" count.
I can't see how comparing an N in 20 number to an N in 100 number makes sense. IANAS.
Depending on what proportion of "Version B" ideas are truly effective, your false positive rate (what the author is concerned about) will be very far away from 5% or 1%.
At the 95 threshold (p-value of 0.05) you could have that >50% of what you identified as an "effective" version are not really effective (false positive). This is a good article to build the intuition around it: https://www.statisticsdonewrong.com/p-value.html
In the end it will also depend on how costly it is to have a false negative. The author hints that in this case is very costly.
This one is important for almost all optimization efforts. Once the low hanging fruit is gone further efforts often don’t produce much but can cause harm instead. Stack ranking is a good example. It may make sense to identify and lay off the bottom 10% for one or two years. The you should stop. But if you keep going you get all kinds of weird dynamics that harm morale instead of improving performance.
Shooting from the hip here possible problems with the "overly optimized" homepage:
1. Long term brand damage? If people go to the site and see something not as "polished" or "professional" looking, the bounces won't leave with just a neutral impression but a straight up negative impression. Sometimes even without a conversion a landing page can sell a visitor on a company and lead to recommendations or return visits down the line.
2. Lower value conversions? Is it possible that people who would immediately sign up without a long sales page end up being less valuable? This is more easily tracked so I'm sure OP would've accounted for this but it's still hard to tell without a ton of data.
A/B testing micro-optimizes an outcome you're searching for. It doesn't tell you the long term trajectory of customers, because that takes way too long to get feedback. An example: "dark patterns". Facebook is deeply damaged, in part by making choices from well-run A/B trials.
Worse, A/B testing searches out local maxima. Maybe you get locked into an approach that precludes you from making the changes that would really drive the visit ultimately.
Performance indicators and properly using quantitative information are really important-- and many executives aren't versed at this. But, conversely, you can't use A/B testing to decide who you are as a company and define your relationship with the customer.
I don’t think any of this stuff is new or undiscovered, really.
1. It specifically mentions another company -> this often feels like a jerk move or from a branding standpoint something that diminishes your product. There's another front page HN article right now that leads with this:
https://alexdanco.com/2019/09/07/positional-scarcity/
2. It mentions nothing from the actual product/features which makes it hard to learn from the tests.
3. Some honest diversity questions with choosing that stock art guy as the face of the company. A tech company I worked at found out that we could get leads about 15% cheaper off of Facebook ads if we targeted only men and excluded women from the audience. We made the deliberate choice to use the less optimized combined male + female audience because it felt so ethically wrong to do otherwise.
“I have a rule-of-thumb for picking A/B test winners: Whichever version doesn’t make any sense or seems like it would never work, bet on that. More often than I like to admit, the dumber or weirder version wins.
This is actually how I tell if a company is truly optimizing their funnel. From a branding or UX perspective, it should feel a little “off.” If it’s polished and everything makes sense, they haven’t pushed that hard on optimization.”
...but if you take this too far you end up with something like Amazon.com, where everything is “a little off”?
I always use Amazon and eBay as go-to examples of Optimization that works as a counter-force to those that want massive redesigns every 6 months because some other new app or competitor has some sexy slick UI.
In such an environment, you're bound to end up with all kinds of weird "winners" but there's no guarantee that (1) they're actually better than the alternatives you tested and (2) even if they are, that the advantage is stable and not just a temporary novelty effect.
With a lead gen funnel, where everyone is going through for the first time, this is much less of an issue than a site with long term users. In the latter case you want to measure learning effects, while in the former no individual is in a position to learn. Novelty could still wear off in lead gen, though, as companies copy each other and your weird new pattern starts to show up elsewhere and become familiar.
https://www.nngroup.com/articles/amazon-no-e-commerce-role-m...
The real thing I was thinking was that productizing the activities of a growth team would be ideal. If you could build a software product that efficiently managed the process of the growth team you could make a good value proposition. e.g. one person ($150k/year) using one tool ($100k/year) is significantly cheaper than the $650k/year he quotes.
Sad but true. And it’s not even about understanding the more in depth math. It’s about having an intuition about the 80/20 rule or Venn diagrams. “If I work on this bug I can make 80% of our users happier but if I work on this bug, which affects no one but our CMO who always complains loudly...” ... is a type of discussion I’ve found happens too rarely
Won't mention the brand but I have in 2019 had in the context of a new website) comments of don't want to much text we will let images lead - which is just not going to work.
This is supposedly the digital native generation who are 20 years younger than me
edited to correct naïve to native
not notpicking; those words have opposite meanings here
>At 99% certainty, you have 99 people saying “yes” and 1 person saying “no.”
The article was worth the read to me just for this part. I thought I had a pretty decent intuitive understanding of probabilities but this put things in new perspective for me.
The rest of the article is marketing gobbledygook to me, though.
At 95% certainty, you have 95 people saying "yes" and 5 people saying "no"?
So it's easier to make an apples to apples comparison? What point is the author trying to make with changing the scale?
> It feels like a difference of four people when, in reality, it’s a difference of 80. That’s a much bigger difference than we expect.
I had to stop reading here.
Your mind sees the 5% going to 1% (or the 95% going to 99%) and thinks it's a small difference. When in actuality it's a big change.
So, what do you want to know, "yes for every 100" or "yes for every no" ? It matters.
What youre struggling with is the counterintuitive nature of applied statistics vs pure math, and this is the point TFA was trying to make.
> You aren’t getting 80 more signups as suggested
TFA isnt saying you "get" more, but just illustrating how different 95 and 99% actually are. Its restating the potato paradox, linked elsewhere in the thread
I wasnt aware "TFA" could be interpreted with a negative connotation, although that seems obvious in retrospect. Just trying to participate in the HN community, and also because its less typing : )
As a matter of writing style, repeating acronyms are invisible to the reader, whereas repeating words are annoying and remove value. This may be an opinion I gained from my military service, which was acronym-heavy.
I appreciate the question and its perspective.
But "TFA", although it derives from "RTFA", never seem to had the same negative connotations. It's just that sometimes you want to refer to the original article in question, but "original article" is long to type and/or ambigious. (Do you mean the news article from the NYT, or the scientific paper the NYT article is reporting on?) And "TA" is too short for people to clearly know what you're talking about (And did you mean "teaching assistant"?) "TFA" is short and unambigious: it always means the article linked to from the main page.
Long story short: Although etymology would suggest that "TFA mentions this" is as aggressive as "Maybe you should RTFA", in actual developed usage, they're very much not the same.
That was my understanding and my intended usage.
An 80% decrease in missed signups only causes about a ~5% increase in revenue. That’s an important point if the team that produced that 5% revenue increase costs a large amount of money to run. At close to a million a year for the team, that’s only going to be worth it for some.
And that was the whole point of the article: this optimization usually isn’t worth the cost.
That's exactly the point. How many people said yes for every one who said no:
50% -> 1 yes for every no
95% -> 19 yeses for every no
99% -> 99 yeses for every no
99.9% -> 999 yeses for every no
This is why you should never stop an A/B test once you've "hit your statistical significance". Always choose the number of tests you'd need to prove the significance before you start, and let it run even if it's "obviously winning" (or losing).
I put 100 lb. of potatoes in the sun. The potatoes were 99% water by weight. After drying a while, the potatoes were 98% water. How much did they weigh at that point?
They offer a series of actions that let you lose or gain x% n many times in a row, with an equal chance of each. The fact that +x% doesn't cancel out a -x% compounds on the difference in lay person expectations to produce horrific odds for microtransactions.
Certainty has nothing to do with how many people said yes or no, it’s about how many people are in the experiment.
If you sample size of 2, and both like the new UI, then you can’t say you have 100% certainty.
I.e. 3/4 people isn’t the same certainty as 75/100
I think the point he’s trying to make is it’s easier to get the statement “95% of people prefer X” from a study then to get “99% of people”. The first requires >20 people, the later >100.
But he’s tying together the concepts of sample size, confidence, and the outcome of the experiment in a way hat doesn’t make sense
>At 5% uncertainty
>At 1% uncertainty
Let me quote from "0 And 1 Are Not Probabilities"[1].
> In probabilities, 0.9999 and 0.99999 seem to be only 0.00009 apart, so that 0.502 is much further away from 0.503 than 0.9999 is from 0.99999. To get to probability 1 from probability 0.99999, it seems like you should need to travel a distance of merely 0.00001.
> But when you transform to odds ratios, 0.502 and 0.503 go to 1.008 and 1.012, and 0.9999 and 0.99999 go to 9,999 and 99,999. And when you transform to log odds, 0.502 and 0.503 go to 0.03 decibels and 0.05 decibels, but 0.9999 and 0.99999 go to 40 decibels and 50 decibels.
> When you work in log odds, the distance between any two degrees of uncertainty equals the amount of evidence you would need to go from one to the other. That is, the log odds gives us a natural measure of spacing among degrees of confidence.
[1]: https://www.lesswrong.com/posts/QGkYCwyC7wTDyt3yT/0-and-1-ar...
Edit: formatting
Your example shows that (non-linear) transformations aren't linear, and that our gut feels about likelihood aren't either, which is true enough.
However, you wouldn't say "101 ˚C isn't a temperature" because it takes much less energy to heat a wet thing from 97 to 99˚ than it does to bring it from 99 to 101˚ because of the phase change.
Of course, they are valid values for the set of probabilities, just like 0K is a valid temperature, and c is a valid speed for massive things. You just will never see any of those.
Sure you can: if you are using inferential statistics and trying to determine the population incidence of some trait and every sample has a sample incidence of exactly 1 then the population estimate will be exactly 1. And that works the same way for 0, or any other value, too.
> c is a valid speed for massive things.
No, it's not. An object with any rest mass would require infinite energy to reach c. It's an excluded upper bound, not a valid speed.