Why Most Startups Are Doing Data-Driven Decision Making Wrong
fforward.ai
fforward.ai
What was it in the initial A/B test that resulted in your 10% improvement in conversion rate to be non-statistically significant?
Would applying a simple statistical test to your decision making process have resulted in a different conclusion?
There might be a deeper, useful insight in answering those questions, but as it stands now I really don't see anything valuable in this article. If anything it leads me to believe that the author is susceptible to p-hacking (see: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3204791)
Hold on - you can't just say ChatGPT "adds precision" with nothing to back it up! Acting as if ChatGPT is capable of thinking about your results better than you seems extremely unwise.
The real way that companies are doing Data-Drived Decisions wrong is by only measuring one thing without correctly thinking of other factors. For example, a website might find that having a pop-up that says "SUBSCRIBE TO OUR MAILING LIST!!!" might increase mailing list subscriptions, but you are forgetting the effects that has on other parts of your business like the website being slower, customers being less happy, etc.
You're totally correct in saying that forgetting the effects on other parts of the business is where a lot of the analysis needs to be directed.
The toy example provided is actually a perfect example of why tools (like chat gpt) can't protect you from fucking this up. Why is the view count so different between versions a and b? Was the test set up as a 60-40 split, or had you intended it to be closer to 50-50? Does the massive increase in conversion rate pass the smell test? If anything seems suspicious, you probably should do some further investigations...
Given any model, a researcher usually doesn't just run it once and accept the output (unless they happen to agree with the output of the first run). If they get an "unexpected" result, they will either "reassess", meaning that they genuinely have learned something new, or "develop the model", re-tweaking the model structure and parameters until they get an "expected" output.
Most of the value in these data-driven exercises doesn't seem to be deciding between A and B. Any executive can do that (and ask the data scientists to justify the decisions after the fact, as others have said). It's in finding C, which you never even considered and takes you off in a new direction. For that to happen, the human mind and the organization behind it needs to be flexible. And this generation of generative AI certainly seems capable of providing or helping to reach novel insights. Or is it? How can we go off of the beaten track using a tool which is built on top of following the track as much as possible?
I like what you said down lower in your comments about looking at data, "Update your gut with an updated map of the state". That's a great description and I think it's exactly what you want from whatever data tool, AI or whatever.
It’s also not clear why you think “most” startups are not doing statistical significance tests on their a/b test results
I tend to agree that data-driven decision making is often deeply confounded and more association-based than causation-based, but I'm not really sure statistical significance is the right way to solve that in every case.
Data comes in a variety of forms - many of which are intrinsically expensive to acquire. It's ok to acknowledge that you don't have sufficient data, and won't be able to acquire sufficient data in a reasonable time/budget - and that you need to use judgement.
I really don't like it.
"Rationality is not just about knowing facts. It's about knowing which facts are the most relevant." - Epictetus
Given that he was paid stupid money, I would be inclined to do the same.
This is also how a lot of finance (especially valuation) and all of management consulting work.
Is it signups that matter? Or DAUs? Or MAUs? Or (and it actually is this one) revenue?
But how do you decide the relationship between each of these? Does an easier on-board mean that you get less "sticky" users? How will you tell?
After we hit that size we would analyze the results. If they were ambiguous we just picked the one we subjectively felt was the better choice.
We also used data a bit less formally. For example, we had something similar to Google analytics' "funnel" feature of how people arrived at some of our pages, and we noticed a usage pattern that indicated we were missing a feature. When we added the feature directly it was one of our most used features.
Constructing experiments is hard, and interpreting data is even harder. PhDs, whose job it is to do these things, who typically have something on the order of a decade of education and even more years of practice doing these things, PhDs mess these things up. It happens all the time.
That's not even getting into the fact that in many times in data driven decision-making, an experiment is constructed to motivate a particular decision. If you have a preferred outcome, it's very easy to deliberately or unconsciously put your hand on the scale. Even without deliberate sabotage, there are many methods to apply to a particular dataset, and if you apply enough tools, sure enough one will support your assertion.
More often than not the result is that data driven decisionmaking is more like an elaborate ouija-board that says more about the people creating the experiments than what they purport to verify.
A vast majority of orgs that claim to be data-driven don't have the correct model