How to better calculate churn rates
catchjs.com
catchjs.com
Subscriptions still active at month x are represented as `l(x)` and subscriptions that "die" (cancel/expire) in a given month are represented as `d(x)`.
This gives you a "life expectancy" and a "mortality rate" (so, churn) for any given number of months that a customer has been subscribed. So I can project how long someone will stay subscribed when they're brand new (at month 0) and how long they likely still have when they're at month 8 (longer than at month 0, funnily enough).
With those subscriptions where the month-specific churn will largely decrease the longer someone's subscribed (after passing the initial high-churn first months), this allows measuring/projecting churn on a much more granular level.
If I'm not mistaken, you're aligning all time periods to the same imaginary start date. This gives you the grand perspective but ignores very real changes in your application, team, marketplace, advertising, and product-fit over time.
Just measuring the grand perspective, as you say, can be a very poor indicator for what is working well, and what isn't.
It's shocking how many data scientist I've known how no idea how to correctly model churn, typically trying to build some predictive model, which is just a more sophisticated extension of first formula in the article.
It's insane how many business problems are really survival analysis problems where you have some hazard function and censored data. In my experience having data scientists that can correctly identify this is very rare (and a great sign of how broken that entire field is right now).
Your setting could be non-contractual. You might not be able to have reliable data to frame the problem as survival analysis with censored data. You might have different groups of customers with different behaviors.
Assuming a constant failure rate is the same as assuming the underlying process is Poisson, which is not an entirely stupid working hypothesis in many cases.
The moral of the story is "by the love of god look at your data before making dumb models".
Again, modeling done right is super hard.
This is because most 'data scientists' have never actually studied much statistics.
And for the rest, it's easy to forgot all the cool tools we learnt about at university when everyone in the corporate world just says 'give me the average revenue per customer', 'give me the average tenure' etc, etc.
People employed as 'data scientists' often come from comp sci/physics/eng/econ etc. They might be great at fitting models that also have great out of sample prediction performance, but they may have never heard of maximum likelihood, or know when to use a Wilcoxon rank sum test.
There's no cohesive group of Data Scientists, and they come from all walks of life. Economics, Business Administration, Life Sciences, Computer Science, Engineering, etc.
There absolutely are professional Data Scientists out there, that haven't touched more than Stats and Probability 101. (But FWIW, a lot of DS - even though they may lack the formal education, pick up these things underway)
What makes it worse is that I've also done survival analysis, but never quite made the jump that churn is just a survival problem, though it's exceedingly obvious once you see it.
Spark even has functions for survival analysis (Kaplan-Meier I think), which means that one can actually run this kind of analysis in a consumer-tech space.
I remember spending weeks trying to get the right sampling/compute so that I could actually use survival analysis, so I would argue that it's not that people weren't aware of this, rather that (some) didn't have tools to allow them to use the methods effectively.
Just like NPS there can be a lot of intellectual dishonesty.
> There’s all kinds of churn — dollar churn, customer churn, net dollar churn — and there are varying definitions for how churn is measured.
I also find it a bit of a disservice they don't discuss other churn types while dismissing the notion completely in the introduction. Statistics is a really "it depends" kind of subject and if you don't put forth good effort to explain the assumptions it can be really hard to follow and even harder to correctly apply.
Next, I find it a bit off-putting they mostly modeled their way out of the problem instead of addressing the bigger meta question "how do I find out if customer churn is a meaningful metric for MY business?" This is a big assumption made by the article (customer churn modeled correctly is useful) but not supported by it.
Finally, I find this in the conclusion to be quite simply HI-LARIOUS for how brazen it is; while, IMHO, kind of missing their own point:
> The numbers we've just computed are perhaps the most central of all for a subscription business, yet it seems like most people have never been thought how to properly compute them.
Calculating the customer churn is valuable for LTV computation but it doesn't seem as useful as a KPI for a growing saas business as net revenue churn rate.
The minutiae of risk modelling for something as inherently volatile seem beside the point. Maybe if you're a stock analyst?
The Lomax distribution (sometimes described as the gamma-exponential) would only be appropriate if we're modeling a continuous time process.
For a contractual product with discrete time intervals (aka monthly contract), it would be more accurate to use a geometric rather than exponential model such as the Beta-Geometric.
https://faculty.wharton.upenn.edu/wp-content/uploads/2012/04...
A friend and I made an implementation in Python a while back:
The article is going towards the right direction but as other commenters mentioned, churn is an example of a survival analysis problem.
A typical model used for modeling churn is pareto/nbd. Wikipedia is a great starting point https://en.wikipedia.org/wiki/Buy_Till_you_Die
It seems this post is trying to make a model to predict churn or guess what future churn might be. But the main case for understanding churn to me is historical.
Last month we had 10 customers paying us $10 each. Two of them cancelled. I have gross churn of $20. That's not a model, just a historical fact. What's "wrong" with that calculation that these models solve?
I will definitely get different churn numbers if I look at groups of customers in cohorts based on when they first signed up, or enterprise-vs-SMB, or whatever. So I have reports to create all of those cohorts and show me the churn for each. I look at that data - with all the historical context - and make a decision about what to do next in the business. "Selling to SMB's was a good idea, but wow they churn a lot more. Let's focus our marketing on enterprises this month."
To me the churn calculation at the top of the article is plenty useful. I'm not sure what value I'd get from these advanced models. If they just exist to predict future churn or LTV, that doesn't seem particularly useful.
How do you compare churn rate of groups that aren't cohorted by age?
In this setting, things like age, gender, or other features of interest are the covariates/predictors. You can then also use indicator features like UK or US for geography. If the geography based features are significant, then the difference in churn between the regions is perhaps not simply due to age etc.