A Bayesian view of Amazon resellers (2011)
johndcook.com
johndcook.com
p(f|good = 85k*0.94, bad = 85k*0.06) ∝ f^79900*(1 - f)^5100 * f^<p_good = prior number of good review> * (1 - f)^<p_bad = prior number of bad reviews>
∝ f^(79900 + p_good)*(1 - f)^(5100 + p_bad)
Recall that the prior parameters mean that the confidence of your prior knowledge of f is equivalent to having observed p_good positive reviews and p_bad negative reviews. So unless the prior parameters are unreasonably strong (>>1000), any choice of p_good and p_bad will have negligible effect on the posterior.The main reason Bayesian statistics is not a magic bullet is because it's up to you to interpret the posterior distribution. What really does it mean that the fraction of positive reviews from seller A is greater than the fraction for seller B with probability 0.713? What if it were 0.64? 0.93? That's for you to decide.
1. Weakly informative priors can be good to regularize and stabilize inference in scenarios with a low amount of data.
2. In case of the review scenario presented by the OP, a hierarchical model could be even better as it would achieve regularization (and shrinking) by borrowing information across different sellers.
In other words, one would learn the overall distribution of seller reliability at the same time as individual ones. This has two advantages: a) Just one (hyper)prior for the overall distribution, but no priors needed at seller level and b) seller predictions are pulled towards the mean in a principled way.
Most scientific publications that find unusually large effects coming from some association turn out to be false. If they used shrinking (2), they would avoid excessively optimistic predictions. See: https://en.wikipedia.org/wiki/Stein%27s_example
Think Bayes by Allen Downey:https://allendowney.github.io/ThinkBayes2/
Bayesian Methods for Hackers by Cam Davidson-Pilon: https://dataorigami.net/Probabilistic-Programming-and-Bayesi...
They're both an excellent warm up for Statistical Rethinking.
Everything else about Bayes’ rule (weighting likelihood by a prior distribution and re-normalizing to obtain a distribution over model parameters) applies just the same.
[0] https://www.statsmodels.org/dev/examples/notebooks/generated...
1) A moving window: you only calculate your updates from the values in the window (say the last n reviews). The downside of this method is that older values drop off precipitously.
2) A forgetting factor: there are many possibilities but one simple one is the EWMA (exponentially weighted moving average). This is pretty standard, and takes the form
Y_updated = alpha x Y_current + (1 - alpha) x Y_previous
with alpha in [0,1]. Applying this recursively, it "forgets" older values by weight alpha. This is also known as exponential smoothing in time series. The advantage of this method is that older values are simply weighed less and drop out more gradually.1. There is a vast literature in "Bayesian Change Point Detection." So you might find a point where something changed and a string of negative reviews began.
https://dataorigami.net/Probabilistic-Programming-and-Bayesi...
2. Similarly, there a bajillion ways to weight recent data. One way is to increase the gain on a kalman filter. That will make recent observations more important. There are bayesian implementations of the Kalman filter.
http://stefanosnikolaidis.net/course-files/CS545/Lecture6.pd...
1. It can provide a probability distribution. So after running bayes, you might have multiple places where it's likely a change has taken place.
2. It generalizes to high dimensional space. It's much harder to tilt the graph in 5-d.
John’s derivation is not based on the Wilson score but a Bayesian update on a beta distribution. They’re actually different algorithms. He starts a beta(1,1) and keeps updating. The advantage of John’s method is you get the variance as well but now to sort you have to technically calculate differences between normal distributions which is more involved (or you can ignore the variance and sort by the mean)
Both work for normalizing sort order so that small samples don’t get biased. But as the sample sizes get larger they both converge to the expectation by the law of large numbers.
I personally use the Wilson score method in my work and it’s easy to calculate and good enough for all practical purposes.
IMO most people think of statistics in a bayesian way, even if they don't know the notation or science behind. It is so intuitive. It is the frequentist approach, taught first in school, that feels alien and forced.