Numpyro: Probabilistic programming with NumPy powered by Jax
github.com
github.com
[1] https://www.goodreads.com/book/show/26619686-statistical-ret...
[2] https://www.youtube.com/watch?v=FdnMWdICdRs&list=PLDcUM9US4X...
https://github.com/pymc-devs/pymc-resources/tree/main/Rethin...
or, for first edition
https://github.com/pymc-devs/pymc-resources/tree/main/Rethin...
In other words, I think I have a practical use case for calibrated confidence scores, which I definitely don't get from my NN classifiers. They are right a certain percentage of the time, which is great, but when they are wrong sometimes they still have high confidence scores. So it's hard to make a firm decision based on the result without manually reviewing everything.
So my question is: is this an appropriate use case for PyRo? Will training my NN classifiers blindly converted to probabilistic classifiers and sampled appropriately give me actually reliable and useful confidence scores for this purpose? Is that the intended usage for this stuff?
Are the two libraries in competition? Or complimentary? I’ve been playing with PyMC for a personal project and am curious what I might gain from investigating (Num)Pyro?
This textbook/walkthrough is great: https://a.co/d/9dWXDTK
One thing I remember that I disliked about PyMC was the PyTensor API, it feels too much like Theano/TensorFlow. I much prefer using JAX for writing custom models.
You can use the Numpyro NUTS sampler in PyMC with pm.sample(nuts_sampler="numpyro") and it will significantly speed up sampling. It is less stable in my experience.
A little to my surprise, despite being a Julia fan, Turing really outperformed both the Python solutions. I think JAX should be competitive in raw speed, so it might come down to the maturity of the samplers we used.
What i mean is that I have a model and I've run my inference on my historical data. Now I have new observations streaming in and I want to update my inference in an efficient manner. Basically I'd like something like a Kalman Filter for general probabilistic models.
These models have what are called sufficient statistics, that can be computed on data, s = f(D) where s is the sufficient statistics, D is the past data. The clincher is that there is a very helpful group theoretic property:
s = f(D ∪ d) = g( s', d) where s' = f(D). D is past data, d is new data, D ∪ d is the full complete data.
This is very useful because you don't have to carry the old data D around.
This machinery is particularly useful when
(i) s is in some sense smaller than D, for example, when s is in some small finite dimension.
(ii) the functions f and g are easy to compute
(iii) the relation between s and the parameters, or equivalently, the weights ⊝ of the model is easy to compute.
Even when models do not possess this property, as long models are differentiable one can do a local approximate update using the gradient of the parameters with respect to the data.
⊝_new = ⊝_old + ∇M * d.
(∇M being the gradient of the parameters with respect to data, also called score)
With exponential family models updates can be exact, rather than approximate.
This machinery applies both to Bayesian as well as more classical statistical models.
There are nuances also, where you can drop some of the effects of old data under the presumption that they no longer represent the changed model.
On the PPL side, I think SMC has often been a secondary target, but there has been good work on getting good, efficient inference be more turn-key. I think this has often focused on being able to automatically provide better/adaptive proposal distributions. For example, "SMCP3: Sequential Monte Carlo with Probabilistic Program Proposals" by Lew et al 2023 (Stuart Russell is among the contributors).
I've seen lots of other approaches proposed for this over the years, here's a recent Stan forum thread with some links: https://discourse.mc-stan.org/t/updating-model-based-on-new-...
It seems like if raw prediction of a data distribution is what you are interested in, explicitly specified statistical models are probably less useful? At least if you have lots of data and can tolerate a model with lots of 'variance'
You can think of a probabilistic programming language as a set of building blocks for building statistical models. In the olden days, people used very simple frequentist models based on standard reference distributions like the normal, Student's t, chi2, etc. The models were simple because the computational capabilities were limited.
In modern days, thanks to widespread compute and the inference algorithms, you can "fit" a much wider class of models, so researchers now tend to build bespoke models adapted for each particular application they are interested in. Probabilistic programming language are used to build those "custom" models.