Pyro: A universal, probablistic programming language
pyro.ai
pyro.ai
Does anyone who works in this area have a sense of why PPLs haven't "taken off" really? Like, of the last several years of ML surprising successes, I can't really think of any major ones that come from this line of work. To the extent that Bayesian perspectives contribute to deep learning, I more often see e.g. some particular take on ensembling around the same models trained to find a point estimate via SGD, rather than models built up from random variables about which we update beliefs including representation of uncertainty.
- Probability math is confusing and difficult, and a base understanding is required to use PPLs in a way that is not true of other ML/DL. Most CS PhDs will not be required to take enough of it to find PPLs intuitive, so to be familiar they will have had to opt into those classes. This is to say nothing of BS/MS practitioners, so the user base is naturally limited to the subset of people who studied Math/Stats is a rigorous way AND opted into the right classes or taught themselves later.
- Probabilistic models are often unique to the application. They require lots of bespoke code, modeling, and understanding. Contrast this with DL, where you throw your data in a blender and receive outputs.
- Uncertainty quantification often is not the most important outcome for sexy ML use cases. That is more frequently things like "accuracy," "residual error," or "wow that picture looks really good".
- PPL package tooling and documentation are often very confusing and don't work similarly to one another. This isn't necessarily the developer's fault, this stuff is hard, and the people with the domain knowledge needed to actually understand this stuff often have spent fewer hours in the open-source trenches.
As a case study, I did most of my grad work on solving Bayesian inverse problems using probabilistic programming for applications in engineering, which is pretty cross-disciplinary. I now work mostly in ML, but I didn't really even touch anything in the ML domain until after I finished school. I could have, the courses were available, but they just weren't relevant to me at the time.
Edit: I wouldn't be surprised if there was a considerable userbase in industries like finance, but in my experience those folks don't share much.
Which largely counts as strong Linear Algebra and Probability Theory background.
CS only comes into the picture at runtime. ML theory is divorced from computability until then.
You don’t need a professional license to do math. Lots of computer scientists to harder and more interesting mathematics than their peers in the math dept. In that respect at least, the main substantive difference between the fields is about $40k/yr.
You're just stating things without justifying them.
What else would you consider a subfield of CS? Finance? Accounting? Logistics? UI design?
What is or isn't a subfield of a given science has nothing to do with the professional qualifications of those who practice it or how the tools may be implemented. We don't call pharmaceuticals "a subfield of robotics" because of how the factories are built.
The same can unfortunately be said of many "statisticians", who use statistics as a big recipe book without understanding the first thing about the mathematical underpinnings of the topic.
Don't believe me?
Go ask the first statistician you run into to give you a half decent explanation of how the Chi-squared distribution and the Chi-squared test works, see what happens.
1. When your main objective is not prediction but understanding the effect of some underlying / unobserved random variable.
2. When you don't have tons data + you have very clear ideas of the data generation process.
(1) is mainly relevant for science rather than private companies, e.g. if you're an epidemiologist, you're generally speaking interested in determining the effect of certain underlying factors, e.g. effect of mobility patterns, rather than just predicting the number of infected people tomorrow since the hidden variables are often someting you can directly control, e.g. impose travel restrictions.
(2) can occur either in academic settings or in private sector in applications such as revenue optimization. In these scenarios, it's also very useful to have a notion of the "risk" you're taking by optimizing according to this model. Such a notion of risk is in the Bayesian framework completely straight-forward, while less so in the frequentist scenarios.
I've been involved in the above scenarios and have seen clear advantages of using Bayesian inference, both in academia and private sector.
With that being said, I don't think ever Bayesian inference, and thus even less so PPLs, are going to "take off" in a similar fashion to many other machine learning techniques. The reason for this are fairly simple:
1. It's difficult. Applying these techniques efficiently and correctly is way more difficult than standard frequentist methods (even interpeting the results is often non-trivial).
2. The applicability of Bayesian inference (and thus PPLs) is just so much more limited due to the computational complexity + reduction in utility of the methods as data increases (which, for private companies, is more and more the case).
PPLs mainly try to address (1), and we do have examples of very successful examples of this, e.g. PyMC3 (they also have a bunch of nice examples of applying Bayesian inference in private sector context), and Stan (maybe more heavily used in academia).
Do you have any good resources/examples for applying these methods effectively? I've read Statistical Rethinking which is a good introduction to these methods at a high level but I find when I dig into an actual problem I have a lot of gaps and wish there were more real world code examples I could learn from.
Not sure if there is a more recent book that's updated to use modern Stan examples, but the Stan user guide itself has developed into a very useful resource on its own. It contains a large number of example models and builds up concepts incrementally. The writing style is also generally easy to follow.
I will check out the stan guide though, thanks!
But those cases are still things were you might have just a dozen variables (though each might be a long vector). It's more the realm of statistical inference than it is general programming or ML.
It hasn't "taken off" in ML because ML problems generally have more specific solutions based on the problem. If you have something simple and tabular, other solutions are generally better. If you have something recsys shaped, other solutions are generally better. If you have something vision/language shaped, other solutions are generally better.
It hasn't "taken off" in general programming because PPLs generally have trouble with control flow. Cutting off an entire arm of a program is trivial in a traditional language, but in PPLs you'll have to evaluate both. If the arm is a recursion step and hitting the base case is probabilistic, you might even have to evaluate arbitrarily deep (or you approximate that in a way that significantly limits the breadth of techniques available for running a program).
AFAICT, a truism in PPL is that there are always programs that your language will run poorly on but a bespoke engine will do better, by an extreme margin. There just aren't general languages that perform as reliably as in deterministic languages.
It's also just really really hard. It's roughly impossible to make things that are easy in normal languages easy to work with in PPLs. Consider these examples:
`def f(x, y): return x + y + noise` where you condition on `f(3, y) == 5`. It's easy.
`def f(password, salt): return hash(password + salt)` where you condition on `f(password, 8123746) == 1293487`. It's basically not going to happen even though forward evaluation of f is straightforward in any traditional language.
Hell, even just supporting `def f(x, y): return x+y` is hard to generalize. Surprisingly it's harder to generalize than the `x+y+noise` case.
I also don’t understand your f example with (x, y, noise) if you fix x and the return value, you still have two unknowns with 1 equation. How is that easy to solve?
Unless you’re considering using parametric inverses to represent the solution — but you didn’t mention this so I assume you didn’t mean this.
def f(x, y):
return pyro.sample(
"z",
dist.Normal(x+y, 1),
obs=5
)
model = pyro.condition(f, x=3)NUTS-based approaches like Stan (and numpyro) have more usage, and I think Prophet is a good example of a generalizable (if limited) tool built on top of PPLs.
Pyro is a very impressive system, as is numpyro, which I think is the successor since Uber AI disbanded (it's much faster).
However, in areas where measuring uncertainty is important, they have taken off. Stan has become mainstream in Bayesian statistics. Pyro and PyMC are also quite used in industry (I have had recruiters contacting me for this skill). Infer.NET has its own niche on discrete and online inference. Infer.NET models ship with several Microsoft products.
Other interesting PPLs include Turing.jl, Gen.jl, and the venerable BUGS.
I'm familiar with the tooling and played with it quite a bit... But never really figured out a practical application.
* Landing the Apollo on the Moon or tracking systems used by e.g. Sidewinder employ Kalman filters. See Example 24.4 (p. 510) [1].
* Predicting ride demand on heavy-tailed time series. Uber does this all the time [2].
* Estimating the effect of some policy on data with hierarchical structures (State > County > Individual observations) [3].
[1] http://web4.cs.ucl.ac.uk/staff/D.Barber/textbook/200620.pdf
[2] https://pyro.ai/examples/forecasting_i.html
[3] https://mc-stan.org/users/documentation/case-studies/radon_c...
Much more user friendly, "good enough", and actually scales to problems of commercial interest.
I think one reason why Bayesian models have not taken off is that representing prediction uncertainty comes at the expense of accuracy, for a given model size. People prefer to devote model capacity to reducing the bias rather than modeling uncertainty.
Bayesian models make more sense in the small-data regime, where uncertainty looms large.
It’s an active area of programming language research — it feels similar to where AD was at for awhile.
I work on this stuff for my research — so I do believe that there is a really good set of abstractions. my lab has had good success at solving problems with these abstractions (which you might not think are amenable or scale well with Bayesian techniques, like pose or trajectory estimation and SLAM, with renderers in a loop).
Other PPLs I’ve studied also have a mix of these abstractions, but make other key design distinctions in interface / type design that seem to cause issues when it comes to building modular inference layers (or exposing performance optimization, or extension).
I also often have the opinion that the design choices taken by other PPLs feel overspecialized (optimized too early, for specific inference patterns). I’m not blaming the creators! If you setup to design abstractions, you often start with existing problems.
On the other hand: if you’re just solving similar problem instances over and over again, in increasingly clever ways — what’s the point? Unless: (a) these problems are massive value drivers for some sector (b) your increasingly clever ways are driving down the cost, by reducing compute, or increasing speed.
I think PPLs which overspecialize to existing problems are useful, but have trouble inspiring new paradigms in AI (or e.g. new hardware accelerator design, etc).
Partially this is because there’s an upper bound on the inference complexity which you can express with these systems — so it is hard to reach cases where people can ask: what X application would this enable if we could run this inference approximation 1000x faster?
(Also note that inference approximations _can_ include neural networks)
That sounds very interesting! Is there something I could read more about it? Perhaps publications by you or anything like that you could recommend?
Modelling uncertainty sounds nice and sometimes is a goal in itself, but often at the end of the day you need a point estimate. And then IME all the priors, flexible models, parameter distributions, just don't add anything. You could imagine they do, with a more flexible model, but that is not my experience.
But then, PPL is just so much harder. The initial premise is nice - you write a program with some unknown parameters, you have some inputs and outputs, and get some probabilistic estimates out. But in practice it is way more complex. It can easily and silently diverge (i.e. converge to a totally wrong distribution), and even plain vanilla bayesian estimation is a dark art.
That said, I'll have to give this a spin sometime soon.
Practically, this means iteratively visualising your data and making informed judgements to even make your model run without humans thinking much about the structure of the data and model.
The ML promise is that there are robust models that you can feed nearly unlimited amounts of data to get better predictions.
Probabilistic modeling is better for people who have a fixed dataset they can visualise and fit an elegant model that incorporates lots of prior information about the problem of interest.
The white elephants are mostly the DSLs/frameworks that would have better off been torch/tensorflow extensions.
PPLs are useful when the data generation process is not easily represented by something like a simple multivariate Gaussian, etc. You find many good examples academic research, e.g. epidemiology.
Not only that, there is no reason why the math can't be done in a library and used in another language in the first place.
Why should they take off? At least for me personally it's not clear what the use case is, and this website answer exactly none of my questions.
There are multiple unbalanced categorical variables, so partial pooling helps a lot to infer the target in regions where the data is sparse.
Had fun using it when working my way through the statistical rethinking series.
[0]: https://turing.ml/
Dirichlet Process Mixture Models in Pyro - https://news.ycombinator.com/item?id=23398746 - June 2020 (2 comments)
Ask HN: What companies are using probabilistic programming? - https://news.ycombinator.com/item?id=17220861 - June 2018 (33 comments)
Uber Open Sources Pyro, a Deep Probabilistic Programming Language - https://news.ycombinator.com/item?id=15637329 - Nov 2017 (22 comments)
Pyro: PyTorch-Based Deep Universal Probabilistic Programming - https://news.ycombinator.com/item?id=15619634 - Nov 2017 (41 comments)
Uber AI Labs Open Sources Pyro, a Deep Probabilistic Programming Language - https://news.ycombinator.com/item?id=15619324 - Nov 2017 (1 comment)