Does this do something similar or is it fancier?
Does this do something similar or is it fancier?
On top of that, if the program branches on random numbers (which is common in simulations), that suffices for the maths to work out and you get an estimate of the asymptotic gradient (for samples -> infinity) of the original program, without any artificial smoothing.
So in short, I do think it is slightly fancier :)
As an aside, the combination "known distributions + automation" is covered in the Julia world by stochasticAD (https://github.com/gaurav-arya/StochasticAD.jl).
If so, does it scale for very branchy programs?
Do you have any comparisons to a Gibbs based approach for any of the use case examples?
We've haven't done a direct comparison to MCMC approaches yet, but it's on the Todo list. My intuition is that MCMC will win out for problems where finding "just any" local optimum is not good enough.