A Visual Exploration of Gaussian Processes
jgoertler.com
jgoertler.com
The article focused on the case where you have a finite number of test points, which is probably a good idea for an article like this. Still, there is another interpretation of Gaussian processes where they are actual stochastic processes (hence the name), a probability distribution over a set of functions.
I would have found an article that covered that interpretation even more helpful, although I'm not sure an easy-to-follow version could exist.
I tried to briefly go over the functional interpretation of GPs in this talk [2], although the book by Rasmussen and Williams does a much more thorough job [3] (free online, check out chapter 2 for this approach).
I'm happy to answer any questions about the differences. If you're a student/academic SigOpt is also completely free [4].
[0]: http://hyperopt.github.io/hyperopt/
[1]: https://sigopt.com/research/
[2]: https://www.youtube.com/watch?v=J6UcAdH54RE&list=PLbSwfqjMfj...
* SMAC (using Random forests) - [1]
* Using Deep Neural Networks - [2]
* Tree Structured Parzen Estimators - [3]
[1] https://www.cs.ubc.ca/~hutter/papers/10-TR-SMAC.pdf
[2] https://arxiv.org/pdf/1502.05700.pdf
[3] https://papers.nips.cc/paper/4443-algorithms-for-hyper-param...
In particular, there's a kernel we could call a "change point," which is a way to have a totally different model fitted before a point in time versus after. It's used frequently in Automatic Statistician fitted models (https://automaticstatistician.com/examples/) which in my opinion are the state of the art of what you can use GPs for. They also developed a "LISP" like representation of the GP kernels, which lets them sample functions, fit them, and publish the simplest ones.
You can see an example of, "Create GP functions and try fitting them" here: https://github.com/probcomp/notebook/blob/master/tutorials/e... . Near the end of the notebook, you can see the "source code" of the fitted GP function. Note this implementation supports change points, but does not happen to need them on the sample data.
But generally, I wonder in which applications GP shines.
On the one hand, with its emphasis on time series data, GP has a lot to offer to finance, especially options pricing. On the other hand, GP boils down to "the near future looks a lot like the near past," which most people already know.
Clearly, what we want to know is: when will change points occur? Whoever cracks that nut has found GP its breakthrough application.
It should be clear from the example that some form of fitting over the generation of "Gaussian Process Programs" is a good first step.
Maybe that's what the above comment meant by Bayesian optimisation.
(another project in this space is https://www.mathjax.org/)
https://planspace.org/20181226-gaussian_processes_are_not_so...
Would love to get more feedback as well!
Shouldn't it be the mode?