Bayesian Optimization Book
bayesoptbook.com
bayesoptbook.com
I believe that modeling the noise directly using a 2nd Gaussian Process [1] could help, but I haven't gotten reliable results. I was hoping this topic would be addressed in the book, but don't see it.
It's a real pity, since a smart optimizer for very noisy functions would be really useful. I was trying to use it for chess engine tuning, since I know Deep Mind used it for tuning AlphaZero. I really wonder how they got it to work well.
[1] https://github.com/thomasahle/noisy-bayesian-optimization
I would also recommend using many more raw_samples and random_restarts in opimize_acqf(). 512 and 20, respectively, are good defaults. The current values are much too small to be able to effectively optimize the acquisition functions.
The SPSA method seems to work quite well with binary outcomes. This is what I was trying to beat. Unfortunately I was never able to converge faster than SPSA (or even close to that) even increasing the number of samples. There is a pretty long thread of us trying to make it work here: https://www.talkchess.com/forum3/viewtopic.php?f=7&t=71650&h...
I got some feedback form the botorch team back then: https://github.com/pytorch/botorch/issues/347#:~:text=thomas...
However, when the noise scale is either variable (but known) or can be modeled with a relatively simple (e.g., parametric) model, there may be some benefit to the added model complexity. Here you could include the parameters of the noise model into the model hyperparameters and proceed following the discussion in chapter 4. In doing so, I would be careful to ensure that the data actually support the heteroskedastic noise hypothesis.
Another approach that might be useful in some contexts is a heavy-tailed noise model such as Student-t errors (§§ 2.8, 11.9, 11.11).
So for example, rather than simulate an entire month in one shot, I'll simulate a day 30 times and therefore have a decent estimate of the noise for that result and be able to clearly distinguish the noise from the covariance of the Gaussian process.
The noise in these simulations can vary dramatically in parameter space (easily 10-100x), so it seems like it would be important to model.
(One might imagine a slightly more flexible model including a scaling parameter, replacing N with c²N and inferring c from data.)
It consists of friendly and readable chapters, each based on an example. Note that the book is very basic, and the mathematics & methods are much less sophisticated than other texts (such as the one posted above), but the concepts are elucidated very nicely.
I have some mathematical background in optimization, and I am quite curious how Bayesian optimization compares to other more established methods.
Powell's methods - COBYLA et al - address only (a) I think. But I might be wrong not having worked on them extensively. Do they also address the other challenges?
A relatively unique trait of Bayesian opt is the modelling of unexplored space. The attraction to exploring that space makes it actually less safe than other methods that do not explicitly care.
One could go through similar steps to model and generate safe steps in other methods. It doesn't seem specific to Bayesian opt. It might even be better since it's less likely to be so computationally expensive as the auxiliary opt within Bayesian opt tends to be.
You're correct that the approach taken in the linked paper could be adapted to increase the safety of other sequential design settings (when needed), assuming you have access to a model that can quantify uncertainty.
For those looking for an easier entry into Bayesian analysis, as a related topic and a starting place before embarking on Bayesian optimization, I would highly recommend "Statistical Rethinking" by Richard McElreath: https://xcelab.net/rm/statistical-rethinking/. Why I really like Richard's book is that it bypasses lot of the heavy mathematical/integral work, and goes straight into sampling - from my experience, you generally can't do integrals by hand, you describe your model in terms of a hierarchy of probability distributions and let an MCMC sampler take care of the rest. Richard's book touches upon causality (important and often overlooked topic in ML!), and you can follow his course online: https://github.com/rmcelreath/stat_rethinking_2022
EDIT: made my recommendation more explicit as a related topic.