For training how do you get any kind of meaningful derivative with it?
(there is still a reduction in memory usage though (just not 24x):
> "Furthermore, Bop reduces the memory requirements during training: it requires only one real-valued variable per weight, while the latent-variable approach with Momentum and Adam require two and three respectively.")
Rule of thumb in optimization: real numbers are easy, integers are hard