Bayesian Flow Networks
arxiv.org
arxiv.org
You find more Twitter threads on the paper.
And here some more on Reddit: https://www.reddit.com/r/MachineLearning/comments/15rrljw/ba...
It's the first paper from Alex Graves since quite a while. Also, he is with NNAISENSE now.
Even quite broadly, Bayesian methods can be interpreted as a rate-distortion problem from Information Theory, which is an approach to lossy compression
https://imgur.com/gallery/kZa6VuZ
Visualization from the paper. Figure 20.
Just like we cure cancer in mice all the time we solve these datasets all the time. And it means nothing. Wait to see if the method is actually useful.
I understood (though only skimmed) that this paper is about generating MNIST and CIFAR-10, not the usual classifying?
I suppose you mean, "Yes you can do "hello world" but now show us something that makes your product worthwhile, a convincing use-case, a record worth noting".
Today, any bad idea you throw at a wall will work on MNIST and CIFAR-10. Even buggy broken code will generally work. We've advanced to the point where these datasets are totally meaningless.
In 1910 if you could show that your airplane idea could fly 100ft, you had something promising. In 2010 we're so used to building airplanes and understand the physics involved to the point where this isn't a meaningful test of the promise of your new aircraft concept.
It's worse than "in mice" in biology/cancer papers. We cure everything in mice. Basically nothing ever transfers to humans. Same with MNIST/CIFAR-10. Everything works, basically nothing matters.
Or another way to put it. A "Hello World"-level compiler would be interesting to report on in 1960. That same compiler would be trivial today and anyone could build it in minutes/hours.
(Note: There are many, many great optimization papers since 2014 - I just don't see them show up in general recipes in open source too often)
Would much prefer to see early work with solid small scale results on arXiV, than have people hold concepts for another 6 months scaling up. Let that be for a v2, if you cannot put early but concrete results on arXiV where else is there?
Recalling that a lot of nice papers are mostly MNIST / CIFAR-10 level results at first, followed by scale (thinking of VQ-VAE, PixelCNN / RNN, PerceiverAR, many others that worked well at scale later). That doesn't mean every result will scale up, but we have a lot of tricks to scale "small-scale" models using pretrained latent spaces and so on. The first diffusion results were also pretty small scale... different time but I don't think things are so different today.
That said, I can agree that you need to be a bit in the weeds on the research side to be diving deep on this - but I expect lots of followup clarifications or blog posts on this type of work.
Even if the methods require scaling up and significant engineering to put into production.
The parent is right to set expectations. Ideas like this often take 5-15 years to be refined into real products, if they make it at all.