The mathematical/conceptual error is that they are assuming each test point is added to the "post-hoc aggregated" prior when they evaluate the bound. This is analogous to including a test point in the training set. Another version of this error would be adding a kernel centered on each test point to a kernel density estimator prior to evaluating test set NLL. In this case, obviously the best kernel has variance 0 and assigns arbitrarily high likelihood to the test data.
You linked paper seems interesting on paper but does it bring any new SOTA?
Can you point me to some example of generated text this model produced? Something similar in quality to that unicorn story from GPT-2?
but does it bring any new SOTA?
Looking at their tables, seems so. The code is open source, and there's demo at https://mosaickg.apps.allenai.org/
So it is a totally new model that will probably keep evolving and being applied to more and more kind of NLP tasks. And it seems that it can have the first place on most NLP tasks, its empirically the breakthrough of the year.
It achieve this while having an extremely small number of parameters which shows: The model is smarter The model has room for more parameter hence even more accuracy!
Finally, theoretically it is a breakthrough as it is a port of a computer vision technology (variational autoencoders) to the NLP world. Actually it might be the successor to the Transformer paradigm. I wonder what pile of incremental improvement will researchers will be able to bring to it like they have on the Transformer paradigm (spanBert, XLnet, etc)
VAEs are not really exclusively a vision thing, they have been used in a variety of settings. Using VAEs for NLP is also nothing new, an early example is Bowman et al, 2015.