It claims that 'while diffusion works in high-dimensional datasets, it struggle in low-dimensional settings', which just makes no sense to me? Modeling high-dimensional data is just strictly harder than low-dimensional one.
Then when you read the intro, it's full of 'blanket statements' about diffusion, which have nothing to do with the subject, e.g. 'The challenge in applying diffusion models to low-dimensional spaces lies in simultaneously capturing both the global structure and local details of the data distribution. In these spaces, each dimension carries significant information about the overall structure, making the balance between global coherence and local nuance particularly crucial.'
I really don't see the connection between global structure/local details and low-dimensional data.
The graphs also make no sense. Figure 1 is just almost the same graph repeated 6 times, for no good reason.
It uses an MLP as its diffusion model, which is kinda ridiculous compared to what's the now-established architectures (U-net/ vision transformer based models). Also, the data it learns on is 2-dimensional. I get that the point is using low-dimensional data, but there is no way that people ever struggle with it. Case in point, they solve it with 2-layer MLPs, and it has probably nothing to do with their 'novel multi-scale noise' (since they haven't compared to the 'non-multiscale' version).
Finally, it cites mostly only each field's 'standard' papers, doesn't cite anything really relevant to what it does.
Overall, it looks exactly like what you would expect out of GPT-generated paper, just reshashing some standard stuff in a mostly coherent piece of garbage. I really hope people don't start submitting this kind of stuff to conferences.
Note that it cites TabDDPM in the related work, but that is for diffusione on tabular data! While most tabular data is low-dimensional, the type of low-dimensional data tackled in the paper is not tabular!
I'm also not quite sure how the linear upscaling is supposed to help, as it can be absorbed into the first layer of the following MLP, so I would rather think that the performance improvement (if any, the numbers are quite close and lack standard errors) is either due to the increased number of trainable parameters or some kind of ensembling effect (essentially the mixture of experts point made by the human authors).