Reflected Diffusion Models
aaronlou.com
aaronlou.com
https://www.google.com/search?q=%22An+astronaut+riding+a+hor...
The internet is already flooded with images generated by these common prompts.
It would be fun to have a space agency do a photoshoot of a real astronaut on a real horse ... just to complete the circle :)
There was a phase decades ago about "googlenope" -- defined as a phrase that google has never seen, so it would return zero results if searched with quotes.
Something similar for GenerativeAI may be worth coining?
I probably sound very negative but I feel that you have some responsibility to your readers if you publish your text.
I'm not, so as a reader here, I'm just happy that I didn't have to pay to get exposed to new challenging ideas. A lot of other science is sadly paywalled inside of journals.
Give it a few weeks for the fast followers to dumb it down for us.
As a researcher in this field, I found the article fairly understandable and accessible. I suspect that was the target audience.
The domain of diffusion models have a lot of terminology and concepts that are not explained here, so maybe that's why it's harder to understand?
I actually wrote a jupyter notebook based off of another notebook that showed that color loss was possible. My jupyter notebook extended the idea by allowing to specify the region and the particular color that the color loss would operate on.
I'm surprised that despite how much tooling currently exists, tooling for specifying colors is not really common or well known. I figured that controlnet would have something for this by now, but here we are...
It works with the Mikubill ControlNet plugin for A1111.
For example x̄ (x-bar) is used in statistics to represent the sample mean of a distribution. It doesn't matter what your distribution is, you'd likely call it x̄.
Whenever you see an equation mentioning "n", this will very often be the number of items in a set/list/array, not matter what that set is counting.
Usually these equations are very general, and they don't care what "n" actually is, and this way it's easy to see where you should plug that number into. If you look at the equation for x̄ (averaging) then you'll see "n" in there, and you already know what it's for, and it's incredibly concise.
If you're familiar with them they actually are named relatively well, and generally they're references that you then implement, they aren't meant to be the implementation themselves.
In the end the symbols on the paper are just a way to guide the reader to his own understanding of the underlying mathematical truth.
I bet most of us do not even think about that. Rho is a density, n is a natural number, epsilon is small, delta a perturbation …
Not quill and ink, but I took several programming exams at uni, where we had to write the answers, including pages worth of Java code with pen on three-layer carbonless copy paper...
This was just 15 years ago.
Not Java, but I have done productive brainstorming sessions where I handwrite Python or R onto paper, to figure out some kind of tricky problem or algorithm. Syntactically correct and indented too. It's not hard to handwrite code at all. Plus my "comments" can have arrows and diagrams and doodles and anything I like.
e^pii = -1
or
theNatural....Constant... = theCircleConstant... theSqureRootOfMinusOne = minusOn
or maybe you replace One with theNeutralElementForMultiplication
but maybe someone is too lazy and you need to give a good clear name for what multiplication is,
Sorry, but I could not help it, I seen this kind of comments multiple time, some people want to jump to the middle or end of a book/course and skip the hard work.
From what I have seen, they should likely be putting their code into ChatGPT and asking it to clean it up because it does a better job of writing functions with properly named variables.
Some dude will say that "i" is also used as the vector for the horizontal axis, i is sometimes used in induction as iterator too, so clearly it should be refactored to sqrtOfMinusOne. Also + and * are overloaded in math, so maybe they should be replaced with a clear name like "realNumberAdition"
(I know these are examples and not what you personally would actually name them, but ugh what a visceral reaction. :P )
but the argument was that what if a random say train driver looks at my code, he will not know what E could be.
This topic is complex. For a more thoughtful take, start with:
If mathematicians worked like software engineers they'd label things like "density[x, y]" instead of "\rho_{xy}" or at least stop using PDFs and use HTML with a tooltip over each variable name like we have in pretty much every modern editor, but we all know they hate readability.
Diffusion models are generative models that work by perturbing data using a stochastic process, then reversing this process to generate new data. However, when discretized into discrete steps (as in most practical implementations), the sampling process can negatively affect data quality. To alleviate this issue, a technique called thresholding is employed, resulting in better samples but breaking the theoretical framework.
Reflected Diffusion Models use a reflected stochastic differential equation (rSDE) to correctly model the thresholding process. This allows the models to train appropriately while respecting boundary constraints, like images with pixel values in the [0, 255] range.
The benefits of Reflected Diffusion Models include improved perceptual quality, correct handling of boundary constraints, better image generation with guidance, and general applicability to different shapes of data domains. This has the potential to expand the applications of diffusion models in fields such as image, language, and molecule generation.
I often wonder if Midjourney, behind the scenes, has already discovered something like this in their v5 release, which is substantially better at modelling things like hands.
It masquerades as summarizing the paper, but it's a random grab of whatever OP decided to jam into the chatbox that fits in the context window, (if you check, its mostly regurgitating the abstract).
However, it presents itself as a summary as if the entire paper was read.
I'm _really_ into AI stuff but the flippant nature its used at on HN has me scared. People here are generally well-educated, considered in their decisions, and are well acquainted with tech. If HN can't avoid glaring issues with it, who can?
The top example that came to mind was someone laboriously sharing a really nice data-driven comment, someone asks source? they say ChatGPT, and acted like we were trolling for saying it hallucinates.
It missed a _ton_ for me, it seems to pick out unique phrases and elegantly create a "skeleton" of the original.
> I did not use GPT-4 to summarize the whole paper.
> I don't think GPT-4 missed anything crucial from the blog.
?