DRµGS: Deep Random Micro-Glitch Sampling
github.com
github.com
As a bit of a layperson / just an AI integrator, I want to get some clarification of how this interfaces with a generative AI. I haven't dug too deep in and I have a mostly elementary understanding of neural nets. No need to _entirely_ dumb it down though, just needing a smarter person to confirm or correct my understanding. :)
Is this influencing the probabilities of token output? My understanding is that currently it is a static value that is essentially "how wild do you want it to be" where each generation has a different consistent static value of "wildness".
So rather than using a static value for an entire generation, this dynamically injects randomness in so rather than the output being 'monotone', it is more dynamic rather than by-the-book that AI tends to be?
From my skim, this proposes to inject randomness throughout the model (as opposed to just during sampling the classifier output) which according to the author (I'm paraphrasing) lets the full power of the model be used to make the most of that randomness, instead of just jittering the output. So the output is still random but it's a function of random noise being added through the model layers instead of just at the output, supposedly demonstrating better properties. I have no opinion about it as I haven't studied it.
Ah ok, so it's like the difference between greedily always taking the shortest / highest value vs using A* and exploring alternatives in hopes of finding something better? That probably doesn't line up as a metaphor as it isn't "exploring" or running multiple parallel generations.
And it employs the method in the layers so that the benefit can be amplified rather than only employing the method in the final layer / output when it selects the output to take
> benefit
In an intellectually honest way, it is not conclusive if there is a benefit to the approach in this repository. However, as a betting man: there is no benefit.
There are a lot of surprises! Riffusion and LCMs for me were super surprising. With no examples or meaningful investigation, I don’t think this is going to yield anything.
This dynamic modulation is seemingly at odds with mathematical soundness. Whether mathematical soundness should be a goal or not is a question to be answered empirically.
I had to chew on this for a minute or two to put it together. This explaination helped a lot.
Current LLMs first find the n most likely next token and only then gets randomness on the choice among the top-n. This injects randomness to the initial search for the top-n.
So I think, from my ignorance, that Stable Diffusion is already doing this in some sense
https://ggwiki.deepfreeze.it/index.php/Eron_Gjoni
Unclear how "the gamergate guy" he was Vs just "accidentally triggered gamergate". I'm not going down that rabbit hole though - where's Internet Historian when you need him?
If you actually googled "Internet Historian nazi" for a guy you don't even know who he is, that would be pretty funny.
It's very obvious from his videos that he is a 4chan troll. Didn't know he was full on Nazi though.
That summary seems a bit tendencious!
I'd think (based on nothing but my experience) that your average nurse might have higher likelihood of taking hard drugs than an average SWE :)
Some people think meth means “amphetamine” (although it usually methamphetamine), and there’s a lot of software engineers who take different types of prescribed amphetamines like Adderal and Ritalin.
And don’t even get me started on weed and millenials.
Coffee, psychedelics and alcohol are drugs just like heroin. Whether you believe they are useful to consume is a different matter.
In that sense, I wouldn't call SWEs a "drug using group."
We all know the pedantic meaning of the word "drug" - in that case, the whole humanity is a drug-taking species, since we prefer drug-altered consciousness to our natural one (e.g. by caffeine).