Riffusion – Stable Diffusion fine-tuned to generate music
riffusion.com
riffusion.com
Meanwhile, please read our about page http://riffusion.com/about
It’s all open source and the code lives at https://github.com/hmartiro/riffusion-app --> if you have a GPU you can run it yourself
This has been our hobby project for the past few months. Seeing the incredible results of stable diffusion, we were curious if we could fine tune the model to output spectrograms and then convert to audio clips. The answer to that was a resounding yes, and we became addicted to generating music from text prompts. There are existing works for generating audio or MIDI from text, but none as simple or general as fine tuning the image-based model. Taking it a step further, we made an interactive experience for generating looping audio from text prompts in real time. To do this we built a web app where you type in prompts like a jukebox, and audio clips are generated on the fly. To make the audio loop and transition smoothly, we implemented a pipeline that does img2img conditioning combined with latent space interpolation.
It would be very interesting to see what the stable diffusion community that is using automatic1111 version would do with this if it were made into an extension.
Guessing at the speed with which AI is developing these days someone is going to have the extension up in two hours at most.
> For every stunning example that gets passed around the internet, thousands of others sucked.
From personal experience this is simply untrue. I don't want to debate it because you seem to have strong feelings about the topic.
I don't feel strongly about this topic at all.
This iteration does, but that's an artifact of how it's being generated: small spectograms that mutate without emotional direction (by which I mean we expect things like chord changes and intervals in melodies that we associate with emotional expressions - elevator music also stays in the neutral zone by design).
I expect with some further work, someone could add a layer on top of this that could translate emotional expressions into harmonic and melodic direction for the spectrogram generator. But maybe that would also require more training to get the spectrogram generator to reliably produce results that followed those directions?
I don’t get the objection.
Prompting is an art and a science in its own right, not to speak of all the ways these tools can be strung together.
In any case, everything is a remix.
…implying there may be an art to AI art. Hmm.
Meanwhile, the degree to which it is off-puttingly hideous in general can be seen in the popularity of Midjourney — which is to observe millions of folks (of perhaps dubious aesthetic taste) find the results quite pleasing.
But the issue is not that the spectrogram is low quality.
The issue is that the spectrogram only contains the amplitude information. You also need phase information for generating audio from the spectogram
Audio spectrograms have two components: the magnitude and the phase. Most of the information and structure is in the magnitude spectrogram so neural nets generally only synthesize that. If you were to look at a phase spectrogram it looks completely random and neural nets have a very, very difficult time learning how to generate good phases.
When you go from a spectrogram to audio you need both the magnitudes and phases, but if the neural net only generates the magnitudes you have a problem. This is where the Griffin-Lim algorithm comes in. It tries to find a set of phases that works with the magnitudes so that you can generate the audio. It generally works pretty well, but tends to produce that sort of resonant artifact that you're noticing, especially when the magnitude spectrogram is synthesized (and therefore doesn't necessarily have a consistent set of phases).
There are other ways of using neural nets to synthesize the audio directly (Wavenet being the earliest big success), but they tend to be much more expensive than Griffin-Lim. Raw audio data is hard for neural nets to work with because the context size is so large.
But the same thing could be done as a post-processing step, finding points where the spectrum is changing fast and resetting the phases to make a sharper transient.
A neural vocoder such as Hifi-Gan [1] can convert spectra to audio - not just for voices. Spectral inversion works well for any audio domain signal. It's faster and produces much higher quality results.
It's definitely a useful approach as an early stage in a project since Griffin-Lim is so easy to implement. But I agree that these days there are other techniques that are as fast or faster and produce higher quality audio. They're just a lot more complicated to run than Griffin-Lim.
Models which operate directly on the time domain have generally had a lot more success than models that operate on spectrograms. But because time-domain models essentially have to learn their own filterbank, they end up being larger and more expensive to train.
The other place where phase is critical is in impulse sounds like drum beats. A short impulse is essentially just energy over a broad range of frequencies, but the phases have been chosen such that all the frequencies cancel each other out everywhere except for one short duration where they all add constructively. Without the right phases, these kinds of sounds get smeared out in time and sound sort of flat and muffled. The typing example on their demo page is actually a good example of this.
The simplest demonstration of it is the doppler shift. But it's not at all that simple because moving relative to the source the sound pressure and thus the perceived loudness also change, distorting the wave form, thereby introducing resonant frequencies. Now imagine that the transducer is always moving, eg. a plucked string.
The ideal harmonic pendulum swings periodically, only losing attenuation. But the resonant transducer picks up reflections of its own signal, like coupled pendulums, which are intractable according to the three body problem.
On top of that, our hearing is fine tuned to voices and qualities of noise.
When you take a window of samples of a signal, and run the FFT on it, for every frequency bin, the calculation determines what is the amplitude and phase of the signal. If you have a frequency bin whose center is 200 Hz, and there is a 200 Hz signal, then what you get for that frequency bin is a complex number. The complex number's magnitude ("modulus") is the amplitude of that signal, and its angle ("argument"d) is the phase.
If the signal is exactly 200 Hz, and if the successive FFT windows move by a multiple of 1/200th of a second, then the phase will be the same in succcessive FFT windows.
But suppose that the signal is actually 201 Hz: a little faster. Then with each successive FFT window, the phase will not line up any more with the previous window; it will advance a little bit. We will see a rotating complex value: same modulus, but the angle advancing.
From how fast the angle advances relative to the time step between FFT windows, we can deduce that we are capturing a 201 Hz signal in that bin (on the hypothesis that we have a pure, periodic signal in there).
How is the phase determined in the frequency bin? It's basically a vector correlation: a dot product. The samples are a vector which is dot-producted with a complex unit vector. The complex unit vector in the 200 Hz bin is essentially a 200 Hz sine and cosine wave, rolled into a single vector with the help of complex numbers. Sine and cosine are 90 degrees apart in phase, so they form a rectilinear basis (coordinate system). The calculation projects the signal, expressing it as a sum of the sine and cosine vectors. How much of one versus the other is the phase. A signal that is 100% correlated with the sine will have a phase angle of 0 degrees or possibly 180. If it correlates with the cosine component, it will be 90 or 270. Or some mixture thereof.
Because a complex number is two real numbers rolled into one, it simplifies the calculation: instead of doing a dot product with a sine and cosine vector to separately correlate the signal to the two coordinate bases, the complex numbers do it in one dot product operation. When we go around the unit circle, each position on the circle is cos(θ) + isin(θ). These complex values values give us samples of both functions. Exactly such values are stuffed into the rows of the DFT matrix: complex values from the unit circle divided into equal divisions.
If you look here at the definition of the ω (omega) parameter:
https://en.wikipedia.org/wiki/DFT_matrix
It is the N-th complex root of unity. But what that really means is that it is a 1/Nth step of the way around the unit cicrcle. For instance if N happened to be 360, then ω is the complex number whose |ω| = 1 (unit vector), and whose modulus is 1 degree: one degree around the circle. The second row of the DFT matrix has 1, ω, ω², ω³, ... the second row represents the lowest frequency (after zero, which is the first row). It captures a single cycle of a sine and cosine waveform, in N samples. The values in that row step around the unit circle in the smallest increment, so they go around the circle exactly once. The subsequent rows go around the circle in skipped steps, yielding higher frequencies: 1, ω², ω⁴ for twice around the circle; 1, ω³, ω⁶ for three times, ... we get all the harmonics up to our N resolution.
That pure sine wouldn't generate any artefacts. It would result in a 200Hz output from the AI if it throws the phase information out. You wouldn't hear a difference unless its an (aptly so called) complex signal. Eg. 200 and 201 Hz layered is an impure signal with a period below 1Hz, far outside the scope. Eventually the signals will cancel out completely. [1]
The important point is, I think, that FFT doesn't simply look at the offset aka phase. Rather, 201 Hz looks like a 200 Hz that is moving. So it encodes phase-shift in the delta of the offset between two windows. For a sum of 200 and 201 Hz it has to assume that the magnitude is also changing, which I find entirely counterintuitive.
From the mathematical perspective, this seems like a borring homework, far detached from accoustics. So, I don't know. The funny thing is that rotation is very real in the movement of strings. If the orbit in one point is elliptic, that's like two sinusoids at different magnitudes offset by some 90 degree, in a simplified model. But it has nearly infinite coupled points along its axis. As they exite each other, and each point has a different distance to the receiver, that's where phase shift happens.
> If you look here at the definition of the ω (omega) parameter
I wasn't going to make drone, but I will take a look.
1: https://graphtoy.com/?f1(x,t)=100*sin(x)&v1=true&f2(x,t)=100...
Looking back at image generation just a year or two ago and people would have said similar things.
Not hard to imagine the trajectory of synthesized audio taking a similar path.
Calling it a Deep Truth might be a bit of an emotional marketing spin but the concept is very exciting nonetheless I believe.
-- [I copied and pasted the below to the above and then corrected it. Below is the original version. This is how I dictate to Google sometimes, on Android. Normally I would have further edited the above but in this case I wanted to show how far basic things like dictation still have to go. By the way I dictated in a completely quiet room. I can't wait for more advanced AI like ChatGPT to take my dictation.]
Alright so this is a pretty amazing our new development period I want to tell you something about out why the state of the heart is is in a i period when you wrote that it is a deep truth it was before I actually listen to The Pieces, I have just read the descriptions period at the time, I thought that you were probably right because I was thinking that music is only pleasing because of the structure of our brains it's not like vision where originally we are interpreting the world and that Where Art comes from music is purely so dove abstract or artistic period however, after I listen to the pieces, I realise that they really sound exactly like the instruments that are making the physical noises period for example it really sounds exactly like a physical piano period so I don't know about out a deep truth karma but it does seem that there is a physical sense that the music are represents which it can successfully mimic using this essentially image generating capability period one thing about all of these amazing AI development, is that I still make some long comments by dictating to Google. When it first got to the point that it was able to catch almost everything then was saying I was absolutely blown away period however, it's really not that good at taking dictation karma and I have to go back and replace each and every individual, and period with with the corresponding punctuation mark period seeing such an amazing developments happening month after month year after year ear makes me feel like we are really approaching what some people have called the singularity period when I read about out net positive fusion being announced my first Instinct was to think oh of course it's now that that chat GPT is available of course announcing a major fusion breakthrough would happen within in days to weeks it just makes perfect sense DJ eyes can solve problems that have have confounded scientists for decades period to see just how far we still have to go take a look at how this comment red before I manually corrected it to what I had actually set
It might be the last step this AI needs to bring some extra clarity to the output.
https://colab.research.google.com/drive/1FhH3HlN8Ps_Pr9OR6Qc...
Example prompt: “deep radio host voice saying ‘hello there’”
Kind of like a more expressive TTS?
That said, our GPUs are still getting slammed today so you might face a delay in getting responses. Working on it!
On a serious note, I'd really love some advice from you on time management and how you get so much done? I love Skydio and the problems you are solving, especially on the autonomy front, are HARD. You are the VP of Autonomy there and yet also managed to get this done! You are clearly doing something right. Teach us, senpai!
How much data is used for fine tuning? Since spectrograms are (surely?) very out of distribution for the pre training dataset, how much does value does the pre training really bring?
One thing that's very important though is the language pre-training. The model is able to do some amazing stuff with terms that do not appear in our data set at all. It does this by associating with related words that do appear in the dataset.
In my experience (CNN based imagery segmentation) proven architectures (e.g. U-Net) performed similar with or without fine-tuning existing models (that have been mostly trained on imagenet, citiscapes, etc.) IF the domain was rather different.
At least in the field of imagery segmentation there is not much of a point in fine-tuning an off-the-shelf model on let's say medical imagery.
So maybe it's the same for the stable diffusion model. I don't see how some knowledge about the relationship between the prompt and given imagery describing that prompt should help this model map the prompt to a spectrogram of the given prompt.
Why was stable diffusion able to generate spectrograms? Because it was fed some. Presumably, those original spectrograms were scraped with little concern over creators' permissions, just like it has been for artists' work in order to produce art-looking image generation. Please, research what has been happening in the art community lately. https://www.youtube.com/watch?v=Nn_w3MnCyDY
A protest on ArtStation has been shown to influence Midjourney's results, proving that huge amounts of proprietary work are constantly scraped without the creators' permission. AIs like these work so well just because they steal and remix real artists' work in the first place. There are going to be legal wars about this.
Stable Diffusion doesn't have an official music generation Ai precisely because it couldn't train it with the same approach without being sued by music labels right away, while isolated artists don't have the same power.
So, back to my question: have you wondered whose work is Stable Diffusion remixing here? Your endeavour is great technically, but as we progress into the future we have to be more aware of the ethical implications that come with different forms of progress.
You could try to base your project on a collection of free-to-use spectograms, and see how it performs. If you do, I think it could actually be very interesting and useful to discuss the results here on Hacker News.
Cheers!
It's impressive what it produces, but I think it probably lacks substance in the same way the visual AI art stuff does. For the most part, it passes what I call the at-a-glanceness test. It's little better than apophenia (the same thing that makes you see shapes in clouds, faces in rocks, or think you've recognised a familiar word in a foreign language; the last one can happen more often though).
So, I think these tools will be used to do background work (ie for visuals maybe help with background tasks in CGI or faraway textures in games). I know less about audio, but I assume it could maybe help a DJ create a transition between two segments they want to combine, as opposed make the whole composition for them, but idk if that example makes sense.
Now, onto a more human point: I think that people often listen to music because it means something to them. Similar for people who appreciate visual art.
I also love interactive and light art, and I love talking to other artists at light festivals who make them because of the stories and journeys behind their art too. Humans and art are a package deal, IMO.
Edit: typos and to add: Also, I think prompt authorship is an art unto itself. I'm amazed what people can craft with it, but I'm more impressed by the craft itself than the outputs. Don't get me wrong, the outputs are darn cool, but not if you look closer. And it's impossible to look beneath the surface altogether, as there is nothing in the output but the pixels.
I don't think this stuff is a threat to genuinely innovative, thoughtful, meaningful work, but that's the top of the market.
That being said the bottom of the market is how a lot of artists make their living, so this is going to deeply impact all forms of art as a profession. It might soon impact programming too because while GPT-type systems can't do advanced high level reasoning they will chop the bottom off the market and create a glut of employees that will drive wages down across the board.
Basic income or revolution. That's going to be our choice.
There will never be a revolution and there's no such thing as late capitalism. Well, not if the Fed does their job.
If AI completely eliminates low skill art labour from the job pool, it's not like those affected by it are gonna disintegrate, riot, and restructure society. They have the choice of filling an art niche an AI can't or they can spend that time learning other, more in-demand skills. This also ignores that fact that some companies would rather reallocate you to more profitable projects even if your art skills don't change.
Selling a product with relative value like a painting or a sculpture will always be an uphill battle. Now that there's more competition from AI, it just gives artists/businesses incentive to find what people want that an AI can't deliver. Worst case scenario, employment rates in this sector are rough while the market recalibrates. Interested to see how these technologies develop.
People don't have unlimited ability to learn new skills. Training takes time, and someone who spent several years honoring their craft won't be able to pick up a new skill overnight.
On top of that, people have preferences regarding their work – even if someone has the ability to do a different work, they might find it less meaningful and less satysfing.
Finally, don't ignore the speed at which AI capabilities improve. Compare GPT-1 with the current model, and how quickly we got here. Eventually we'll get to a point where humans just won't be able to catch up quickly enough.
I don't know how many custoners feel the same way, but I won't be purchasing any AI art or music or knowingly giving it any of my attention.
Although personally, I think using "AI art" to create impossible photographs is more interesting and doesn't compete with illustrators as much.
Maybe a parallel would be furniture; there are people who buy hand crafted furniture but it's kind of a luxury. Most people just have Ikea and wouldn't pay more for the same (or have less good furniture) just to get some artisanal dinner table chair.
I'm afraid in near future we will all bombarded with AI-created music, art and text whether we want it or not.
The big problem is that it's hard to filter through the huge amount of content being produced by humans to find things that you'll like, so we rely on kingmakers curating the culture. This means a few huge winners taking all and a lot of great creative work at the same level going unrewarded. If we can solve the content discovery problem in a more personalized and fair way and make it easier for people to support creators they like that would go a long way towards cushioning the job losses that AI will create.
I’ve been trying to play through the scenario in my head. At least in terms of software developers being replaced by AI, I think we’re going to first see AI doing work in parallel or under monitor by humans. Basically, Google will take AI and send it off to do work that they lack the staff to do. Now, on the other hand, they could also temporally play it out where first they feign an inability to staff people due to finances so there are layoffs/terminations, and then maybe a quarter later they replace those people with low cost AI compute time that is orders of magnitude more productive.
In any case, AI disrupting people’s ability to feed, shelter, and clothe themselves is sure to trigger a pretty brutal and hostile response, which would be grounds for legislation and perhaps a class war.
The weird part is that if the potential of AI is truly orders of magnitude expansion beyond what we already have, then the longterm surely has room for a tiny little mankind fief. But, in order to get to the long term our hyper-competitive technocratic overlords may strangle out part of or all of the rest of us while justifying accelerating through the near-term window to achieve AI-dominance.
So many menial jobs are kind of like basic income anyway - you put in 2 hours of actual work to pad out the entire day at some shitty low end job, knowing all the time that your contribution isn't valued and that if your employer ever got their shit together your job wouldn't even be needed, and the robots are coming for it anyway. You get paid a small amount for doing nothing much useful.
The rich today are rich largely because they or their ancestors were plunderers. Perhaps they plundered the planet, exploiting the cheap energy that fossil fuels provide. Perhaps they plundered our social cohesion building skinner boxes that manipulate the minds of millions just to gain eyeballs and clicks.
Why should the bill for these past excesses fall on those who never benefited from them? In previous times, a young person of average intellect could get a job on a farm or factory and be a valued contributor. What happens when automation removes the last of these jobs - do we really expect people to put up with more and more menial and slavish existences?
Basic income is, like carbon taxes, an obvious solution. Maybe it will take off when a tipping point arrives - when the rich class decides that their repugnance to giving someone a "free ride" is overtaken by their need to have masses dulled and stupified, sitting at home with blinds drawn in front of their playstations, so they don't revolt at the obvious unfairness of the world.
Work gives people dignity. And idle hands are the devil's plaything. Put that together and UBI would be a disaster. Also, it's not "the rich class" who decides whether or not to bestow such a lifestyle on the masses... that in itself is a conspiratorial line of thought. Right down that road is the thought "hey, this UBI isn't enough!"
Jobs disappear. Other jobs replace them. Often, jobs are not fun, and often they feel meaningless, but working is still much more dignified than not working. Raising generations who've never worked and simply take their UBI and breed - what would even be the point of educating such people? Eventually they'd just be totally disposable and, no doubt, be disposed of.
I fear you are right. But neither of those is going to be an easy transition, if only because the effects of all this innovation is felt disproportionally by people in countries where such a revolution will not do anything to give solace.
Basic income assumes that the funds to do this are available and revolution assumes that the powers that be are the parties that are in the way of a more equitable division of the spoils. Neither of those are necessarily true for all locations affected.
Evolution.
We have such vast wealth and our historic methods for trying to make sure most people are taken care of are failing us. Those methods were rooted in the nuclear family with a head of household earning most of the money and jobs designed with an assumption that he had a full-time homemaker wife buying the groceries, cooking the meals etc so he could focus on his job.
We need jobs to evolve. In the US at least, we need to move away from tying all benefits (such as medical benefits and retirement) to a primary earner. We need to make it possible to live a comfortable life without a vehicle. We need to make it possible for small households to find small homes that make sense for them, both financially and in terms of lifestyle.
There is a lot we can do to make this not a disaster and make it possible for some people to survive on very little while while pursuing their bliss so that we stop this trend of pitting The Haves against The Have Nots and make the current Have Nots a group that has real hope of creating their own brilliant tech or such someday while not being utterly miserable if they aren't currently wealthy.
Knowledge is power.
Basic income sounds good in theory in some imaginary futuristic society of harmony and grace.
In real life, it's a way for the masses be controlled down to your very substinence by the state. Where the state is basically an intermediary for big private interests and lobbies.
Third option. Mandatory 4 day weeks.
Although I'd specify it as no one can work more than X hours a week.
And then adjust X down - or up in short timescales but likely down overall - as needed.
The competition is for "work". If AI is taking large chunks of "work" off the table. Spread the rest of it around.
Now notionally people will tell you that there is no finite "work" limit. You are effectively limiting competition.
To which I say - good. The rat race IS the competition. Don't we all want to slow it down a little? If F1 can put limits on a race, we should too for humanity.
Work smarter, not harder.
> Basic income or revolution. That's going to be our choice.
I'm definitely pro basic income, but I've heard an interesting remark a few weeks ago. And that's that COVID was kind of a UBI experiment (in the US), albeit very limited, and it turned out that if people don't have to worry about making a living and don't have a job to work in then they'll start do stupid things on the internet. Like make up stupid conspiracy theories about vaccines. I can't remember who said this, it was one of the guests on Lex Fridman's podcast. I'm also not sure if it's a valid analogy but reminds me of Vonnegut's Player Piano.
What this means for the future is maybe a little more unsettling however.
I don't know of existing tech that can generate actual good mashups in realtime given arbitrary mp3s, but this has promise!
neither a set of MP3 nor a set of spectrograms from MP3s supplies the function arguments
or a connection to a path that uses that function
Regarding "can you use this to help you through?" - yeah, you could probably use it as a source of inspiration, but at the risk of getting sued by someone whos music you didn't even know you were copying...
With Stable Diffusion and similar generative systems we have seen a leap in generative art/media, partially with significant improvements within a few months. What makes you think this was the last or only leap in the next 5 to 10 years? As if progress would just stop here? Huh?!
Do you think we hit a ceiling were progress is only tangential? A line which is impossible to cross? Otherwise I dont get this mindset in the face of these modern generative AI systems popping up left and right.
Having said that, it's not my job, and I can see where the issues lay there.
Raves that have no human DJs and never stop.
Good? Bad? Not for me to say, really, the most I make at it is a couple hundred bucks for two nights in a bar playing FM radio hits, and there's lots of people younger than me who like that music, so obviously I'm doing it for different reasons and I don't anticipate losing access to as many bar gigs as I want for the rest of my life.
But certain genres are very tolerant of low-effort music, and I think the people who are monetizing low-effort music are gonna lose their income streams. I do different things than those people, but I still consider them compatriots, even if I don't care for their art.
Well, they will, if the AI plus 1 human deciding "where and how to use them" can replace producers and musicians playing...
For some reason people tend to think that these tools/AI/ML systems will never be good enough to do their job (or a specific job). This argument can take different forms, sometimes stating that it will just do the boring part of the work (e.g. with programming) or that it will still need human creativity (maybe, but not necessarily and that's not the point) or that it will just replace low level, unskilled or mediocre professionals. And somehow everyone thinks they are not mediocre (i.e. average). But even these assumptions are unfounded. Why would anyone think that these systems will top out below their skill levels? Why would anyone think that they can't become superhuman?
They did in chess, go, I think poker too. Not to mention protein folding. And without much of a hitch between mediocre/good enough and superhuman. Because that difference is just interesting for us, but doesn't necessarily mean that there is huge step, that the system needs to undergo serious development and that it would take a long time. (Like decades or so.) People thought that was the case when AlphaGo beat Fan Hui saying that Lee Sedol was a completely different level. Which, of course, he is. Still, it just took DeepMind half a year to improve alphago to that level.
So yeah, you can be pretty sure that if this track (no pun intended), if this solution is good enough then it will quickly evolve into something that will replace some music creators.
It'd be really cool if you could implement an MS paint style spectrum painting or image upload into the web app for more "manual" sound generation.
Have you already explored doing the same with voice cloning?
2 amazing AI projects. Huge respect :)
How many paired training images / text and what was the source of your training data? Just curious to know how much fine tuning was needed to get the results and what the breadth / scope of the images were in terms of original sources to train on to get sufficient musical diversity.
- My strongest suggestion is finding some strategy for smoothing over the sometimes harsh-sounding edge of the sample window - Perhaps it could be filling in/passing over segments of what is sounded to user as a larger loop? Both giving it a larger window to articulate things but maybe also showcasing the interpolation more clearly… - Tone control may seem challenging but I do wonder if you couldn’t “tune” the output of the model as a whole somehow (given the spectrogram format it could be a translation/scale knob potentially?)
Also the key question is - would something like this ever produce something as hauntingly beautiful and unique as classical music pieces?
Absolutely blows my mind.
Random forests: Ko and Breiman's, not really Bell Labs and UC-Berkeley
Transistors: Bardeen, Brattain, and Shockley, not really Bell Labs (thank the Nobel Prize for that)
UNIX: Primarily Bell Labs, but also Ken Thompson and Dennis Richie (this is a hard one)
GPT-n: OpenAI, not really any individual, and I can't seem to even recall any named individual from memory
I had a look around several months ago, and it seems like everything is locked behind SaaS APIs.
Steve Blum: https://fakeyou.com/tts/result/TR:xmjjq9ty0hnsyjrjnw806k6rnp...
Furiously working on voice-to-voice (web, real time, and singing!) Should be out the door tomorrow!
Have you released a tool for volumetric capture? I'm applying this to LED lighting fixture setup for tv/film/live shows and 3D positioning is the last step to fully automated configuration.
My goal is real-time sync between 3D model and real world.
How do you think human endeavours progress other than by small steps?
I can't wait to hear some serious AI music-making a few years from now.
The amazing thing is that the current diffusion models are so good that the spectograms are actually reasonable enough despite the small room for error.
I think this will be particularly useful for musical compositions in movies and film, where the producer can "instruct" the AI about what to play, when, and how to transition so that the music matches the scene progression.
If you're interested, the idea of applying Image processing techniques to Spectrograms of audio is explored in brief in the first lesson of one of the most recommended AI courses on HN: Practical Deep Learning for Coders https://youtu.be/8SF_h3xF3cE?t=1632
Also, the music doesn't sound "Lofi" because it's generated by algorithms. A lot of hard work and software goes into taking a clean, pitch-perfect digital signal and making it sound like something playing on a record player from the 70s.
edit: added a bit more to the thought
I just ran a quick Google Scholar search, and the first result is https://ieeexplore.ieee.org/abstract/document/5672395
This is from 2010. I didn't go looking, but it wouldn't surprise me if the idea is older than that.
https://en.wikipedia.org/wiki/UPIC
Edit: my favorite of all these systems was Chris Penrose's HyperUPIC which provided a lot of freedom in configuring how the analysis and synthesis steps worked.
Sure, AI can do lots of things well. But would you rather live in a world where humans get to do things they love (and are able to afford a comfortable life while doing so) or a world where machines do the things humans love and humans are relegated to the remaining tasks that machines happened to be poorly suited for?
I think that kind of attitude is defeatist - it's implying that humans will be stopped from making music if AI learns how to do it too. I don't think that will happen. Humans will continue making music, as they always have. When Kraftwerk started using computers to make music back in the 70s, people were also scared of what that will do to musicians. To be fair, live music has died out a bit (in a sense that there aren't that many god-on-earth-level rockstars), but it's still out there, people are performing, and others who want to listen can go and listen.
Maybe consumers will start consuming more and more AI music, instead of human music [0], but the worst thing that can happen is that music will no longer be a profitable activity. But then again, today's music industry already has some elements of the automation - washed-out rhythms, sexual thematics over and over again, re-hashing same old songs in different packages... So nothing's gonna change in the grand scheme of things.
For me, the worst that could happen is that people spend so much time listening to AI generated music, that human musicians can no longer find audiences to connect to. It's not just about economics (though that's also huge). It's the psychological cost of all of us spending greater and greater fractions of our lives connected to machines and not other people.
Black metal community, for example, has always rejected all forms of "automation" and considers it not kvlt - rawness is a sought-after quality, defined as having people performing as close to the recording equipment as possible.
There's also a rapper named Bones (Elmo O'Connor) who's never signed a contract with a label, does only music he wants to do, releases a couple albums every year. There's something about his approach that makes his music sound very organic and honest. I listen to him more than I listen to any mass produced rapper.
So in conclusion, music was always about people. Unless AI reaches AGI level, I don't think it will ever impact music enough to kill all audience.
Advancing AI capabilities in no way detracts from this. You talk about humans being "relegated to the remaining tasks" - but that's a consequence of our socioeconomic system, not of our technology.
Those two are profoundly intertwined. Our tech affects our socioeconomic systems and vice versa.
I don't see this as a threat to human ingenuity in the slightest.
It absolutely sucks at cymbals, though. Everything sounds like realaudio :) composition's lacking, too. It's loop-y.
Set this up to make AI dubtechno or trip-hop. It likes bass and indistinctness and hypnotic repetitiveness. Might also be good at weird atonal stuff, because it doesn't inherently have any notion of what a key or mode is?
As a human musician and producer I'm super interested in the kinds of clarity and sonority we used to get out of classic albums (which the industry has kinda drifted away from for decades) so the way for this to take over for ME would involve a hell of a lot more resolution of the FFT imagery, especially in the highs, plus some way to also do another AI-ification of what different parts of the song exist (like a further layer but it controls abrupt switches of prompt)
It could probably do bad modern production fairly well even now :) exaggeration, but not much, when stuff is really overproduced it starts to get way more indistinct, and this can do indistinct. It's realaudio grade, it needs to be more like 128kbps mp3 grade.
Well no wonder, it has absolutely no concept of composition beyond a single 5s loop, if I understand correctly.
> It absolutely sucks at cymbals, though. Everything sounds like realaudio :)
> It could probably do bad modern production fairly well even now :) exaggeration, but not much, when stuff is really overproduced it starts to get way more indistinct, and this can do indistinct. It's realaudio grade, it needs to be more like 128kbps mp3 grade.
I haven't sat down yet to calculate it, but is the output of SD at 512*512px at 24bit enough to generate audio CD quality in theory?
And I suspect this will always have phase smearing, because it's not doing any kind of source separation or individual synthesis. It's effectively a form of frequency domain data compression, so it's always going to be lossy.
It's more like a sophisticated timbral morph, done on a complete short loop instead of an individual line.
It would sound better with a much higher data density. CD quality would be 220500 samples for each five second loop. Realtime FFTs with that resolution aren't practical on the current generation of hardware, but they could be done in non-realtime. But there will always be the issue of timbres being distorted because outside of a certain level of familiarity and expectation our brains start hearing gargly disconnected overtones instead of coherent sound objects.
What this is not doing is extracting or understanding musical semantics and reassembling them in interesting ways. The harmonies in some of these clips are pretty weird and dissonant, and not what you'd get from a human writing accessible music. This matters because outside of TikTok music isn't about 5s loops, and longer structures aren't so amenable to this kind of approach.
This won't be a problem for some applications, but it's a long way short of the musical equivalent of a MidJourney image.
Generally we're a lot more tolerant of visual "bugs" than musical ones.
But of course something like this, which only thinks in 5s clips can not generate a larger structure, like even a simple song. Maybe another algorithm could seed the notes and an algorithm like this generates the sounds via img2img.
The timbral qualities of the posted samples remind me of some of the stuff I heard from Aphex Twin, like Alberto Balsalm. Not accessible by a long shot but definitely human
This show me that Stable Diffusion can create anything with the following conditions:
1. Can be represented as as static item on two dimensions (their weaving together notwithstanding, it is still piece-by-piece statically built)
2. Acceptable with a certain amount of lossiness on the encoding/decoding
3. Can be presented through a medium that at some point in creation is digitally encoded somewhere.
This presents a lot of very interesting changes for the near term. ID.me and similar security approaches are basically dead. Chain of custody proof will become more and more important.
Can stable diffusion work across more than two dimensions?
One feature of SD not discussed much in my opinion, is the deterministic key it provides to the user. This is what enables the smooth transition in every second of music it generates, and the next second in time. Moving the cursor of latent space between in a minimal way, creating the next piece of information and change it ever so slightly, it definitely sounds good to the human ears.
[0] https://www.soundonsound.com/techniques/synth-school-part-7
Agree totally. Before SD was created, i thought that it is impossible to replicate a prompt more than once. Deterministic/fixed seed is a big innovation of SD, and how well it works in practice, it is simply amazing.
From the article: >Those who tried this method, however, soon found that, without analogue filters to run through the harmonic content of waveforms, picking out and exaggerating their differing compositions, most hand‑drawn waveforms sounded rather ordinary and often bland, despite the revolutionary way in which they were created.
Yes, the technique which the people of riffusion created, displayed to everyone, and shared it as well, it is the holy grail of electronic music synthesis. I would imagine it is has some way to go before it is applied to electonic music effectively, integration with some tools, practice of musicians on the new tool, some fine tunning etc.
Something likely to be affected: anything using biometrics as a password instead of a name.
For example, you could take half the previous spectrogram, shift it to the left, and then use the inpainting algorithm to make the next bit... Do that repeatedly, while smoothly adjusting the prompt, and I think you'd get pretty good results.
And you could improve on this even more by having a non-linear time scale in the spectrograms. Have 75% of the image be linear, but the remaining 25% represent an exponentially downsampled version of history. That way, the model has access to what was happening seconds, minutes, and hours ago (although less detail for longer time periods ago).
But perhaps plain stable diffusion wouldn't work - you might need different neural networks trained on each "zoom level" because the structure would vary: music generally isn't like fractals and doesn't have exact self-similarity.
One of the samples had vocals. Could the approach be used to create solely vocals?
Could it be used for speech? If so, could the speech be directed or would it be random?
How many paired training images / text and what was the source of your training data? Just curious to know how much fine tuning was needed to get the results and what the breadth / scope of the images were in terms of original sources to train on to get sufficient musical diversity.
If we could measure certain kinds of productivity it might even be useful as a way to "extend" certain highly productive ambient environments a la "music for coding".
Or at a house party, club or restaurant... as more people arrive or leave and the energy level rises or declines..or human rhythms speed up or slow down...so does the music...
A couple of ideas that come to mind:
- I wonder if you could separate the audio tracks of each instrument, generate separately, and then combine them. This could give more control over the generation. Alignment might be tough, though.
- If you could at least separate vocals and instrumentals, you could train a separate model for vocals (LLM for text, then text to speech, maybe). The current implementation doesn't seem to handle vocals as well as TTS models.
But what if what if a model was trained not on single images, but animated sequential frames, in sets, laid out on a single visual plane. So a panel might show a short sequence of a disney princess expressing a particular emotion as 16 individual frames collected as a single image. One might then be able to generate a clean animated sequence of a previously unimagined disney princess expressing any emotion the model has been trained on. Of course, with big enough models one could (if they can get it working) produced text prompted animations across a wide variety of subjects and styles.
Reminds me of the soundtrack to Nier Automata which did a similar thing: https://youtu.be/8jpJM6nc6fE
Who else will AI make looking for a new job?
GPT-3, what policy should we apply to increase tax revenue by 5% given these constraints?
GPT-3, please tell me some populist thing to say to win the next election, or how should I deflect these corruption charges.
One of the automated prompts was "Eminem anger rap", I'm confident if you had showed me the audio without the prompt I could identify which artist it sounded like.
And this is just a basic first attempt at reusing a tool not even designed for audio. I can only imagine how powerful it could be after some trivial revisions like using GPT-3 to generate coherent lyrics.
While I've seen a lot of cool stuff which helps generation for hobby projects or smaller indie games it's nowhere near the quality and consistency needed to come close to the work of a skilled human artist at a larger studio.
To give you an example join the public discord of https://www.scenario.gg and check out results. Come back and tell me those aren't on a professional level.
I am not saying that designers won't be needed anymore but AI is definitely able to replace jobs and speed up progress in game development.
Great work!
...OK, need to create a band name generator to work in tandem with this thing. Let's see what one of its brethren in ML makes of it...
- "Echoes of the Past": This name plays on the idea of Gregorian chanting, which is often associated with the distant past, and combines it with the intense and aggressive sound of death metal.
- "The Order of the Black Chant": This name incorporates elements of both the religious connotations of Gregorian chanting and the dark, heavy sound of death metal, creating a sense of mystery and danger.
- "Foretold in Blood": This name evokes both the ancient, mystical nature of Gregorian chanting and the violent themes of death metal, creating a sense of ancient prophecy coming to pass.
- "Crypt of the Silent Choir": This name brings together the eerie, otherworldly sound of Gregorian chanting with the underground, underground feel of death metal, creating a sense of hidden secrets and forbidden knowledge.
"The Order of the Black Chant" it shall be.
Edit: After a bit of playing around I at least got some credible results with "electric guitar solo, glam rock". It also understands grunge.
Interesting that it gets so good results with just fine tuning the regular SD model. I assume most of the images it's trained on are useless for learning how to generate Mel spectrograms from text, so a model trained from scratch could potentially do even better.
There's still the issue of reconstructing sound from the spectrograms. I bet it's responsible for the somewhat tinny sound we get from this otherwise very cool demo.
I'm not so sure. Considering how successful AI-driven social media feeds are, which already include substantial AI-generated content, why would a feed consisting entirely of such content be any less successful? The quality will only keep increasing.
> Art will not be challenging but designed by the algorithm to get us to like it.
I don't think these advancements are a threat to art created by humans, just as any art created by humans isn't a threat to other art. It's just... more art.
Eventually, AI will be capable of being truly creative, instead of being trained on human art and producing permutations of it, which will also be wonderful.
The role of humans will be to train these models to produce art we find enjoyable. Imagine if your AI media feed was an infinite stream of artworks personalized just for your taste. It will be TikTok on steroids. I can't say I'm thrilled by that prospect, because it will also be used for exploiting users, but the entertainment potential is huge.
This is the most exciting thing I’ve seen in ages as it shows we may be on the verge of the next wave of new technology in music that will allow all sorts of weird and wonderful new styles to emerge. I can’t wait to see what these tools can do in the hands of artists as they become more mainstream.
Might be a traffic thing?
Edit: Works now. A bit laggy but it works. Brilliant!
{"data":{"success":true,"worklet_output":{"error":"Model version 5qekv1q is not healthy"},"latency_ms":530}}
There was a purple victorian house in Colorado Springs where the living room was converted into a record and cd store called Life By Design. I picked up these albums and a ton of other obscure music there. I was so happy to not have to drive all the way up to Wax Trax in Denver to be able to discover new artists.
The audio quality is surprisingly good, but does sound like it's being played through an above-average quality phone line. I bet you could tack on an audio-upres model afterwards. Could train it by turning music into comparable-resolution spectrograms.
I'm quite impressed that there was enough training data within SD to know what a spectrograph looks like for the different sounds.
In theory the tone quality is not an objection here. When it sounds bad it's because it's 512x512, because the FFT resolution isn't up to the task, etc. People cling to very inadequate audio standards for digital processing, but you don't have to.
Played a bit with the very impressive demos, now waiting in queue for my very own riff to get generate.
Great as this is, I'm imagining what it could do for song crossfades (actual mixing instead of plain crossfade even with beat matching).
IE, use a polar coordinate system where angle 12 oclock is 440hz, and the 12 chromatic notes would be mapped to the angle of the hours. Maybe red pixel intensity is bit mapped to octave, IE first, third and eight octave: 0b10100001.
Time would be represented by radius. Unfortunately the space wouldn't wrap nicely like if there was a native image format for representing donuts.
I think I'm beginning to crack the code with this one, here's my attempt at something like a "DJ set" with this. My goal was to have the smoothest transitions possible and direct the vibe just like I would doing a regular DJ set. https://www.youtube.com/watch?v=BUBaHhDxkIc
I wonder if this could be the future of DJing or perhaps beginning of a new trend in live music making kind of like Sonic Pi. Instead of mixing tracks together, the DJ comes up with prompts on the spot and tweaks AI parameters to achieve the desired direction. Pretty awesome.
Could anyone help me understand whether using SVG instead of bitmap image would be possible? I realize that probably wouldn't be taking advantage of the current diffusion part of Stable-Diffusion, but my intuition is maybe it would be less noisy or offer a cleaner/more compressible path to parsing transitions in the latent space.
Great idea? Off base entirely? Would love some insight either way :D
Personally I like the results. I'm totally untrained and couldn't hear any of the issues many comments are pointing out.
I guess that all of lounge/elevator music and probably most ad jingles will be automated soon, if automation cost less than human authors.
Plenty of times this has been called before, and it's certainly possible that this wont be the last time this is called, but allow me to declare this the end of the bedroom producer (stage performers remain unaffected). And no two elevators will ever sound the same again, on any day.
FYI, other apps that use more classic but still complex Spectral/Granular algos :
https://www.thecargocult.nz/products/envy
https://transformizer.com/products/
Made this: https://soundcloud.com/obnmusic/ai-sampling-riffusion-waves-...
I made a video specifically in reply to you if you want to see exactly how I did it (3 mins): https://www.youtube.com/watch?v=69Q-cseNCI4
40 sec clip of uk garage/todd edwards style track made with riffusion -> serato studio with todd beats added
(1) Is there a corpse the keywords were collected from?
(2) Is it possible to model the proximity of the image to keywords and sets of keywords?
[1] https://www.riffusion.com/?&prompt=Post-avant+jazzcore&seed=...
[2] https://www.riffusion.com/?&prompt=Progressive+dreamfunk&see...
Edit: To be honest, I find something like 'Band In A Box' to be more impressive and actually useful, I don't understand how I would ever use this or listen to this. To me, it's further proof that Stable Diffusion really just doesn't work that well
https://www.youtube.com/watch?v=nUXgJkkgWCg
Not simply SOUND like a Bill Withers song, but to have the same depth of meaning and feeling.
At that point, even if we lose we win because we'll all be drowning in amazing music. Then we'll have a different class of problem to contend with.
Wondering if it would be possible to create a version of this that you can point at a person SoundsCloud and have it emulate their style / create more music in the style of the original artist. I have a couple albums worth of downtempo electronic music I would love to point something like this at and see what it comes up with.
prompt: relaxing jazz melody bass music negative_prompt: piano music
They even talk specifically about about applying stable diffusion and spectrograms.
Would you be willing to share details about the fine-tuning procedure, such as the initialization, learning rate schedule, batch size, etc.? I'd love to learn more.
Background: I've been playing around with generating image sequences from sliding windows of audio. The idea roughly works, but the model training gets stuck due to the difficulty of the task.
Me in the Stable Diffusion discord, 10/24/2022
The ppl saying this was a genius idea should go check out my other ideas
1) Write text to describe the problem 2) Generate an image Y that encodes that information 3) Parse that image Y to do X
Example: Y = blueprint, X = Constructing a building with that blueprint
and then things are going to get wild
I am so blessed to live in this era :)
Would be nice to have a channel where people can share Riffs they come up with.
I'm new but this is something that would get me going.
I'd want to train it to include foreshadowing, suspense, relatable characters and perhaps a twist ending that cleverely references the beginning.
The reason is, the models don’t understand content or even context, they just recognize patterns and can generate similar patterns which we then interpret.
Case in point, a photo of an Astronaut on a horse is not actually an astronaut on a horse, it’s a 2D pixel map of light that our eyes then interpret to mean an astronaut on a horse. It’s even easier to understand this listening to the generated singing. It sounds just like singing - but it’s not language at all and doesn’t mean anything, it’s just a very similar pattern.
These AI models are great at making things that look like a pattern that to us means something, but they don’t generate semantically meaningful content itself.
So when you do this with GPT3 or other models and long-form narrative it falls apart pretty fast as the model can’t keep straight things like characters and their internal personalities and motives, nor the overall plot arc.
But! You can definitely feed in prompts and ask questions to get ideas and boilerplate - descriptions of people and places come out especially well - and then you can edit that and use it to accelerate your writing process.
"If you have a GPU powerful enough to generate stable diffusion results in under five seconds, you can run the experience locally using our test flask server."
Curious what sort of GPU the author was using or what some of the min requirements might be?
It's possible a 3060 might work, depending. in my experience the 3060 is about 50% slower than the 3070, but that may be a bad 3060 in our test rig. but a 3060 gets pretty close to 5 seconds for an image, so try it, if you have one.
just tested prompt "a test pattern for television" on both cards and 3070 took 1.87s and the 3060 took 2.93s. Similar results for the prompt "an intricate cityscape, like new york"
edit: i should note we're using SD 1.4, not 1.5, although i think that just has to do with the checkpoint of the model, not the algorithm, but i could be wrong.
Also the model is over 14GB, so perhaps the 3070 can't do it after all. i'll test it later as soon as the local admin wakes up and downloads it onto our machine.
Tried getting something in an odd timing, but still is 4/4.
Who's going to be first to take this approach and use it to generate human speech instead?
This is already better than most techno. I can see DJs using this, typing away.
If I understand these AI algorithms at a high level, they're essentially finding patterns in things that already exist and replicate it w some variation quite well. But a good song is perfect/precise in each moment in time. Maybe we'll only be ever be able to get asymptotically closer but never _quite_ there to something as perfectly crafted a human could make? Maybe there will always be a frontier space only humans can explore?
There’s always that immortal randomly typing monkey with a typewriter thing [1]. And, in our case, it seems to be better than random.
So, yes, perhaps. But perhaps we could instead build and create things that are yet unimaginable upon it. We’ll see.
Consider sounds over 12khz. On a spectrogram during a chorus or drop that area is lit up, with so many things changing from millisecond to millisecond. A lot of AI samples really struggle at high frequencies, or even forgo them entirely.
Midi based approaches have been really great though, and an approach like in the OP is fascinating (and impressive).
Multi-instrument Music Synthesis with Spectrogram Diffusion:
* Don't do the alpha fade to the next prompt - just jump straight to alpha=1.0.
* Pause the playback if the server hasn't responded in time, rather than looping.
How, when everyone and their dog can generate such music? It's gonna be like stock photography in the age of SD.
Also, spectrographs will never generate plausible high quality audio. (I think)
So I think the next move is to map the generate audio back over to synthesizer and samples via midi …
Bracing myself for when major record labels enter the copyright brawl that diffusion technology is sparking.
It seems like you’d need very high resolution spectrograms to get a consistently snappy drum sound.
I'm guessing if one had access to one of those nvidia backplane rackmount devices one could generate 8k or larger resolution images.
It isn't perfect and it takes a lot of fiddling and sysadmin stuff, but it is way better than nextchar = CHR$(RAND64); color = (rand65535); of yore.
In summary: The search for AGI is dead. Intelligence was here and more general than we realized this whole time. Humans are not special as far as intelligence goes. Just look how often people predict that an AI cannot X or Y or Z. And then when an AI does one of those things they say, "well it cannot A or B or C".
What is next: This trend is going to accelerate as people realize that AI's power isn't in replacing human tasks with AI agents, but letting the AI operate in latent spaces and domains that we never even thought about trying.
I predict some kind of testing, validation or ranking be developed to filter out generated contents. Each domain has its own rules - you need to implement validation for code and math, fact checks for text, contrasting the results from multiple solutions for problem solving, and aesthetic scoring for art.
But validation is probably harder than learning to generate in the first place, probably a situation similar to closing the last percent in self driving.
As to why, music is fun to create, and this is just a tool.
You'd have to combine it with these guys https://www.youtube.com/watch?v=WqE9zIp0Muk
I did.
> The future seems depressing when I look at all this AI generated art.
You should talk about your concerns with an AI psychotherapist.
It doesn't make anything new or fresh. It doesn't pull any real-life emotions or experiences into a synthesis that a person can relate to. It's more like asking a teenaged comedian to imitate numerous impressions of music styles. e.g. in Clerks when the Russian guy does "metal": https://youtu.be/7gFoHkkCaRE?t=55
Of course the modern conception of music in the West is as an accompaniment to other, mostly drudging, activities, as opposed to something to be paid singularly attention. Therefore, there are many "valuable"(*) occasions to produce "impressions" of music. E.g. in advertisements and social media flexes where identity and attitude are the purpose of music. For these, a shallow interpretation or reflection of loosely amalgamated sound clips will suffice. But we don't just attend concerts or focus sustained energy on sonic impressions. We listen to lyrics and give over our consciousness to composed works because we want to find secrets others give away in dealing with this crazy thing called life- ideas to succeed, admissions of failure, and what the expected emotional arcs of these trajectories looks like. This lofty goal is to date not within the scope of AI stunts.
As Solzheinetysn said, "Too much art is like candy and not bread."
...now do stock charts
- No new result available, looping previous clip
- Uh oh! Servers are behind, scaling up
I hope Vercel people can give you some free credits to scale it up.
... nothing? like, there's no playback at all. Is that expected?
Perhaps AI can be trained to create music in different ways than generating spectrograms and converting them to audio?
It's all about empowering artists to explore more possibilities.
I can totally see how it can help prototyping or exploring ideas and visions.