Visualizing Music with GANs
twitter.com
twitter.com
So the glasses are always present when the bass/808s are hitting, so is there something that maps the sound to the images?
What is it about the algorithms that make the images 'dance' so quickly between the 3.5 beat and the 1? Is it because there are static risers that move so quickly through the wave spectrum?
Wait... is light skin mapped to when highs dominate and dark skin to lows?
In Phantom Part II they mostly have their mouths closed. In La La Land it varies but the mouths are mostly open. If you focus on the mouth you'll get little mental radar blips where the mouth could be tracking what is being said.
"Pretty girl and you let go" - https://youtu.be/52qWiLoOeIQ?t=18
"If you wanna waste time baby" - https://youtu.be/52qWiLoOeIQ?t=44
"Yeah i met her at a one oak" - https://youtu.be/52qWiLoOeIQ?t=84 (esp the one oak part)
Anyone that actually watches these will probably just think I'm low on sleep, but it's kind of interesting.
I'd be delighted to learn how you translate musical structure to a point in latent space!
(Try this with different kinds of songs if you can!)
Imagine using a Vive to explore a GAN interactively. You'd be able to control the GAN using vive controllers and by walking around your room.
Right now it takes 163ms to render a 1024x1024 frame on a K80 GPU. That's 6 FPS, which is within an order of magnitude of 60FPS.
I haven't timed a 256x256 GAN, but presumably it would be 16x faster to generate. If so, then you'd be able to achieve 98FPS.
The above timings are based on the 1024x1024 FFHQ GAN model, which generates portraits of humans. https://github.com/pbaylies/stylegan-encoder
And indeed, it looks like the author uploaded an FFHQ music video 14 minutes ago! https://www.youtube.com/watch?v=3TLEfOMBbMw It looks cool.
Someone should train a 256x256 FFHQ and make a 90FPS interactive renderer for it.
Unfortunately it's not possible to take a large GAN like 1024x1024 FFHQ and only generate a 256x256 image. Each GAN is trained for a specific size, so you're stuck with 6 FPS at 1024x1024. I wish the FFHQ authors had saved a 256x256 checkpoint during training.
Training a 256x256 GAN from scratch costs somewhere in the range of $150 GCE credits. But you might be able to bootstrap a 256x256 FFHQ using the weights from the 1024x1024 FFHQ (aka transfer learning). That might train a lot faster.
There is also the recent NoGAN technique, which skips progressive growing by pretraining the generator: https://github.com/jantic/DeOldify/#what-is-nogan Supposedly it speeds up GAN training by a huge amount.
[1] https://twitter.com/goodfellow_ian/status/108497359623614464...
I've implemented music visualizer ages ago using similar concepts (pure algos though, no real images). It happened when nVidia released the first affordable consumer video card with decent shader support. I think it was 6600GT . My animation part that made video dance to music was a bit more sophisticated though.
However, in terms of graphics, this strikes me as different from anything that was possible before the recent advances in GANs. During the era you're talking about, the art of shader-based music visualizers was being pushed by projects like Milkdrop 2, and nowadays a lot of similar research still happens on Shadertoy, and the demoscene, of course, hasn't stopped blowing people's minds.
But this is on another level entirely. It's as is the content and seemingly human concepts themselves are being smoothly animated.
- well this is because you did not see my vis. It looked just like the one you saw on GAN's related link with similar transitions. Except that all imagery was generated by math formulas running in pixel shaders instead of ready bitmaps/videos.
Here is the actual screenshot: https://exsotron.com/exs_files/exvis-0003.jpg
Actually I played with the actual music video clips as a source of the imagery and the results were really cool but obviously other then experiment at home could not really do this part due to copyrights etc.
@jonathanfly has been posting some interesting machine learning effects: https://twitter.com/jonathanfly/status/1185843103271444480
I guess one way to make these effects dance to music would be to make a Mel spectrogram of the audio, then somehow use the shapes in the spectrogram to apply deltas to the rendered frames.
1) Each of my pixel shaders was driven by let's say 32 parameters (do not remember the exact value)
2) The code would generate first set of said parameters and the second set (values were random) and start transitioning ( lerp ) between 2 with the length of transition of about 60 seconds.
3) Upon completion of the transition the first set would be replaced by second set and the second set would be replaced by freshly generated third set and an infinum.
4) Seps 2 and 3 allowed for non stop fluid motion.
5) Lerp value the degree of transition between sets for each parameter would be modulated by sound (one FFT band for each also passed through synth like attack / decay.
6) Finally there was beat detection part which upon detecting a beat would invert lerp direction
There were more steps and various tricks to make it more interesting and non repetitive but I am not writing article here ;)
The end result was quite artistic. The visualizer was part of much bigger enterprise grade media playback / management / delivery /scheduling platform I've developed for hospitality industry
[1] https://twitter.com/xsteenbrugge/status/1188798045158293505