So the glasses are always present when the bass/808s are hitting, so is there something that maps the sound to the images?
What is it about the algorithms that make the images 'dance' so quickly between the 3.5 beat and the 1? Is it because there are static risers that move so quickly through the wave spectrum?
Wait... is light skin mapped to when highs dominate and dark skin to lows?
I'd be delighted to learn how you translate musical structure to a point in latent space!
In Phantom Part II they mostly have their mouths closed. In La La Land it varies but the mouths are mostly open. If you focus on the mouth you'll get little mental radar blips where the mouth could be tracking what is being said.
"Pretty girl and you let go" - https://youtu.be/52qWiLoOeIQ?t=18
"If you wanna waste time baby" - https://youtu.be/52qWiLoOeIQ?t=44
"Yeah i met her at a one oak" - https://youtu.be/52qWiLoOeIQ?t=84 (esp the one oak part)
Anyone that actually watches these will probably just think I'm low on sleep, but it's kind of interesting.
(Try this with different kinds of songs if you can!)