A hidden gem in sound symmetry
soundshader.github.io
soundshader.github.io
Music is a temporal ornament. There are many types of ornaments, e.g. the 17 types of wallpaper tesselations, but few of them look like music. However there is one particular type of ornament that resembles music a lot - I mean those “mandala” images. I don’t know how those are produced, but I noticed a connection between those images and music:
- The 1st obvious observation is that a mandala is drawn in polar coordinates and is 2PI periodic. Sound is periodic too, so I thought the two facts are related.
- The 2nd observation is that patterns on those images evolve over the radial axis. Ans so is music is a sequence of evolving sound patterns.
- The 3rd observation is that a 2PI periodic function trivially corresponds to a set of frequencies. We usually use FFT to extract the frequencies and another FFT to restore the 2PI periodic function. Thus, a single radial slice of a mandala could encode a set of frequencies. If this is correct, a mandala is effectively an old school vinyl disk.
Putting these observations together, we naturally arrive with the idea of using ACF. More details in the linked github project.
(Edit) also this is a really fascinating and enlightening way of viewing it - didn’t mean to imply otherwise with this quibble :)
(Edit) After reading a bit more, it seems to make more sense that it's a combination of both, as the resonance on the cochlea is likely not 100% accurate, and conversely the impulses from peaks would tend to be around the areas of resonance, so rather than being mutually exclusive it makes sense these two effects work in parallel.
Modern research suggests that the perception of pitch depends on both the places and patterns of neuron firings. Place theory may be dominant for higher frequencies.[4] However, it is also suggested that place theory may be dominant for low, resolved frequency harmonics, and that temporal theory may be dominant for high, unresolved frequency harmonics.[5]It is very exciting to come across others who are also interested in this topic. I am also very interested in the shape of sound but I have spent less time on empirical observations and more on imagining an abstract logic of numbers which can be visualized and heard. Real sound visualizations are also interesting to me but I decided to focus on abstract ideals because I thought it would be appropriate for a video game.
Hope you don't mind me sending some emails.
I take the FFT of the phase component, which is very similar to ACF; it’s the FFT of an FFT, but preserves phase. It even takes abs(), which might be mostly equivalent to your squaring operation.
Weird. I am really not trying to claim that I discovered ACF — quite the opposite. My result was shockingly different to what you found, even though the operations are so close to identical.
I think phase is extremely important in visualization. You can see why here: https://twitter.com/theshawwn/status/1176070853819342848?s=2...
You’ve come up with one of the most gorgeous visualizations I’ve ever seen for signals in general!
One way to incorporate phase: turn the angle into an x,y coordinate using atan2, then shade red and blue based on x and y. E.g. x of 1.0 is “full red”, x of -1.0 is no red; ditto for y, but with blue.
The other trick I used was to un-interleave the lines. Basically I noticed that every other line has a strong correlation; therefore drop every even numbered line to remove the aliasing artifacts. Then suddenly you get nice and smooth phase interpolations.
How did you compute FFT of the phase? The thing is, phase is discontinuous or multivalued function if we represent phase as a real number. We could also represent phase as a complex number of unit magnitude: exp(i phi). It would be continuous, but complex-valued.
And phase is indeed important for hearing:
https://auditoryneuroscience.com/vocalizations-speech/speech...
I didn't quite get the trick with uninterleaving the lines.
As you can see, the raw phase waveform is very "wavy", as might be expected. It oscillates rapidly, making it hard to see the patterns. But if you go to the tweets I linked above, you'll see the phase is much smoother in those images. How did I do it?
The key is to focus on every other line. Notice that if you simply pay attention to every odd row, it will be smooth.
I think I simply did "row 0, row 2, row 4, ... row n" followed by "row 1, row 3, row 5, ... row n + 1"
As for the fft of the fft trick for phase, I'm rsync'ing all of my old code and demo images to here:
https://battle.shawwn.com/sdb/voicecloning/
You may be interested in the png images, in particular the ones with "phase" in the names. You can probably ignore all the code except repl2.py.
Those images were generated via unknown methods -- sadly my repl sessions weren't saved. But, I happened to write down in repl2.py how the tweet images were generated:
cv2.imwrite(os.path.expanduser("~/Downloads/mel-phase-spectrogram-phase-fft-abs.png"), np.abs(np.fft.fft2(-1+2*1/255*cv2.imread(os.path.expanduser("~/Downloads/mel-phase-spectrogram-phase.png")))))
cv2.imwrite(os.path.expanduser("~/Downloads/mel-phase-spectrogram-phase-fft-abs2.png"), -1+2.0*np.abs(np.fft.fft2(-1+2*1/255*cv2.imread(os.path.expanduser("~/Downloads/mel-phase-spectrogram-phase.png")))))
So, input: https://battle.shawwn.com/sdb/voicecloning/mel-phase-spectro...Then, using the code above, the result: https://battle.shawwn.com/sdb/voicecloning/mel-phase-spectro...
I've verified that it still works. I think you can wget those images and copy-paste that code into a python repl.
So the only remaining question is, how was mel-phase-spectrogram-phase.png generated? Unfortunately that seems to be lost with the sands of time. But, as a hint, I think it was simply a matter of turning the phase component into x,y using atan2, then turning it into blue and red.
Also, completely unrelated, but I once made a super high resolution mel spectrogram that looked way cool and I can't resist showing it off: https://battle.shawwn.com/sdb/voicecloning/ultra-mel.png
I did all this when making 'Dr Kleiner sings "I Am the Very Model of a Modern Major General"' around a year ago.
https://www.youtube.com/watch?v=koU3L7WBz_s&ab_channel=Shawn...
Kinda funny that all of this visualization work was just to make memes, but the quest to meme turns out to be surprisingly motivating.
https://battle.shawwn.com/sdb/voicecloning/demo_output_101.w...
https://battle.shawwn.com/sdb/voicecloning/demo_output_75.wa...
Anyway, I think there's a lot left to discover in terms of audio visualization! I would definitely encourage you to play around with the phase component. The results can be pretty striking, as you can see from the "Result" image above (https://battle.shawwn.com/sdb/voicecloning/mel-phase-spectro...).
Sorry for the scattered explanation -- it's 4am here, but I wanted to give you some kind of writeup, even if it's rather disjointed. If you have more questions, be sure to ask! I can give better details tomorrow.
Chladni Plates? A few links from past research:
https://www.comsol.com/blogs/how-do-chladni-plates-make-it-p...
Video example - https://www.youtube.com/watch?v=dPTnGEEoFf4
https://www.youtube.com/watch?v=CR_XL192wXw
interactive version - https://www.dynamicmath.xyz/calculus/chladni-patterns/
(1) Making a stable and fast solver is very difficult. A simple solver for the canonical wave equation is fast, but unstable, so the solution has to be periodically adjusted to avoid NaNs. A stable solver, even for the simplest equation, would be 20x slower. For complex cases, we'd have to involve the Floquet theory, but that would bring an already slow solver to a halt.
(2) A wave diff equation can barely visualize a select frequency, not even a simple mix of frequencies or let alone music. The thing is such wave equations and their boundary shapes have a few select "resonance frequencies" that produce semi-stable patterns. Even a tiny step from a stable frequency, e.g. 6.1 Hz vs 6 Hz, and the solution turns into a mix of unstable patterns morphing one into another, which is cool, but not visually appealing. Mixing multiple frequencies together often produces an unstable mess and even if a pattern forms, you're never sure if it's the pattern for that frequency or just a transient shape, and if it's transient, you can't know if it's due to numeric errors or due to the nature of the equation.
(3) Limited resolution. The rule of thumb is that on a 1000x1000 px screen, the densest Chladni pattern would make 500 full wave repetitions, one pixel per positive and negative sides of the wave. This means we can render only the 500 different frequencies, with 500 Hz slowly turning into a mess due to rounding errors (2 pixels per wave period isn't really enough). Increasing the internal solver buffer to say 4096px brings fps down to 3-4.
However, despite all this, Chladni patterns are hiding something very remarkable, that seems to be glossed over in technical papers. If you imagine that a Chaldni pattern is a water or glass surface, with reflective and refractive properties, and look at the reflection of a simple symmetric object, e.g. a ring, you'd see something resembling a 3d hologram: all these inter-reflections will produce a "virtual image" or remarkable complexity. This picture can be taken by a hi-res camera, but visualizing it with GLSL is again very difficult: the raytracer needs to be outrageously precise.
Right now, the visual experience is like watching movement through a high-speed tunnel. I was expecting the "mandala" you mentioned in the sense that the end result is the accumulated visualization of all waves.
The sound representation would not disappear out of the borders. The first sound recorded would be stored as a narrow outside ring right next to the circle limit. The next sound would be stored as another narrow ring right before the first one. And it would continuously being accumulated sound after sound. The final result would be like a tree cut. It would have a final image representation the whole song, not just sequential snapshots of the sounds included in the song as it is now.
Anyway, congrats for the project! It is awesome and inspiring!
soundshader.github.io/?s=acf3&n=512&fps=1&acf.decay=0
It effectively computes the bispectrum as B(p, q) = F(p)F(q)F^(p+q) and runs the inverse 2D FFT to restore the triple autocorrelation. The results are interesting, but not impressive and very GPU intensive (NxNxlog(N) per frame is slow). In any case, I strongly believe that bispectrum is hiding something interesting and I just haven't figured how to see it.
Images are Fourier transformed, and the result is transformed to log-polar coordinates. This turns rotation and scaling in the source image into translations in the resulting log-polar data.
Anyway, fun stuff, thanks for the share!
[1]: https://sthoduka.github.io/imreg_fmt/ (follow link to the pipeline description)
Your comment reminds me of the search for the mandelbulb fractal. It seems a fitting comparison, a bulb being a sort of ornament.
Anyway, interesting work.
> AudioContext.createMediaStreamSource: Connecting AudioNodes from AudioContexts with different sample-rate is currently not supported.
Edge worked though. Didn't try Chrome.
Edit: This is beautiful!
I have checked out a variety of songs, and I feel the visualization is rather dominated by whatever frequency is loudest. (E.g. bass sound -> three- to five-fold symmetry and most detail obscured by it).
Have you considered applying something like the https://en.wikipedia.org/wiki/Equal-loudness_contour somewhere in the process to more evenly weight the frequency contributions according to human hearing perception? Not sure if would have the intended effect, but I'd be curious what happens.
I've in fact tried implementing the equal loudness contour - try adding ?acf.aweight=1 to the URL. However the result is mediocre. I've also tried applying a few bandpass filters for low, mid and high frequency ranges, rendering them separately with different colors and then mixing images together. The result is, again, medicore. I've been entertaining the idea that ACF waves are ought to be rendered like ocean waves: via light reflections.
Very nice job, congratulations!
It interprets the trumpet in Miles Davis - It Never Entered my Mind as dark ripples and is beautiful in it's own way.
It's a long way from the last time I used a visualizer on Winamp, nice!
Edit. The generated images in fact contain the small ripples, but our eyes don't notice the 0.1% modulation of color. If they did, we'd see bright orange-blue waves with a fine pattern of wavelets colored with a slightly different shade of orange and blue. People who can see 100 shades of orange would see this pattern.
Still, great experiment, and interesting results!
Can’t try it on my dated iPhone - you need to vendor prefix the AudioContext with webkitAudioContext if AudioContext is undefined.
That's really it. The (4, 2, 1) is the oversaturated orange color, i.e. it would progress as (1, 0.5, 0.25) -> (1, 1, 0.5) -> (1, 1, 1). None of the tricky HSL/HSV schemes worked better than this.
I really enjoyed your work: the write-up was clear and the demo page worked well.
(also note that your site does not currently work in Firefox, which would be nice to fix)
The demo works on Firefox 78, Ubuntu. However you'd have to set correct sample rate with ?sr=44.1 to match the mp3's sample rate: Firefox won't do resampling.
What you can do though is look at the first few bytes of an .mp3 file (since it's a file drop/file load) to just directly read the sample rate from the MP3 block header[1], where you directly check the value encoded by [data[19], data[20]]: if it's [0,0] that means it's 44100, [0,1] means it's 48000, [1,0] means it's 32000 and that's it. There are no other sample rates allowed for MP3.
[1] http://mpgedit.org/mpgedit/mpeg_format/MP3Format.html for the full block format)