Show HN: Sound visualisation, better than FFT (iOS)
vsound.app
vsound.app
While the post is still somewhere around the front pages, I wanted to ask what do you think about the following potential applications for monetisation:
1) Desktop app with various visualisations suitable for displaying on screens at bars / clubs. Charge for subscription for extra (good) visualisations (made with proper graphics, not what I did).
1.a) Same but for making videos for YouTube etc.
2) Phone/Tablet app that can consume sound from other apps. Charge for extra 2D/3D dynamic visualisations (without user being able to move around), one time fee per extra visualisation.
3) 3D environments on Phone/Tablet, where things "vibrate to music", like leaves on the trees, waterfalls, etc. Charge for extra 3D environments. Very hard to make, but possible.
4) Apply machine learning to extract components of the tracks, and allow mixing different components from different songs using a new kind of user interface. Extremely hard, but potentially could lead to a DJ revolution, if successful.
5) Any other ideas?
Thank you!
The visualizations could be more details to show various frequencies or zoom in on specific areas etc..
I would love to chat with you about an audio app that I'm working on. Thanks!
I also like looking at spectrograms of birdsong. I had a setup where I would see the spectrogram of sound coming through my microphone in realtime. I would place it on the windowsill and a few times I would see very faint birdsong I hadn't noticed with my ears, and then I would "tune in" my ears to that frequency range and I would hear it.
Thinking about it the ideal form factor for a realtime spectrogram for the hard of hearing (or the curious) would be something like Google Glass. (Originally thought of a smartwatch but that doesn't make much sense for an additional sense).
There's a project that's the opposite of this where a colorblind man attached a sensor to his head that converts color into sine waves he can hear. Adam Montandon. Eventually this electronic sense became a natural part of his experience and he began to dream in "color".
Funny think about birds! Yesterday, as I was preparing to release the app, birds were singing outside of my window (coronavirus lockdown, so quiet), and I noticed how fun to look at their songs on my app! I've spent literally 30 minutes looking at that :) I've posted a link to birdwatchers reddit :)
Normally this would affect his ability to hear certain frequencies, but he’s so accustomed to working with raw waveforms, he’s not too concerned. He knows exactly what the sound he wants should look like.
Am I looking bat the wrong place? Can someone explain the math used here?
No need to publish a scientific report.
You mention that your wavelet transform is "real-time", but the normal wavelet transform already is real time.. or what am I missing ?
It’s interesting as I just worked on FFT and sound visualization with Python, Seems quite troublesome to circumvent the need for a sample buffer or time-windows.
If you see a way to monetize I would definitely do so. Only you can decide how unique your method is. If it is very unique, patent it? Perhaps Shazam or other apps could use your method?
Being able write a lib to plug into TouchDesigner or OF would make it fun to write other kinds of visualisations.
http://hplgit.github.io/wavebc/doc/pub/._wavebc_cyborg001.ht...
We only need to apply it to a circle and rewrite the solver in C. This will give us a visualizer for a single frequency.
Then we run FFT, run the solver for all frequencies, observe that the final solution is a linear combination of the individual solutions and apply the rainbow coloring to corresponding frequencies.
This solver needs to be kinda fast to run at 15 fps, but luckily, different frequencies can be solved in parallel. Most likely, changing the input frequency a bit will change the solution only a little, and so we could pre-solve the frequency range with sufficient density, cache them and rapidly derive actual solutions by interpolation.
Bonus points for using complex numbers. The boundary condition in a singing bowl is u(x0, t) = A sin(Bt) where x0 denotes the circular boundary. Since real numbers are boring, we could expand the problem into the complex plane: u(x0, t) = A exp(iBt). In this case the solution u(x, t) would be in complex plane also, where the absolute value |u| is the amplitude, or pixel opacity on our visualization, and the angular coordinate arg u would be maybe color of that pixel?
The real physical solver doesn't do FFT, though. It makes the boundary circle vibrate with the input sound wave and effectively solves the wave equation where the boundary condition is u(0,t)=f(t) - the input sound.
I've run some calculations that solving the wave equation in real time would be infeasible. The convergence depends on the Courant number, which basically says that the grid step dx must be c*dt, i.e. the sound must travel on grid step per one time step. Since sound travels at 1.5 km/s in water, the grid needs to be super dense and 1 second of sound would need around 1 petaflops of calculations.
(I spent like 15 minutes playing a 'simulated' church organ and watching the frequency pattern with the dots. It appears that each line is an octave with the leftmost dot in each row tuned to a G. Super fun!)
Edit: I'm thinking how cool would it be to add some kind of visualization like you see here: https://paveldogreat.github.io/WebGL-Fluid-Simulation/
Is the chromatic scale you align your visualizations to just or equal tempered?
Some of the world's finest audio devs hang out there - Urs Heckmann/u-he, Andy Simper/Cytomic, Alexsey Vaneev/Voxengo, Dave Gamble/DMG - and so do a bunch of amateurs and rand0s. Good discussions about both classic/novel audio DSP algorithms, and the challenges of monetization.
This may help answer another question: What type of visual would allow feeling the music without hearing?
I think this intermediate representation that I've extracted using this algo would be much better for machine learning on the sound data than either 1) raw sound or 2) frequencies extracted with FFT. But there is an engineering difficulty to overcome: the frequency data that the algo extracts doubles in its amount (more frequent frequency samples) with each octave... It's like wavelet data... It's a challenge to feed this data to standard ML algorithms, need to think of how to configure the inputs, it would have to be highly hierarchical.
My initial goal was to "learn" different instruments from raw sound data, and this intermediate representation is good, because it allows "translation invariance" across frequencies. I've described it here: https://vsound.app/high-precision-in-frequency-domain.html
It's a work in progress... This app is just for me to see if people are generally interested in this area, I don't want to do something for months only to discover nobody wants it (although it's been a few months I've worked on this app LOL).
I might go further by letting people hear the difference between them (by transposing the visuals back into audio, to ground our sense of quality loss in the original medium, sound).
In both cases though it's possible to losslessly round-trip the audio within numerical precision (neither the STFT nor the Wavelet transform lose information).
If you are using a normal FFT this is no problem. You can reconstruct the original signal with a very small amount of error. However, this works because the FFT preserves phase data for each of the buckets. The spectrogram does not preserve phase data, so it will really mangle things.
(I’ve tried this, but it’s been a while. You get a sound which is recognizeable but total garbage otherwise.)
One way to get high frequency resolution from the STFT or Wavelet transform is to use the phase derivative within each frequency band, which is usually called "instantaneous frequency", and is closely-related to the phase-vocoder (different from the daft-punk-famous vocoder).
The main issue with this kind of instantaneous phase estimate is that it assumes there's only one dominant frequency within each band - does your method improve on that?
About the question of "assuming there is only one frequency within each band" -- yes, it's a problem, and that's what's causing the "frequency leakage" that you see in the app with the "high-precision: off" (and FFT). To solve this, I've added some extra math on top, which is available via "high-precision: on" in the app. I will release more details when I open source the algo. Please subscribe to email list if you want to receive an update, I am not super active on social media otherwise :).
I would say that the spectrogram is a visualization of the STFT magnitudes, in the same way that a line plot is a visualization of the FFT magnitudes. The STFT and FFT are both invertible though.
edit: wording
Cool project, good luck!
Just playing with it now, but some quick feedback:
- Disable the lock screen while watching the visualization. I have to keep tapping it while music is playing.
- Is it possible to select the input source as the "speaker out" (or whatever ios calls it) rather than the microphone so only the music playing is analyzed rather than the microphone. My clicky keyboard is messing up the music visualization.
Nice work!
Why am I using IOS 9, you ask? Because that is the latest my device supports. Why not upgrade you ask? Because my device works fine and is in great shape, so I'd rather not destroy the planet for for that sake of my own greed, vanity and stupidity.
a) you are running on a platform that it is impossible to load a previous OS onto an actual device, so if the developer bought their device after iOS 9 (quite likely), that actually can't put iOS 9 onto their device for testing.
b) would you excited about testing 4 years' worth of operating systems over at least 3 devices (iPhone 7 size, iPhone big size, iPhablet size), many of which you can only do on the simulator ...
c) ... fixing bugs on obsolete versions of the OS that few users will even encounter because they cannot run those versions
c) ... for free
d) ... just to accommodate 1.5% of users who insist on living four years in the past?
If you are the kind of person that likes to do all that for free, though, perhaps you could offer your assistance to the author?
Btw, Apple will recycle your old phone for free, if the planet is the only thing holding you back.
[0] https://gs.statcounter.com/os-version-market-share/ios/mobil...
If it doesn't work for your setup then just move on instead of telling people to spend more of their time to make it work for your edge case.