Wavelets: A mathematical microscope [video]
youtube.com
youtube.com
I've used wavelet transforms on and off for years for things like image compression, harmonic analysis, and currently brain research, and I've always been somewhat disappointed in the quality of material that's out there explaining how wavelets work and how to use them effectively.
There's a small amount of "abstract math" heavy material but relatively little that's focused on using it for engineering, in contrast to all of the amazing material on Fourier transforms.
Random shot in the dark. What do you think of the Free Energy Principle?
My view of neuroscience is that we're decently good at figuring stuff out at the very very low level (e.g. behavior of complex molecules, classification of cell types and basic low level functions), and at the very high level (e.g. "this region is responsible for speech"), but haven't really made the right breakthroughs for figuring out the mid level (e.g. "how exactly does this ball of 50 million neurons perform its function, architecturally"). This theory seems to lie in that middle area.
Agreed completely!
That's why I submitted it to HN! <g>
The better approach with better resolution is to use native time-frequency technique for example non-linear Cohen class time-frequency but currently it requires massive computing resources even for a short duration [2]. Hopefully with the available of cheaper RAMs in the order of Terabytes it will be more feasible to use proper time-frequency techniques for big data signal processing.
[1]Wavelet:
https://en.m.wikipedia.org/wiki/Wavelet
[2]Bilinear time–frequency distribution:
https://en.m.wikipedia.org/wiki/Bilinear_time%E2%80%93freque...
Its not quiet done yet and misses some polishing but I find it already helpful rightnow.
I'm not actively involved in the package development anymore but it is still maintained [3] and there is a great chance that you have it already in your Python environment as a dependency of scikit-image/scikit-learn. Just give it a try, it's very simple:
>>> import pywt
>>> cA, cD = pywt.dwt([1, 2, 3, 4], 'db1')
[1] https://en.wikipedia.org/wiki/JPEG_2000
[2] https://pywavelets.readthedocs.io/
[3] https://github.com/PyWavelets/pywtThat only works in the finite case though.
You have a database of k reference signals you know about vrefₖ ∈ Rⁿ.
What you want is to match the signal you observe to one of the reference signal.
This is called template matching : you pick the closest reference signal to the observation :
You compute argminₖ distance( v, vrefₖ)
This is great but it scales badly with respect to the number of reference signals.
So instead of trying to find the closest signal, you try to decompose the signal as a sum of reference signals :
You compute argmin over λₖ distance( v, Σₖ λₖvrefₖ)
This decomposition is not unique, so you regularize it (if your distance and norm are well picked the following is convex so you can get uniqueness guarantees):
You compute argmin over λₖ ( distance( v, Σₖ λₖvrefₖ) + Σₖnorm(λₖ) )
This is called sparse coding.
The rest (fft, stft,wavelet transform,...) are mathematical tricks (called transforms) to compute this more efficiently by choosing certain special reference signals (picking a basis).
The first math trick to be aware is (a-b)²= a²+b²-2ab : This is the bridge that relate what is called correlation to the distance :
for vector a,b ∈ Rⁿ : Σₖ(aₖ-bₖ)² = Σₖaₖ²+ Σₖbₖ²- 2* Σₖ(aₖbₖ)
With judicious pick the first two sums Σₖaₖ² and Σₖbₖ² can be made irrelevant to the minimization problem because they are either constant by construction (normalized to 1) or don't depend on the variable that we minimized over. This trick applies both to template matching and sparse coding (even though there is more terms they all simplify).
If the two first terms can be made to disappear then minimizing the distance has the same result as maximizing the correlation.
Then you have a bunch of sliding sums and recursive summation tricks which all depend on the specific shape of your data. Notably with fft there is the well known butterfly algorithm.
A century of techniques and theory since Hilbert have been devoted to finding various smart variants. And they allow you to compute this decomposition fast usually in O(n) or O(nlog(n)).
The alternative to hand-crafted techniques, is learning the reference signals from the data (Sparse dictionary learning). You can also brute-force your way through with deep-learned filter banks. It works better but require more compute, and have less guarantees. It also let you deal with nitty gritty details like missing data and masks.
The physical significance really arises from a place called the Huygens principle in mechanics, which supports claims about spatial wavefront propagation to state that at each timestep and point on the wavefront, another similar wave propagates, these subwaves are called wavelets.
If anything, the uncertanity principle is due to the time-frequency resolution trade off.