I can roughtly understand the latter, but the former remains a mistery.
I can roughtly understand the latter, but the former remains a mistery.
The discrete cosine transform is awfully similar to the discrete fourier transform. They are conceptually the same and just have minor technical differences. The DCT and DFT's dictionaries are made up of sine and cosine wave atoms at different frequencies. These waves are all infinitely long, so they work best at representing signals which are repetitive over their whole length. They don't do a very good job at noticing where interesting parts of a signal are.
A wavelet is like a small chunk of sine/cosine wave. Instead of being infinitely long, it only exists over a small length of time/space. The wavelet dictionary contains wavelets in different places as well as at different frequencies. So when you decompose using wavelets they do a better job of encoding where the interesting bits are. Wavelets assume your signal might be periodic in particular areas, where the DFT tries to find periodicities over the whole signal.
so f = sum_n a_n*e_n
It's not at all obvious you can find any set of "e_n" so that this works, but it turns out you can. Not only that, but there are many ways to do it!
It turns out the set of functions has some very particular properties - without details, scaled sin & cos work (in complex variable, you get the FT), but so do step functions (Haar basis) and weirder functions (Wavelets), as well as many others if you relax some constraints (Frame theory, things don't have to be a basis any more. This is useful but has consequences like energy may not be conserved). Calculating "a_n" scalars takes a little of the details too, but is straightforward in practice.
The discretization of many of these things is quite useful in practice, but they all have continous variable equivalents. The DCT is basically a "trick" too, use twice as many samples as the DFT, but use only the real part (no complex numbers).
F(x) = ∫f(x)g(x)dx
for some arbitrary g, then the short time fourier transform is
F(x, t) = ∫f(x)w(x - t)g(x)dx.
So the mental model is your windowing the function you're taking the transform of. Remember that's just multiplication, so if w(x) is 0 outside of some interval, then so is f(x)w(x), and you use t to slide the window around.
But multiplication is associative, so you could easily think of it as windowing g(x) instead! Windowing the complex exponential gives you a wavelet. Rather than leaving it there, wavelet transforms add a scaling factor as well, giving
F(x, t, s) = ∫f(x)g((x - t)/s)dx
where g is now some 'mother wavelet'. Which could be the complex exponential, windowed or otherwise. If you remember that the complex exponential maps R onto the complex unit circle, with period 2pi, then increasing s increases the period (with respect to x), allowing for larger frequencies, and decreasing s gives you more detail in the frequencies you can see.
So, you'd want a wavelet that lowpasses the signal to avoid aliasing (which you can do by windowing!); and by the nyquist theorem in the discrete case you then need less samples to represent it fully. Taking that idea forward leads you to the whole filterbank thing you see all the time
I think the key thing is that wavelets aren't transforms of one variable, in the discrete case indices, but of three; index, time (location) and scale.