Shazam: not magic after all
blog.revolution-computing.com
blog.revolution-computing.com
http://news.slashdot.org/comments.pl?sid=7310&cid=823710
(from 2000).
I remember this because I got in a big argument with someone about whether this could possibly work. Of course, I never got off my ass and implemented it, which I guess makes me a huge loser.
"Unfortunately, there's no indication in the paper of what software was used to develop the process (although the scatterplots in the paper do look decidedly R-like)."
I used to work with Avery Wang, the guy who devised the algorithm. He used Matlab.
By the way, the psycho-acoustically spectral measurements referred to in the article are called MFCCs[4] - basically an FFT reading weighted according to the sensitivity of our ears. They are often used in both music and (especially) speech recognition because they tend to accurately sum up the timbre we perceive in a given sound. Timbre is much easier to extract from a digital audio file than pitch or vocal information, hence why it tends to be successful in applications such as this.
Shazam is still pretty cool too
[1] http://jmir.sourceforge.net/
[2] http://libxtract.sourceforge.net/
Then, you look up those peaks in a book, which has compounds ordered by the wavelength of the highest peak.
It takes minutes to do it by hand, I'm not surprised computers can do it better.
5-6 hrs of a good developer == 1 month of 3 ops in India or elsewhere. Except that getting ops to work itself can be painstaking.
Airtel, a leading telecom provider in India had a SongCatcher service long back (3 yrs ago) http://www.techtree.com/India/News/Catch_a_Catchy_Song_with_... I never tried it - may this one worked for a predefined set of songs.
"Specifically, a fixed length of audio is converted to audio DNA; this conversion process extracts certain features from the signal based on the psycho acoustic considerations. The system has two components, one that enables the extraction of Audio DNA from a few seconds of recording, and the other is an efficient search engine that finds the exact match for the DNA.
The audio DNA is based on extracting 64 sub DNAs every 3 seconds. The sub DNAs are generated by looking at the energy differences along the frequency and time axes. These 64 sub DNA form the chromosomes of the system, which enables the system to uniquely identify the chosen song."
Looks to me like pretty much the same technology. This is again not a surprise. Most implementations of this idea will be using similar techniques. What I am amazed at is that somebody thought that all this was feasible.
Anyway, I think the idea of query by humming is not a dead end. However, such a hypothetical service should somehow collect and use a database of different "hums".
#DDUUUUDDDUUSSUDDDDUUSDDD
Many, many tunes can be separated with the first 20 symbols.Anyone have a reference? I'd like to acquire a copy ...
here is an overview:
http://cosmo.nyu.edu/hogg/research/2006/09/28/astrometry_goo...
It's fun to read about but way over my head mathematically.
I'm tired of these software 'engineering' types who insist that computers are run by using 'maths' and 'numbers' (whatever those are).
Clearly, computers are run by aphasic tonally-separated spinning disks. These disks fire puffs of air out the sides of the computer, creating little tiny tornados, which summon air spirits to call the fire spirits, which causes the screen to light up and the keys to make tappy-tap noises.
anyway. to be clear: not statistics. not math. not regulated pulses of electrons. MAGIC!
While less technical people don't understand, they have 'faith' that there's a logical, scientific explanation for how computer's work.