The Mathematical Genius of Auto-Tune
priceonomics.com
priceonomics.com
I think this is the article's way of dramatizing the standard way of calculating the autocorrelation using the convolution theorem: https://en.wikipedia.org/wiki/Autocorrelation#Efficient_comp...
Assuming that it's the trick that the author was referring to, the article is actually a bit sparse on details.
My (very non-expert) reading of the patent is that it uses sample reduction and some specific features of the periodic form of the human voice to simplify the math involved in the auto-correlation routine. Although this routine does seem to be unique as far as I know, I wouldn't be surprised if it shares similarities to other techniques to reduce correlation complexity.
Doesn't sound like FFT.
Just cos you've got enough grunt to do it the dumb way doesn't mean the smarts that produced this are somehow unimpressive.
http://theweek.com/articles/467109/sad-songs-made-happy-amaz...
edit/tidbit: he told me that the interview process for hiring new developers goes like this: all the questions are straight out of K&R's C Programming Language, you have to get all of them correct, and so far only one programmer has :P
He's obviously referring to himself.
Seriously though, I’m tired of folks doing stupid counterproductive interviews and then parading them as a good thing.
That is dramatic hand-waving, but it does help convey to non-programmers how dramatic an improvement the right algorithm can implement.
It's true what the article says that pitch tracking was considered a difficult "holy grail" in 1995. I attended a talk about pitch tracking at Interval in 1996 by somebody whose name I don't remember any more. The speaker made the point that pitch is a perceptual -- not a mathematical -- concept, and it was hard but possible to do it in real time on a typical PC at the time (i.e. 90 MHZ IBM ThinkPad 760C). But he had done it, and everyone seemed impressed by his demo! ;)
The article mentioned formant analysis. Maybe that's related to cepstral analysis, which is another way of tracking the pitch and formants of voice, and has its own cool nomenclature of secret code words. "Cepstral liftering" is basically two FFT's followed by an inverse FFT.
If you take the complex FFT of a voice signal, the formants show up as two or three large "hills" in the spectrum, but the pitch manifests as higher frequency repeating furrows in the hills. (Which you want to filter out so you can analyze just the formants, or synthesize a different pitch into them (i.e. auto-tune), so you need to know the frequency of the change in the spectrum magnitude over time: another FFT!)
http://www.phon.ucl.ac.uk/courses/spsci/matlab/lect10_files/...
So you take a second FFT of the log of the first FFT, to get a "reverse spectrum" or "cepstrum" in the "quefrency" domain. The fundamental pitch shows up as one big spike in the cepstrum (with smaller spikes for its harmonics). Just "lifter" out the pitch spikes, then perform an inverse FFT to get back to the smooth low frequency formant hills with the high frequency pitch furrows removed.
I'm sure there's a lot of "special sauce" in getting the math tweaked and tuned right so it actually sounds good and runs fast.
https://stackoverflow.com/questions/4583950/cepstral-analysi...
The patent seems to be about autocorrelation, which is something different than cepstral analysis, but maybe they could be used together to get even better results.
I don't know but would love to learn what the trade-offs and limitations the two techniques have.
https://en.wikipedia.org/wiki/Cepstrum
>The name "cepstrum" was derived by reversing the first four letters of "spectrum". Operations on cepstra are labelled quefrency analysis (aka quefrency alanysis), liftering, or cepstral analysis.
https://surveillance7.sciencesconf.org/conference/surveillan...
"The original application was to the detection of echoes in seismic signals, where it was shown to be greatly superior to the autocorrelation function, because it was insensitive to the colour of the signal."
Also related to this, at least I think, is the issue of turning recording into midi or sheet music. Now that would be a killer app... There is some good software out there, such as Melodyne, but it requires a lot of manual work and tweaking.
The lead singer's voice just sounds so thin, and unnaturally on-key (no vibrato).
I think you'd be hard-pressed to find verifiable examples because those higher settings are probably mostly used to cover up poor singing, and so nobody involved is going to volunteer that information.
FWIW, I don't think modern auto-tune plugins have a problem pitch correcting singers using vibrato.
I think auto-tuners get a bad rap. As an amateur musician, subtle corrections in post-production can save me days of recording to get the right take.
Not a pro editor, but I know my shit.
Outside of the a cappella domain, I believe it's pretty prominent in most pop music, but my understanding is that the technology's gotten to the point where 95% of the time you can't tell apart machine correction from good tuning of the vocalist.
Of course a lot of folks cannot tell the difference, but to me it's night and day. I couldn't stand Glee for instance because of how auto-tuned the voices were.
This was pretty obnoxious when I noticed it; I don't mind if they use auto-tune for a one-off musical episode of a show, but for a show where every episode is a musical to use it (and not very subtly) was very off putting.
I hate it because it's throwing off kids' sense of what a natural human singing voice sounds like. It's photoshop for the voice and is similarly damaging to one's image of self and of others.
Songify / Auto-Tune The News / Schmoyoho: https://www.youtube.com/user/schmoyoho
No Fair - Trump ft. Eric & Don Jr. | Songify This: https://www.youtube.com/watch?v=8NkNcwzGCfQ
Obama Mic Drop: 1999: https://www.youtube.com/watch?v=NbqtAuT4zbk
I'll try to walk you through what I'm actually hearing that makes me think there's Auto-Tune there:
- 2:50: Subtle, but as the first "oh" starts, it sounds like it quickly jumps from a slightly lower note to the correct one.
- 2:58: The word "part" has a classic Auto-Tune sound. The real recording probably went slightly off pitch as the note was held and Auto-Tune has made it perfect and a little robotic.
- 3:07: "Where" has a similar sort of robotic pitch slide sound as it starts as the "oh" did earlier.
- 3:13: "Discover" has a sort of a glitch in the middle as Auto-Tune tries to track across the 'k' sound in the middle of the word.
As someone else already said, songs in TV show Glee used it all the time in a fairly obvious (but still subtle compared to intentionally sounding Auto-Tuned) way.
[0] https://www.youtube.com/watch?v=zmhTmKD6rXc
Same man, same song, in 1978, 1986, and 2000:
[1] https://www.youtube.com/watch?v=fcLV6gE6THg [2] https://www.youtube.com/watch?v=Hwwu5vYSShY [3] https://www.youtube.com/watch?v=lQkXdRKMB9c
And now in 2015:
[3] https://www.youtube.com/watch?v=ovGwkScJRZ4
Now, while it's possible that he simply became a better singer between 2007 and 2015, subtler auto-tuning of the type conventionally used in pop production generally just makes the voice sound cleaner, thicker, more polished. Note that his style has also changed to favor staccato phrasing so as to limit stairstepping artifacts.
I think here you may be confusing Auto-tune with "modern" voice processing, which while happened together are quite different. Modern pop singers are overdubbed ridiculously with added distorsion, resulting in a "massive" (and slightly "robotic") lead voice.
Edit: You can clearly hear this "saturation" in the song Hello by Adele (https://www.youtube.com/watch?v=YQHsXMglC9A), especially in the chorus. Compare with her natural voice: https://www.youtube.com/watch?v=-yL7VP4-kP4
*Sax was the only non-fretted/pitched instrument in that recording.
I'd imagine its popularity is more due to convenience rather than demand.
https://en.wikipedia.org/wiki/Barbershop_music#Ringing_chord...
I've been working on a relatively simple real time music/audio processing project on an Arduino (identifying tempo and using it to create interesting lighting effects for a Halloween costume) and it's an interesting challenge. Extracting any kind of useful information about the underlying musical structure from polyphonic audio is an incredibly hard problem. Add to that limited hardware and the kind of sampling rate you need to capture music (upwards of 40kHz if you want to capture everything you can hear) and you have to get creative.
http://www.antarestech.com/products/detail.php?product=Auto-...
Also as a $349 rack device : https://www.amazon.com/Tascam-Producer-Processor-Antares-Aut...