Utau – a Japanese singing synthesizer application
en.wikipedia.org
en.wikipedia.org
Here's the most impressive Vocaloid song/video I've seen:
http://www.youtube.com/watch?v=zkLJoFp2UAE
EDIT: From the artist's bio:
"Mitchie M is a relatively new producer who became popular due to his very realistic tuning of Hatsune Miku. His name is derived from The Jimi Hendrix Experience's drummer, Mitch Mitchell, and originally was "Mitchiell Mitchie" before shortening it. He started composing songs in high school in a band which made covers. However, he wanted to write his own songs. His biggest influence is Tom Jobim. He uses Logic Pro 9 to compose his songs. Mitchie M's popularity is steady with most of his videos quickly gaining at least 200,000+ views."
This hints at part of the magic of Vocaloid: It gives the inexperienced, unknown, or unconnected a pitch-perfect (if a bit robotic-sounding) vocalist to experiment with, significantly expanding the range of genres a lone musician can cover. The take a relative "outsider" can have on pop music is very interesting, something of a middle ground between "mass-produced"/"commercialized" pop and the more experimental, "underground" electronic music you usually hear from solo acts.
Vocaloid English: http://www.youtube.com/watch?v=S_YNQzcRhmE
Utau English: http://www.youtube.com/watch?v=Cm7o5ptHwb0
For a head-to-head comparison, we'll need to switch languages. Listen to Hachi's song Matroshka rendered by:
Vocaloid Japanese: http://www.youtube.com/watch?v=_JGaQ3g8WU4
Utau Japanese: http://www.youtube.com/watch?v=nCqy27QrqZA
Here's something somebody spent way too much time working on.
https://www.youtube.com/watch?v=t8GKzBhAgQo
and another
UTAU doesn't do anything algorithmic to help with these word-splitting tasks; even if it could, there's a lot of nuance and judgement that depends on everything from the tempo to the specific vocalist's pronunciation. For instance, the word "slam" is actually best rendered as "sle, e m" when working with Kumi because the "a" sound is more hollow, and the point at which you switch from the "sle" to the "e m" depends a lot on context. A 2-minute song with chorus and verses can take hours to get perfect.
And of course, pronunciation is only the beginning. Mitchie M (one of the most highly regarded vocaloid artists, mentioned elsewhere in this thread) stated in an interview that he can spend up to 2 weeks programming intonations for rap segments in his songs. [4] It's very interesting to think about how spoken word would be coded as pitches, and how we are able to detect deviations from common pitch-patterns as "unnatural" when listening to spoken passages.
[1] See http://en.wikipedia.org/wiki/Katakana
[2] Why unreleased? Still working on making lyric videos to accompany the recordings; for that, I need to shut down my VMs to launch After Effects, and I can't justify procrastinating that much!
[3] Links at http://utau.wikia.com/wiki/Kumi_Hitsuboku
[4] https://www.facebook.com/vocaloidism/posts/568055823290010
The note on katakana isn't very helpful -- but Japanese does indeed have quite few phonemes, 21 according to various random searches (sounds about right). Spanish has 24, French 34 and English 44 according to [1] (No source for Japanese).
Apparently Cantonese has 609-630 or so (due to being a tone language).
Actually that number for English sounds a bit high, I believe 40-41 is more reasonable. Perhaps it includes "all" dialects of English.
My completely non-Chinese intuition (and having asked the question a couple times to native speakers, including a couple who speak Cantonese), is that the number of phonemes for tonal languages tends to reduce vastly when singing. Since you can't hold a tune and produce tones at the same time.
Vocaloid's Hatsune Miku is pretty impressive. "She" has even done "live" shows in Japan and has opened for Lady GaGa in the US.
> Also, while "Chinese" (Cantonese or Mandarin) are quite complicated
Not sure what you refer to in terms of how complicated is the language. Phonetics? Grammar? Vocabulary ?
> I'm not sure if they're particularly complicated from the point of view of emulating singing?
The pronunciation in chinese is not as detached/split as in Japanese, that's for sure. This being said, it's certainly possible if you put sufficient effort into it.
> Not sure what you refer to in terms of how complicated is the language. Phonetics? Grammar? Vocabulary ?
I meant that Chinese have many phonemes, but part of that comes from being a tone language -- and if you're going to emulate singing, you already have to deal with tone change.
However when I think more about it, I really have no idea how phonemes are "tuned" when singing a melody. For a language where tone doesn't carry meaning[1] (or hardly ever does) I would think we simply change tone according to the tune we sing. But what happens if changing the tune of the phoneme changes it meaning? I suppose it is a little like how we must fit the words to the rhythm...
[1] By "carry meaning" I mean in the phoneme sense: how in British English you can have "red" and "read" both be pronounced /red/ -- and there would be no way to answer the question: "Is /red/ a verb or a colour?" (without context). Well you could of course answer "both"...
Relevant clip from wikipedia's summary:
"In the post Tokyo/San Francisco earthquake world of the early 21st century, Colin Laney is referred to agents of the aging mega-rock star Rez of the musical group Lo/Rez for a job using his peculiar talent of sifting through vast amounts of mundane data to find "nodal points" of particular relevance. Rez has claimed to want to marry a synthetic personality named Rei Toei, the Idoru (Japanese Idol) of the title, which is apparently impossible and therefore questioned by his loyal staff, particularly by his head of security, Keith Blackwell."
[edit: While Hatsune Miko's first release was in 2007, although she is portrayed as being 16 years old]
> Written in VB6
That says you don't necessarily need a great programming language to do something interesting, which is a nice change from the norm.
I must admit, off the top of my head the projects that come to mind when thinking about Japan and programming are Ruby, Nilfs2, TOMOYO MACs for Linux and game/system development (Sony PS3/Nintendo etc) -- and more recently Shougo's[1] vim plugins.
I have a "host-family" cousin that was studying at a vocational college in the late 90s, learning game development, and I recall he complained that learning c++ (as a first language) was hard. But that seemed kind of standard for game dev anywhere at the time.
I admit, I have a knee-jerk negative reaction when someone generalizes across an entire nation without much explanation.
Voice "synthesizers" generally use specially developed algorithms. See: https://en.wikipedia.org/wiki/Speech_synthesis#Formant_synth...
Pitch shifting samples is one way to do it. A singer is recorded singing a syllable and that is shifted up or down by software. Artifacts creep in relatively quickly, especially with something as nuanced as the human voice. A variety of pitches and syllables can be sampled and the pitch shifting manually tuned to minimize audible artifacts.
Modeling could also be used, from simplistic models not far removed from the ADSR envelopes of basic synthesis to advanced physically based models. Samples and modeling could be combined to expand the palette of syllables.
Our ears and neural processing of speech and singing are finely tuned to process subtle shades of difference so any technique often sounds artificial. Fortunately this can be exploited musically and great music can be made with these 'artifical' sources.
The oscillator includes the vocal chords as well as a model of the lips for fricatives, plosives, etc.
The resonator includes several major tunable cavities, from the longs to the trachea to the sinuses, nasal cavities and mouth. These resonators form filters called formants which have default configurations for every vowel sound, however they are highly customized for every singer and express a terrific degree of nuance. Synthesis requires a multi-dimensional score somewhat like a speech synthesizer. The score can be dimension-reduced, but it will sound like crap. I would expect it to take about as much time to enter the data as it would to learn to perform it.
While the song is playing you can see the settings. I've played with it once but it's hard to get it right.
First you place the notes to add pitch and length. Then you attach phonetic codes to the notes. So it's not like you are adding words to notes. It's all about how it should sound. Then you also can add things like amplitude settings.
What's also intersting: the image characters of the vocaloid packages are under a CC license and can be used freely for non commercial work, and quite easily for commercial work as well as it seems. The availability of the character and all the community surrounding is often cited by the son makers as a major attract to show their work to a larger public, and not have to take too much time on the image, "marketing" side of it.