Vocaloid 6
vocaloid.com
vocaloid.com
My Vocaloid song recommendation: Ungray Days by the producer Tsumiki. Tsumiki creates a sharp, aggressive sound that is disagreeable at first but really addictive. https://www.youtube.com/watch?v=UvF3Mwj5d4E
Nope, certainly stays disagreeable to me. I wonder what makes people enjoy weird stuff in so many different ways. I might not like this, but I enjoy white noise artist Merzbow [0] or breakcore from Drumcorps [1]
As a small form of resistance to the surveillance state I partake in, I have taught kids to ask any nearby personal assistants to play woodpecker #2 and they find it hilarious.
https://www.youtube.com/watch?v=Y2k8EOBL75o
Not great, not terrible.
On the western side in a similar vein you've got hyper pop coming up from 100 gecs and laura les and what not. This kind of sound, hypertuned and almost as incomrehensible, sounds better to me. You do still get a vein of emotion. I love this sound.
In Japan, the term is "denpa" (電波ソング). Denpa music is intentionally strange as it is catchy, and hypnotic as it is awkward. There are many producers creating high-BPM electronic vocaloid music that is chaotic for effect. It is a bit more twee than the western sounds, as you mentioned, but it can be quite enjoyable if you're in the right mood.
More on denpa music: https://en.wikipedia.org/wiki/Denpa_song
Nanahira playlist, an example of a vocaloid character: https://www.youtube.com/watch?v=NHIyvhJadXM
Explaining Vocaloid in 3 minutes: https://www.youtube.com/watch?v=GODXMGAMpVc
Also, I think you'd enjoy the Song Exploder podcast. If you haven't heard it already, check out the episode where 100 gecs break down how Money Machine was created:
In Japanese, there is no distinction between syllable-final [n] and syllable-final [m]. But in English there is. Traditional romanizations of Japanese will transcribe this as "dempa", for the obvious reasons that (a) that is what the Japanese spelling says; and (b) that is also how the word is pronounced.
I often see English speakers get very confused over exotic modern transcriptions such as "denba" or "senpai", believing there must be a reason they are written that way. But I'm not sure what that reason is supposed to be.
Attempting to approximate pronunciation is a valid theory of transcription, but one which also ought to prescribe that 電気(でんき) be transcribed as dengki; English is not much less discerning of syllable-final [n] vs [ŋ] as it is vs [m]. This is not a position I've ever seen anyone defend in earnest, though.
(Romanization for anglophone is a bit of a lost cause anyway, since we're going to fuck up the vowels no matter what you do.)
That is blatantly incorrect. English converts syllable-final [n] to [ŋ] when followed by a velar exactly the same way Japanese does, and English spelling reflects that. Consider the English words "think", "clunky", or "handkerchief".
How would that suggest that it's reasonable to spell the Japanese word "dempa" as "denpa"?
For demonstrating lack of assimilation of /n/ to following bilabial, there are a couple distinct questions you might ask. It's very frequent for people to preserve the tongue gesture associated with /n/, because a bilabial stop doesn't use the tongue and so [n] is easily coarticulated. But that turns into /mp/ or /mb/ over time because the difference is not easy to hear. In contrast, for a word such as "impossible" where this process completed many hundreds of years ago, the tongue is not used at all in the pronunciation of /mp/. This is a kind of lack of assimilation.
You can also see lack of assimilation in the very people who go to special efforts to pronounce [n] in Japanese words where that is inappropriate.
Note that the English and Japanese phenomena you're talking about are very distinct. This is a fact about the historical development of sounds in English (and Latin...) that doesn't apply to current English, where a sequence like /ng/ will often be preserved across word boundaries. ("One ghost"; this is the only context in which such a sequence can occur at all.[1]) English maintains a robust distinction between /n/ and /m/ and a weaker one between /ŋ/ and the other two.[2]
In contrast, Japanese ん assimilates to whatever follows it, and in the case that nothing follows it it may (rarely) be realized as nothing more than nasalization of the preceding vowel. Word boundaries are not relevant. Japanese does not have a phonemic syllable-final /n/ or /m/ (or /ŋ/). It has a single sound (usually indicated /N/ by specialists, apparently, due to even more weirdnesses that it involves) that gets realized differently in different contexts.
So again - what would justify representing the Japanese sound as "n" regardless of context in languages where, unlike in Japanese, the distinction between "n" and "m" is meaningful?
[1] You say that most would-be /np/s are already spelled "mp", but this is false - the words that are spelled "mp" changed long ago, and do not represent attempts by modern speakers to pronounce an /np/ sequence. They represent attempts to pronounce an /mp/ sequence.
[2] Why weaker? /ŋ/ doesn't have the status the other two do; it cannot begin a syllable. And it makes for a less than perfect contrast with /n/ and /m/ because it has a fairly pronounced effect on the vowel that precedes it, which makes drawing a clean contrast difficult.
Consider "inpainting", "unbiased", and (as suggested earlier) "government", each of which is a synchronically transparent /n/ across a morpheme boundary, yet a cursory survey of recorded English speech suggests that it's pretty common for these tongue gesture associated with /n/ to be absent—infamously, the second syllable of the last routinely loses its coda altogether. This occurs across a transparent morpheme boundary, even with affixes productive in the modern language, even in learned usage.
English does have a lot more wrenches to throw in this, like producing nuclear nasals in a range of situations and not always assimilating across prosodic word boundaries—heck, it probably goes both ways in an utterance like "in my main menu". Words spelled "mp" are reliably [mp] in the modern language, but it's not a simple case as "mp" spelling /mp/ read [mp] and "np" spelling /np/ read [np]; English phonotactics also coerces the nasal in /np/ to a bilabial realization.
> because a bilabial stop doesn't use the tongue and so [n] is easily coarticulated
That doesn't sound quite right—this assimilation surely wouldn't be nearly as globally prevalent as it actually is if that were true.
Try it. While you'd think from the descriptions that a bilabial stop shouldn't care where the tongue goes, I think you'll find it quite challenging to coarticulate [n] with [b]—tongue positioning at lower teeth is pretty obligatory—and much easier to sequence them or produce [mb].
Clearly you can see the unnaturalness of lack of assimilation to call the attempt to do so "special effort"! So of course, the typical anglophone is not going to try to realize [n.p], they'll just see the <np> and read [mp] because that's what they would with any other internal /np/.
> So again - what would justify representing the Japanese sound as "n" regardless of context in languages where, unlike in Japanese, the distinction between "n" and "m" is meaningful?
Now, this gets to an entirely different issue: the purpose of the transcription. You seem convinced that the main goal of romanization is to provide a pronunciation guide for anglophones. But in the context of discussing a niche musical genre on the internet, that's not necessarily a high priority in the first place; you might care more about, say, searchability: we're looking for https://en.wikipedia.org/wiki/Denpa, not https://www.worldbank.org/en/programs/debt-toolkit/dempa.
And in a wider context, the principal users of romaji Japanese aren't anglophones; they're Japanese-speakers who for some or other reason need need to coerce Japanese text into an ~ASCII-subset representation, targeting primarily computer systems with that sort of limitation (most common case being keyboards via IME, hold that thought) and secondarily other people who can read Japanese; and naturally they make the distinctions Japanese makes and largely don't make the distinctions Japanese doesn't make. So unless backed by a marketing department, they tend to produce n (or nn as needed) for ん, because they have a tenuous grasp on how anglos spell [mp] in the first place and でmぱ is garbage that their IME won't convert into the right word, so why type that?
(This is also why pinyin can be the way it is, yet their IMEs have routinely have modes to ignore s-sh/n-ng/n-l distinctions.)
> That doesn't sound quite right—this assimilation surely wouldn't be nearly as globally prevalent as it actually is if that were true.
> Try it.
You know, I mentioned a specific theory here that you've completely ignored. The coarticulation is easy. But it is difficult for a listener to tell the difference between coarticulated [nb] and [mb]. If you're willing to let multiple generations pass, this means that /nb/ will become /mb/ regardless of how easy it is to pronounce.
You will also note that this theory of what's happening mostly cannot be disproved by recordings, which you appear to want to do. You'd want an X-ray or MRI study, something which shows you what the tongue is doing.
> I think you'll find it quite challenging to coarticulate [n] with [b]—tongue positioning at lower teeth is pretty obligatory
This is just obviously false. You have no problems producing [b] with your tongue positioned however you like. You can position it for [t], you can position it for [tʃ], you can position it for [k]. And of the three coarticulations I just mentioned, all of them are well attested, though only the middle one is attested in English ("pshaw", a scoffing sound).
> and much easier to sequence them
This is worthy of comment; there is a linguistic concept called "coarticulation", but all cases of coarticulated consonants seem to have a conventional sequence associated with them. I have no real knowledge or opinion on how real the conventional sequencing is, or how much sequencing is allowed before you stop calling the sounds coarticulated. I suspect that indeed it is easier to sequence two events than to coordinate them to occur at exactly the same time; this is true for all types of events, not just language-related ones. I don't think that the linguistic concept requires absolute synchronization of particular points in time; my understanding is that producing any given phoneme requires some motion and therefore takes place over a nonzero span of time, and "coarticulated" consonants are those for which the durations overlap, not necessarily those for which the durations perfectly coincide.
But I will note that while sequencing of /nb/ is obviously necessary in a way that is not true for /pt/, since /n/ must have nasal airflow and /b/ must not, there is no reason for "coarticulation" of /nb/ to be more difficult than it is in the attested coarticulation /tm/ (exactly the as /np/ for our purposes; /tm/ also features a voicing difference between /t/ and /m/).
> Consider "inpainting", "unbiased", and (as suggested earlier) "government", each of which is a synchronically transparent /n/ across a morpheme boundary
I don't think "government" is a valid example, and you should stop trying to lean on it. In my view, the pronunciation of "government" has as much to do with the morphemes suggested by its spelling as the pronunciation of "comfortable" does with the morphemes suggested by its spelling.
I have no problem with "unbiased"; that's a great example of what we're talking about.
> Clearly you can see the unnaturalness of lack of assimilation to call the attempt to do so "special effort"!
I don't agree with this. I claim that it is common for Anglophones pronouncing "unbiased" to make contact between the tip of their tongue and their alveolar ridge while they pass over the /n/ in the word. (And here, we're on firm ground saying that the internal phoneme is /n/ and not /m/, since it's part of a productive prefix un-.) I further believe that they make no special effort to do so. They may or may not allow a longer duration of nasal murmur than they do in other contexts, to make the /n/ clear; doing this would constitute a special effort. I believe that some speakers will do this and some won't bother. Of those who do, only a small amount of effort will be given to the task.
But the case of English speakers attempting to pronounce Japanese is different. They will go to great lengths to demonstrate that they want to comply with the bizarre textual representation they see. They are happy to produce highly unnatural speech in order to do so. (Which isn't really a problem; they don't really have an alternative to producing unnatural-sounding speech in early attempts to pronounce a foreign language. But this is something they shouldn't encounter problems with.)
> And in a wider context, the principal users of romaji Japanese aren't anglophones; they're Japanese-speakers who for some or other reason need need to coerce Japanese text into an ~ASCII-subset representation, targeting primarily computer systems with that sort of limitation (most common case being keyboards via IME, hold that thought)
> (This is also why pinyin can be the way it is, yet their IMEs have routinely have modes to ignore s-sh/n-ng/n-l distinctions.)
This isn't a flattering comparison for the all-n Japanese transcription system. The pinyin for 吕 is lü. Chinese people don't use German keyboards, which makes the pinyin impossible to type. So where ü contrasts with u, pinyin input methods require you to input V. And Chinese people have responded to this by adopting v-based spellings; it is common to see pseudo-pinyin like "lv" where that pinyin has been generated by an ordinary Chinese person for their own purposes, such as a sign over their business or an online username.
But the letter V is formally not a part of pinyin at all, which means that text generated by the government never uses it and neither do instructional texts.
It is true that this situation is the reverse of the one we're discussing - the Chinese are making a distinction that is required by their language but forbidden by their keyboard, and the fact that they are aware of the distinction makes it easy for them to know what to do. The Japanese are failing to make a distinction that doesn't exist in their language but does exist on their keyboard; this is precisely parallel to the pinyin IME settings you note that will allow the user to ignore phonemic distinctions that they don't make. Again we see that the system maintains the distinction and it's the job of the input method to interpret what the user wants to say.
Chinese IMEs also offer a "double pinyin" input method, in which you type one letter to indicate the onset of a syllable and a second letter to indicate the rime. All syllables are two input-letters long; this model matches the traditional Chinese view of their own phonology. You could just as easily base your system of English transcription on this: instead of "Xi Jinping", 习近平's name would be "Xi Jnp;". Instead of "Sun Yat-sen", we'd talk about "Sp Yixm".
That's what it looks like when you base spelling on what it's convenient for foreigners to type as an intermediate input to their own, different spelling. (As is the case with Japanese input methods.) There are zero people who believe it's a good idea. It's not a better idea in the Japanese case.
ななひら (nanahira) is probably the most well known denpa artist, but she mostly sings normal songs now I think (and has a lovely voice doing so).
ココ is my favourite denpa artist https://youtu.be/2wl8Ofce8TE
To me the most interesting part to vocaloid is the ability for a sole producer to make a complete song without any external help. The vocal parts have always been a barrier, and while emotionless and still lacking in some areas, vocaloids are “good enough” to support a well produced song.
We’ve seen creators rise through the ranks through vocaloid, get experience and exposure, to then move to full professional production with a staff and an actual singer (who’s voice will also be heavily processed, but they have a ton of tuning experience at that point)
I also agree with the parent comment that some creators do benefit from the “mechanical” part. Throwing more links, Giga works with both singers and vocaloids and is pretty good at extracting the best of boths: https://youtube.com/c/GigaVideos
An early example of this was the debut album of Boston which was mostly recorded in Scholz's basement with him on every instrument except drums, then the tapes were mailed to LA for Delp to record vocals. I think it's rather funny in particular that Rock and Roll Band was written and mostly recorded before the band even existed.
CHO-DARI- - Hatsune Miku https://www.youtube.com/watch?v=DU1HjAPvHG8
Mum / 雄之助 feat. flower https://www.youtube.com/watch?v=IjAcngUNiZ8
Hana to Nare / Yunosuke feat. KAFU https://www.youtube.com/watch?v=XqKbuEDvaf8
IA - Conqueror https://www.youtube.com/watch?v=C3E5fb39xcs
Is there any examples of songs out there going the opposite way, trying to use something like Vocaloid to make the voice and singing as realistic and human-like as possible?
Most vocaloid definitely leans into the robotic quality.
It's also not particularly humanoid, just closer than most vocaloids I've heard. Going off of my memory of the song, it's more the intonation then the timbre that stuck out as being more realistic.
I'm afraid I'm not willing to listen through the song currently, I usually dredge up some stuff I'm not wanting to deal with right now when I listen to this artists music.
Twitter Land - STEAKA : https://www.youtube.com/watch?v=e_qQEU_uGjw
Chimera - DECO*27 : https://www.youtube.com/watch?v=c6HKcNVbByc
Start Up! - Nariyama Ryo : https://www.youtube.com/watch?v=LFOV9NbkiJM
My name is - yanagamiyuki : https://www.youtube.com/watch?v=1hj3BDehQGc
Dance with me - Osanzi : https://www.youtube.com/watch?v=n37kZTKbpSM
Highlight - KIRA : https://www.youtube.com/watch?v=AYUNaQaDfa8
Ghost city tokyo - ayase : https://www.youtube.com/watch?v=lWl5viCqGSc
Aqua illumination - PedestrianP : https://www.youtube.com/watch?v=F02fIei8gZU
These songs might be pedestrian (heh) to you, but there is so much niche and experimental all using the same voice--I find this highly fascinating.
The above songs I think give a pretty wide longitudinal view of Vocaloid music and the variety you can find in the fandom, from just Hatsune Miku as the vocal.
This one is incredible! It's like rap+vocoder. I listened to all of these and loved it, thanks!
Never thought I'd find something this interesting from a HN thread, thanks!
Luka(?) sounds pretty natural here, too, especially for 2009. I still remember thinking at first that it was sung by a human.
/jp/ themesong - anonymous ft. Luka, Len, Rin, Miku https://www.youtube.com/watch?v=SCSM4W8vk3Q
/jp/ themesong 2 - anonymous ft. Luka, Miku https://commons.wikimedia.org/wiki/File:Jp_themesong_2.webm
Edit: better audio on the first one: https://www.youtube.com/watch?v=UC2QrK4c3Qw
He also has some funny parodical bits he does, like rapping about having a lot of money/jewels/etc and then the vocaloid characters rap about having a lot of RAM.
The track you reference sounds like chipmunks sped up 2x; it's not unpleasant to listen to, and fun, but I feel it could be made just like that (record at 80bpm, high pass filter, maybe transpose 1 octave, and speed up to 180), no "AI" involved.
It's a synthesizer. It's an alternative to human singers. I can imagine someone seeing a digital piano for the first time. "I'm not sure what it even does. I could just use an acoustic piano. It sounds the same."
This one if an official track for a popular vocaloid rythm game.
Also, at this point the “chipmunk” sound is part of the brand and will be kept to some extent for tracks labelled as vocaloids (it’s kind of a market on its own)
There's also plenty of music directly derivative of the vocaloid scene that maintains a similar aesthetic with 'organic' vocalists and dispenses with some of the awkwardness of vocaloid-oriented compositions. Example: https://www.youtube.com/watch?v=hjJMIWyl_l4
the audio is generated from a voicebank that is a database of prepared phonemes recorded from a voice actor. some packages come with multiple variants of voicebanks, like you could have a "soft" voice and a "vivid" voice.
Human Japanese singers, especially women, tend to operate in a higher octave range than what is common in the west. It's slightly culturally insensitive to take shots at vocal pitch when talking about J-Pop. Pitch is largely a social/cultural construct, and Japan generally leans into the idea of higher pitch -> polite or cute and lower pitch -> aggressive or rude. (e.g. you raise your pitch when talking to your boss, and drop it to express your disgust with someone.) Just putting that out there, not trying to be accusatory or anything. It's just always good to keep in mind that western cultural norms are hardly universal.
For the record, I was responding to the gp saying
> ... pushing the boundaries of pop music in a way that wouldn't be possible with a real singer
=> I felt it was possible to do what the example does by singing slowly and speeding it up afterwards.
The only conceivable surprise is a crude chromatic key change to the minor version of the raised mediant.
You'd think the precision of those dynamic envelopes and timbral games would push the artist to venture out and explore that mediant relationship to create quicker and more jarring harmonic progressions and modulations. But no-- it turns out to be less inventive than the mediant chains emanating from, say, Joni Mitchell and her acoustic guitar over fifty years ago:
https://www.youtube.com/watch?v=3q2jiRUVLgI
(I find some of the lyrics apt, too.)
Compared to the cookie-cutter harmony and melody of the music you linked, even Mitchell's augmented triad in the melody at the end of the chorus sounds like the musical equivalent of solving fast homomorphic encryption.
It's the the audio tech that is on display in the music you linked, so every other musical consideration shifts to the background to illuminate that tech. I get that. But holy shit why does that baseline have to be stuck in the fucking 1650s? While I love the "electrified Vivaldi" hack that is heavy metal from the late 70s/early 80s (Master of Puppets et al), I question whether we really need more than one musical genre based on that parlor trick.
It would be like every stand up comedian ending their set with increasingly theatrical pyrotechnic pull-my-finger jokes. I could laugh my ass off at the absurdity for a year, maybe two. But forever?
My (possibly wrong) impression of your comment is that you seem to have made the mistake of associating complexity with quality in music which is extremely common in those who’ve just started looking into music theory.
Most music needs only the smallest dash of novelty to achieve the perfect mix of the new and familiar to its target audience. If you start attempting to evaluate popular music on what about it is inventive or new, you’re likely to find yourself unable to appreciate most of what people are enjoying and cut yourself off from loving a broad spectrum of musical expression.
You might also find yourself unable to express why you enjoy the music you do like in a way that doesn’t come across as if you’re arguing an objective scientific point——an approach which might undercut your argument by making you unintentionally come across as someone who has just learned a lot of fancy theory jargon and is eager for an excuse to wield it.
What I did find more interesting was the AI "sung" version of Joelene that was doing the rounds a few days ago, based on the voice of Holly Herndon:
https://youtu.be/kPAEMUzDxuo
Interested to see where that goes, although I've got to admit, I'm a purist, and any type of digital vocalist is going to make me go "meh" sooner or later when compared to even a half decent human singer.I even exploited this fact as a way of staying awake a couple times while taking long road trips, as a stand in for caffeine.
Having high energy music is OK, allowing it to disturb the peace is not, time to teach the lessons about manners and being considerate, I suppose (buy a pair of headphones for her, limit her volume so she doesn't suffer early onset hearing loss).
No guarantee she won't turn out to be obnoxious as an adult, but that's the genetic lottery, I'm afraid.
I'm being genuine when I say I'm interested to hear what about this moves you. I almost always get it even if I don't like it. This...I don't get it.
[1] King Gnu - Hakujitsu: https://www.youtube.com/watch?v=ony539T074w
[2] LiSa - Gurenge: https://youtu.be/MpYy6wwqxoo?t=45
IMHO, the best gateway drug is this album/series this is from, though that may just be my personal nostalgia:https://www.youtube.com/watch?v=s_lGrcOtzck
Another interesting vocaloid artist is SOOOO. However, anyone who has struggled with depression/self harm or has suffered abuse should not look them up if they think there is any possibility of being triggered by a mention of it (to the degree where I opened youtube in a private tab to find these links so that I wouldn't risk them being recommended while I'm in a bad headspace).
That being said, https://www.youtube.com/watch?v=RUIelJYMO4U, https://www.youtube.com/watch?v=0OOWWNFTguY, and https://www.youtube.com/watch?v=uGZ0I71Yawo encompass many of those feelings better than any other songs I've heard.
I personally like the chaotic quality that many vocaloid songs carry, and much like other somewhat experimental genre's you start to be able to follow more after listening to it over time.
Why?
Diversity is important, but it has the drawback that compatability suffers. Monoculture is no better, but the tendency to dismiss what others find unique is a recognition of the general (biological) strategy, to conserve some status quo to build community.
General advice, don't have strong feelings about what others like - if they enjoy similar things fine if not their opinions are not worth much in the first place. At the same time, don't be afraid to enjoy what you like or find your tastes change over time, it's natural and not necessarily wrong.
Worst thing you can do to a young person, be old, and tell them you love everything they like - either they will think you a foolish old person or be devastated they aren't as hip as they thought they were...
If you examine this stuff on the basis of harmonic structure, rhythm or arrangement you're basically going down the path of discounting most of electronic music, which is discretized into microgenres just on the basis of using a faster tempo, a different snare hit sound or an unusual mixing strategy. You have to really lean into timbre and texture to find what to appreciate.
Such an excellent and on-point description for 99% of Vocaloid content.
I'm also a musician, there's a lot I can enjoy with vocaloids and utaus.
Also fwiw, 'rhythms stolen from other genres' is a really weird comment for a musician to make.
EDIT: That being said Miku can also be amazing: https://www.nicovideo.jp/watch/sm23655091
The original in a more Miku style: https://www.youtube.com/watch?v=PqJNc9KVIZE
Or a more jazzy version but not as good as the first: https://www.youtube.com/watch?v=CwG9viczjhs
Compare to a human cover: https://www.youtube.com/watch?v=fEsyBaG-uNw
Here is a different example:
https://www.youtube.com/watch?v=gIBdpzporFs
Of course the original version which is true to Vocaloid style is like this:
https://www.youtube.com/watch?v=Mqps4anhz0Q
Even if Vocaloid is capable of much more than just Miku, Miku is immensely influential in the subculture.
Does it sound better live ( https://www.youtube.com/watch?v=K_xTet06SUo )? Or for a closer comparison with the first link https://www.youtube.com/watch?v=nepNc0Gk1E8 ? Sure, even though the Vocaloid style has its own charm. But you also do not need to be able to be able to sing in order to create a song with Vocaloid, so overall it's a great tool. If the song is good, someone will eventually cover it live.
Vocaloid allows many who want to compose but do not sing to participate in a remixing ecosystem. The collaborative nature of the community is an incredible strength.
It's definitely not my thing, but Tsumiki's use of vocals is interesting and well executed.
Efficient Pitch Detection Techniques for Interactive Music
https://ccrma.stanford.edu/~pdelac/PitchDetection/icmc01-pit...
New Phase-Vocoder Techniques For Pitch-Shifting, Harmonizing and other Exotic Effects
I know of Sinsy [0] but I couldn't get it working. eCantorix [1] is very old and rudimentary (it uses espeak underneath [2]).
Searching just now I see OpenUtau [3] but I have no experience with it.
Seems crazy there isn't a good FOSS solution for this.
[1] https://github.com/divVerent/ecantorix
Also has integration with NNSVS (neural net based vocal synthesizer) and the entire UTAU ecosystem.
I would say this is the good FOSS solution!
https://pages.cpsc.ucalgary.ca/~hill/extra-synthesis-example...
I have recently found Synth V, which uses AI trained on a singer and it can produce shockingly good results. It can also still used to produce similar sounds to vocaloids as well.
Here are some of my favourite Synth V covers and songs: https://www.youtube.com/watch?v=jU_CG_FF6WI https://www.youtube.com/watch?v=EKOSQGKn5Cw https://www.youtube.com/watch?v=ShG8Ij6_Hbo https://www.youtube.com/watch?v=cXv_vKX6Y-0
And knowing the limitations of vocaloid never stopped anyone from trying. Mitchie M is a great example.
As far as the iconic characters go, they're moving to an entirely different engine (Piapro NT) anyways, so I wonder how future works created using them will sound.
Synth V is extremely fast and the output is shockingly good with some tweaking - good enough to be indistinguishable for many people.
The audience wanting more vocaloid-like sound and the one wanting more realistic sound aren't really the same, and the overlap between them, I suspect, is not large. So it makes way more sense to capture the latter group by creating that more-realistic-voice spin-off product line, as opposed to being forced to choose between the realistic and vocaloid-like target demographics.
We already know the size of vocaloid-sound target audience, but I bet the audience for realistic-sound synthesis is going to be magnitudes larger (mostly because of versatility of where that tech could be useful, while with vocaloid it is mostly constrained to music production and vocaloid-related visual arts accompanied by a typical vocaloid voice).
I kind of suspect Piapro NT is going to end up being a bust, with a pivot back to Yamaha's platform. We're a few years in and they've still only released Miku, none of the other Crypton Future Media characters, and a lot of people are sticking to the Vocaloid 4 release because they're not fans of how NT sounds. Now V6 is out and the technological gulf is widening.
[0] https://www.youtube.com/watch?v=341IsnWdaT4
Vocaloid was developed as a DSP project at a Spanish university, supported by Yamaha, and first marketed in the UK - with almost zero success.
It wasn't until it was personalised/mythologised by combining it with anime in Japan that it really exploded. Crypton have very deliberately milked it for everything it's worth. Making the output branded but royalty-free was marketing genius.
One bizarre thing about it - it's converging with autotuned vocal stylings applied to organic human vocals, especially in genres like hyperpop. It's becoming increasingly hard to tell them apart.
Another bizarre thing - there's a kind of corrosive psychedelic "What does human mean?" aesthetic that applies to AI art in general. Vocaloid music is a subset of other emulated artforms: simultaneously cute, spectacular, and naive, but also disturbing, overwrought, and uncanny.
This looks very much like a new era, in the manner of Baroque, Romantic, and Modern. It's going to be explosively transformative for all of the arts, and it's not obvious yet that there's going to be much left that's still recognisable after the dust has settled.
It's a great example of a company seizing an opportunity and not stepping on the fan communities that built them said opportunity.
For example DiffSinger sounds better (but not perfect). It's using diffusion model like the popular AI image generators. I cannot find English demos but these are not too bad:
https://www.bilibili.com/video/av599316695/ (unmute the player)
https://www.youtube.com/watch?v=hJ0wNFZGECo (bad mixing but the voice is still better than vocaloid)
code and huggingface demo:
https://github.com/MoonInTheRiver/DiffSinger
https://huggingface.co/spaces/Silentlin/DiffSinger
and other techniques (demos at bottom):
https://r9y9.github.io/projects/nnsvs/
Of course these are not complete user friendly softwares but I would expect Vocaloid would have something like these implemented.
They took the credit for your second symphony /
Rewritten by machine and new technologyVocaloid was the only real competitor in vocal synths for so long they completely stopped innovating. Very IE6-esque.
[1]: https://vi-control.net/community/threads/synthesizer-v-vocal...
The "realism" of Vocaloid is not exactly a priority--a lot of people enjoy the robotic sounding voice. Making ridiculous statements like "blow it out of the water" really rubs me the wrong way of the typical techbro looking for objective KPIs instead of what the general vibe of the fandom.
To me, this is like saying "DAWs blows physical instruments out of the water because it can create a lot more sounds!"--you're right, but you're totally missing the point.
Here's a KPI: nobody is making music from Synthesizer V, and it's not exactly very popular.
> Vocaloid was the only real competitor in vocal synths for so long they completely stopped innovating. Very IE6-esque.
This is probably the only fair argument here. The Vocaloid editor has always been quite difficult to use, and recent competitors in that space has made improvements to it. Hatsune Miku's new release moves away from Vocaloid in favor of a Piapro/Crypton voice engine with a better editor, for instance.
However, I do think that it's a lot more limited in what it can do. The voices in Vocaloid sound quite synthetic, and there's only so far you can go if you're not going for a stereotypical "Vocaloid-sounding" vocal. IMO, Synth V can get pretty close to the Vocaloid sound with Eleanor Forte, but in addition can also produce some much more realistic-sounding vocals, especially with Solaria, Eleanor AI, or the newer Dreamtronics AI voicebanks.
People are making music with Synth V! Check out AIKA who has made some amazing songs using Synth V.
From Wikipedia:
> In August 2010, over 22,000 original songs had been written under the name Hatsune Miku. Later reports confirmed that she had 100,000 songs in 2011 to her name.
(Admittedly, those are not unlikely all offshoots of the anime/manga/etc. fandom though.)
Vocaloid music has a huge following in the US. It's just a subculture thing and not part of the mainstream sphere...which is a good thing.
I'm not sure you can paint the Japanese music industry like that? It's more like, this gives indie songwriters an easily accessible tool.
I didn't mean to. I was just wondering why Vocaloid is so popular specifically in Japan.
It initially became a thing in Japan because Japanese has fewer phonemes than most mainstream languages, so it was easier to make. Then it gained popularity for the cute characters and the freedom it gave to people who would go on to be music composers. Many popular Japanese composers got their start making and uploading Vocaloid songs on Youtube/NicoNico. On top of that, they have had a series of well designed rhythm games on several platforms for decades now which were pretty popular. These days they have a fairly popular mobile gacha/gambling game too.
Nostalgia is probably also a big driving force since many people grew up with either the music or the games.
These days it's still pretty popular within the anime fanbase outside Japan. Pre-Covid there used to be a vocaloid concert series every year which would alternate between Europe and US tours. The reason it hasn't gone too much more mainstream is likely that English and related languages have thousands of phonemes, so making a good voicebank is significantly harder.
All the stats about the rise of the old music are supporting my claim:) Young people, invest in your talents, with analogue processes in mind.
This will be a huge market.
If you want to synthesize more realistic sounding voices there are better options.
The purpose seems as usual to get rid of as many human musicians/singers they can so that everything can be made by a single person (or AI in a few years), therefore saving money. In a different context, the transition from multi elements bands to one man bands with keyboard, then finally karaoke, in many cases is motivated by costs as well.
There's hardly any evidence automation ever destroys jobs; it seems to actually create them. It's very silly people just keep claiming this.
People do art because it's fun. Nothing will stop people from getting together in a band and jamming, because it's fun.
And, this is subjective, but a lot of those tool-enabled artists seems to do better even in absence of such automagic enablers than non-enabled. Good AIs sharpen humans into unassuming Olympians, bad AIs just fall out of the Internet attention span.
All jokes aside, I will throw serious money at the first streaming service that implements an 'autotuned' tag, and lets me filter anything tagged with it out of my stream. Like, $100 a month. Maybe more.
Well, I hope you don't like listening to any music made after ~2010, then. Melodyne is completely standard for any modern vocal processing, in all genres of music.
I get where you're coming from. I remember Celemony announcing their polyphonic tuning engine a little over a decade ago, and remember buying it as soon as it came out and re-tuning a load of Imogen Heap tunes and loving it, in something like Reaper v2. I know how prevalent Melodyne is, I know the commercial and production related justifications for its use. But I also know for a fact that this is not the case 'in all genres of music'.
Sometimes, I want to listen exclusively to music with real vocals. I love the first CHVRCHES album, it's one of my favourite albums of the last decade. As it happens, it took me ages, a couple of years, to figure out that the reason I liked it as much as I did was the lack of autotune. I (ironically) figured this out when they released their second album which _did_ utilise autotune. I never made it all the way through listening to that second album.
You may as well say "I don't listen to music that uses EQ" or "I don't listen to music that uses compressors."
Not the same as either of those things. You might as well be arguing that I'm claiming that volume knobs shouldn't exist.
Pick an equivalent like photography. In the photography world, EQ is like a fill light. A compressor you could compare to a polarising filter. Autotune (or drum quant) though, is like photoshopping. Removing all of the skin blemishes and imperfections at best, and at worst and more often than not, it's fake disproportionate waist and ass booty enhancement.
Compression is an effect that's applied. Selective tuning, like quantisation, is the result of an _interactive, selective edit_. It's the equivalent of photoshopping out the spots on my ass and making my hips a little narrower. Sure I could have not had that ingrown asshair, and maybe my hips would be narrower if I worked out. But the reality is different.
(Also IMHO ratio is a better single indicator of severity of compression, rather than dB)
> Compression is a uniformly applied filter. It's like a contrast adjustment or colour curve. It also affects amplitude only.
Very incorrect. Many famous compressors are known for the color (saturation, technically speaking) that they add to the sound. Hell, some are known for how badly they destroy the sound, like the Level-Loc. Analog-style compressors (which most engineers still use for vocals, in particular) also react very differently to different input gains, so it's not a uniformly applied filter.
> (Also IMHO ratio is a better single indicator of severity of compression, rather than dB)
It is not. Talking in terms of total gain reduction is more indicative of the effect on the sound. Using a 100:1 ratio (in practice, a limiter, something like Pro-L on Safe mode) is very common on vocals to catch quick peaks, but only catching a few dB of gain reduction. You won't notice that a limiter is being used on vocals this way. But you would notice if I used a 2:1 ratio on a vocal and set the threshold all the way down, crushing the dynamic range. You also can't talk about ratios when talking about using compressors in serial, which again, is standard vocal processing.
Recommend one, I'd appreciate it.
> Many famous compressors are known for the color (saturation, technically speaking)
Are you comparing colouration / saturation from a classic compressor to autotune?
I agree with you on most of the last paragraph. Thanks for the explanation.
It’s very difficult to tell if autotuning is used, if it is used well.
I don't want to listen to artists who use autotune. I _do_ want to listen to artists who don't. I know that there definitely is a set of artists who explicitly don't use it, who are more concerned with accurately expressing their creative intent and musical virtuosity than they are with gaining popularity and mass appeal. What I am saying is that I would like to be able to consciously choose to only listen to and support those musicians, and I will gladly pay a disproportionate amount for it. Anecdotally, I know for a fact that I am not the only one who feels this way.
Also btw -
> when artists actually release music that uses no autotuning, it tends to be less popular
Normally when you make a statement like this on HN I'd expect to see a citation or reference to where that statement came from.
I want to be able to choose to consciously support artists who don't use autotune. I want to be recommended and discover new artists, who consciously don't use autotune. I want to be able to choose to listen to real vocals as a genre. I want that choice.
im making fun of you for being silly
Note that many recording artists do not actually make that choice; it happens further up the chain. Regardless, whether or not a singer or producer uses a particular effect on their voice does not distinguish between the vocals being 'real' or not. If you simply don't like the way it sounds, on the level of artistic taste, well, you can make that judgement for yourself, but, to claim it's more profound than that is just pure pretense.
Also, why are we talking about autotuned vocals on a thread about a speech synthesizer? Claims of 'real vocals' are already out the window at that point.
If they don't make such choices for themselves, they're not much of an artist, more like mass-produced manufactured candy pop/r&b/etc...
Because of course there can't be a cutoff point, with my argument applying to the level of involvement I described (if one's vocals are to be processed with autotune or some major effect is a major thing for a vocalist to leave to others, even level of reverb or echo will often be a thing to debate with the producer/engineers) and not to ridiculous meta-levels like building your own DAW.
Perhaps there is also a standard mic for Vocal Realness. I assume it must be a SM58; it is, after all, well-known that having a low-pass switch on the microphone lowers the industry-calibrated Realness Score by at least 250 mSpr(ingsteen). More if it's on.
edit: A friend also pointed out to me the inherent Springsteen ceiling of computer-reproduced audio. And you know, he's right, I'm going to go find a chamber and hire some monks for the true realness that only an authentic Gregorian chant can provide. Denon sells them in twelve-packs, you know.
It's not based on some technical argument or law of physics. People either get why, or they don't.
If you want some kind of technical justification, those effects you mentioned just add some sparkle (flanger) or fixes the balance (often just to make a recording, which begins deader in a sound-treated studio, sound closer to real life environment (reverb) and dynamics (compression).
Autotune, on the other hand, changes the pitch and vibrato, you know two of the main things a singer is supposed to produce. And if overdone as "effect", it also fucks the timbre.
And let's be real, nobody says this about some singer using autotune to fix a flat note or two. It's the autotune-as-effect (whether T-pain levels or more subtle) that people complain about.
It's gatekeeping bullshit.
Put frankly: your take is a small one. It's one that weakens the idea of art, and, no less importantly, is cruel and sabotaging to people. You should change your mind, but you're kind of glorifying in that cruelty and that smallness throughout this thread (which is gross!) so I will not be holding my breath for it.
Except if you thought that when complaining about autotune being bad, we were talking about some rare band that uses it as a creative tool, and not about the millions that use it as a clutch or for the 1000000th recreation of the same BS sound...
If I think it's "bad" then it has _zero_ value to _me_. Why does this bother you or anyone else?
And people can still be able to differentiate between stuff they merely don't like (taste) and stuff they consider detrimental to music in a larger way. In fact they might even like the latter and still consider them detrimental (e.g. I find some commercial pop tunes catchy, but consider them a bad musical and societal influence).
Do you have anything in mind that sounds misinformed? Or just can't fathom that anybody who can tell what a granular delay or an LFO or an automation envelope is can't possibly dislike Autotune?
> But the idea that they're not "real vocals" is an attempt at shitty gatekeeping that has no place in music.
I'm not sure whether you see the irony in this statement. You're saying my opinion has no place 'in music'?
> This wannabe arbiter does not get to decide what "real" is or what "art" is.
I never claimed to be the arbiter of what art is. Art is subjective. It's a matter of personal preference. Like not wanting to listen to autotune.
Lastly, you seem really angry. I'm really sorry if anything I've said has upset you.
Man, this is basic Popper stuff. If you'd said "I don't like autotune", I'd have probably agreed with you. If you're saying that what others are doing aren't "real" because you don't like it, that's a whole different kettle of fish.
> I never claimed to be the arbiter of what art is.
You picked the word "real". Words mean things, and "real" does not mean "to my preference". It is a an assertion of legitimacy, it is that assertion to which all of my comments in this thread are directed, and it's something that neither you nor I get to take away from somebody.
> Lastly, you seem really angry. I'm really sorry if anything I've said has upset you.
I wouldn't say that I'm angry, I've been on the internet a long time and random posts have to be really special to do that, but I do write sharply when I care about something. If one believes genuinely in the openness and democracy of art--and I do--I don't think there's a properly strident reaction to the implications you laid down that wouldn't be a little bit testy.
There is nothing wrong with that. It is still art, I never claimed it wasn't. Deepfaked actors can still constitute 'art' (see Sassy Justice with Fred Sassy for a notable example). Photoshopped photos are still 'art'.
However, we don't refer to those 'real actors' or 'real photos'. They are examples of creative expression through manual editing. Photoshopped photos can be real art without being real photos, but there's a reason why National Geographic photographers don't crazy with the spot healing brush.
There are absolutely edits so subtle that without having seen or heard the original you'd have no way of knowing it was modified at all. Pitch correcting someone's voice up 1/1400th of a step is not going to be noticeable no matter how perfect one thinks their hearing is. These kinds of subtle changes are far more common than the drastic and noticeable edits or even smaller but still quite large edits where people with a trained eye/ear will notice but the average person wouldn't.
Perhaps so. I'm not one of those people that picks up on (or claims to) these subtleties, so I wouldn't know. I wonder if any studies have been done on this.
You're obviously entitled to share your opinion, but also I also find your strawman a little offensive. Vocal pitch correction is a very real and very noticeable category of processing, absolutely not the same realm of snake oil bullshit as $30,000 gold plated wires. If it was, it just wouldn't exist. Why would anyone ever take the time to write software that correct vocal tuning if it made zero perceivable difference to the output?
> There are absolutely edits so subtle [...] Pitch correcting someone's voice up 1/1400th of a step is not going to be noticeable no matter how perfect one thinks their hearing is.
Nobody ever 'corrects someone's voice up 1/1400th of a step'. It just doesn't happen, human voice can't consistently hold a frequency to that resolution for more than a few milliseconds. For reference, a 1/4 step vocal oscillation is only considered 'moderate vibrato' [0]. Even a 1/128th pitch variation is rarely considered noticeable or consistent enough to correct. This is miles away from a '1/1400th' (which is also a very strange fraction to choose btw, even when you take harmonics into account).
I posted this in another comment, but you should take a couple of minutes and take the MusicLab Tone Deafness test [1]. It'll give you an idea of what 1/64th variance sounds like, and how noticeable it is or isn't to your ears.
I'm not sure the point of the Tone Deafness test - I scored a 31/32 but on many of the questions if one tone had been replaced with the lower/higher tone I'd never know it had been replaced despite my ability to tell them apart from one another. My point wasn't that you couldn't tell A from a very similar B (you can!) but that if A had been replaced by that similar B in the first place you'd never have known (because you couldn't!). You'd need to have been privy to the editing process to actually know if it had been edited at all.
Imagine an A/B test where A has been removed entirely and you've been asked which edits B has made to A. It's a very different test from an A/B test where you can compare A with B. While a trained ear may be able to hear obvious edits there are literally hundreds of non-obvious edits made to songs during production that you couldn't possibly know about without hearing the original recording to compare against.
There's a pretty interesting test[0] on themusiclab.org, that will test how good your microtonal perception is. See how you score - it might align with how well you're able to perceive autotune. Here's my score [1]. This was from my first attempt a few months ago, but I've done it a couple more times since and my score is pretty consistent - whether I like it or not.
[0] https://www.themusiclab.org/quizzes/td [1] https://imgur.com/a/5R3K43Z
Then you have problem of the budget. Studio time is quite expensive and after nth take hard decisions have to be made. Booking another session or fix the tuning with software?
Also take into account that vocals that you hear in songs, that are not obviously autotuned may actually be composed of dozens of takes. That technique is called comping and the mixer (sometimes together with the artist) would choose the best take out of dozen often for each phrase or even single word. Sort of like natural autotune, when e.g. only phrases that sound in tune are picked.
Any examples? Are they all fairly small artists, or are there some big names who explicitly don't use it?
Not sure what you have in mind, but people don't need to listen to the top-20 or BS R&B.
Even if you want to listen to electronic music, there's a big universe of artists who have nothing to do with the "autotune" sound and modern commercial productions.
Tom Waits and Coldplay (or specifically, their mixing engineers) have gone on the record saying they use Melodyne in their marketing brochures.
Again, it is literally impossible for a human to know if pitch correction was used on a song. But if it's a song that was released after 2010 and mixed by an engineer, they used pitch correction, guaranteed.
[0] https://musicmarketing.ca/DNET/rack/brochure_melodyne_3.pdf
Not exactly. A guy who worked with them said it. They also have worked with dozens of others which would be where they used it. Also Waits recorded for 25+ years before Autotune (much less Melodyne) was even a thing.
And instead recorded the same piece several dozen times and for the final recording spliced together what sounded best.
This is like complaining that you don't want authors using spell check, they should have to retype the entire paragraph every time they make a mistake!
The end result is the same, the only difference is the path to get there.
There may be some fair complaints about usage at live shows, but for recordings, the end result is going to be the same if melodyne is used or if the artist is recorded again and again until everything is "perfect".
I think you are forgetting that the autotune IS their creative intent. T-Pain is probably the most famous heavy autotune user of all time, despite having an amazing voice without it. It's an intentional effect - like how electric guitars aren't "worse" because you aren't hearing the raw sound of the string.
If you can’t tell the difference, why do you care? If you can tell the difference, why do you need the label?
Speaking as an amateur musician, I think these arguments reflect a lack of understanding of how music production works, and how that’s connected to the artist’s creative vision. I’ll say that a big part of the blame lies with poorly used autotune. Just to pick an example, the entire first season of Glee is especially bad, to the point where I want to leave the room. Then there’s various places where autotune has been overused on singers who don’t actually need it to begin with (Bublé comes to mind), or where autotune has been used to cover up some sloppy singing. (Bublé is particularly illustrative—it is known that he uses autotune, but he has said that he doesn’t… I suspect that Bublé is simply unaware that autotune is being used. I suspect that many other singers are also unaware that they are being autotuned—but you can sometimes find clear evidence for it when you analyze their songs with a computer.)
But autotune is also used, manually, by producers, to make small adjustments as needed to improve a take. It can mean that the singer does fewer takes to nail the song, because with your comping and autotune choices, you get what you need faster. The amount of comping and autotune that you do is a matter of style and situation—producers are free to do comping and autotune as they see fit, to capture what they think is the best version of each part.
It means that you can say, “take 3 has the most beautiful phrasing and a lot of soul to it, it’s just a little sharp” and then manually adjust the pitch by the amount you think is appropriate.
> Normally when you make a statement like this on HN I'd expect to see a citation or reference to where that statement came from.
I’m not sure why you have that expectation or think that it is at all reasonable. Like everyone else, I have a lifetime of experience. Not everything that I tell you was written down in the first place, and I can’t reasonably be expected to remember the exact source for every piece of information, nor should I be expected to shut up just because I don’t provide a source for something I say.
I don't want to get halfway through listening to a song to hit autotuned vox. I just don't, I'd like the option to filter on that. I would like a recommendation/discovery engine that eliminates that anxiety.
> a lack of understanding of how music production works, and how that’s connected to the artist’s creative vision
Quite the opposite. I understand music production extensively. I studied it, I have used it, half my social circle are studio engineers or conservatoire grads. I just want the option of filtering on artists whose creative vision _is_ the original vocal recording.
> I suspect that Bublé is simply unaware that autotune is being used
The only reason that man has a career is because his geriatric target audience is too old to know what autotune is (hence his denial is plausible). I'm pretty sure his record sales depend almost entirely on the abuse of autotune for the hard of hearing. He isn't an example I'd have used in this argument.
> But autotune is also used, manually, by producers, to make small adjustments as needed to improve a take. It can mean that the singer does fewer takes to nail the song, because with your comping and autotune choices, you get what you need faster
I totally know and appreciate all of this. Recording without it is considerably more effort and money, it's often not economically viable. Nevertheless, I want to be able to sometimes filter on it.
> I’m not sure why you have that expectation or think that it is at all reasonable. Like everyone else, I have a lifetime of experience. Not everything that I tell you was written down in the first place
When your argument is your opinion or something drawn from your personal experience, that's fine. However your words were: "Fact is, people say this, but when artists actually release music that uses no autotuning, it tends to be less popular". Stating a 'counterintuitive fact' the way you did suggests it was proven in a study or experiment of some kind. When putting arguments like that forward, most posters on HN tend to reference their sources. Otherwise your 'fact' is just an opinion. There's nothing wrong with that, but it is not a 'fact'. If you state something is a fact, I think an expectation that you can back it up with an external reference is entirely reasonable.
> nor should I be expected to shut up just because I don’t provide a source for something I say
Not sure where I said that you were expected to shut up. If I implied it I apologise. However, I stand by my expectations of providing external references for things stated as fact.
Is “anxiety” the right word? This is an unexpectedly serious way to phrase things and I’m not sure I understand you correctly. If you’re viscerally bothered by the presence of autotuning, I can see why that must be very frustrating.
> I just want the option of filtering on artists whose creative vision _is_ the original vocal recording.
I think the idea of “original vocal recording” gets weaker the closer you look at it. You want something that’s not autotuned, fine. You want to logically explain your personal preferences in terms of preferring the “original” audio recording? The logic doesn’t hold up. That’s okay, you don’t need to logically explain your personal preferences.
In fact, I’d rather you didn’t. Pet peeve of mine. Maybe if you get a visceral reaction to bad autotune, I get a visceral reaction when someone explains the “logic” behind their personal preferences. It’s aesthetics. Maybe someone can bridge the gap between logic and aesthetics someday, but I haven’t yet met somebody who has succeeded.
> The only reason that man has a career is because his geriatric target audience is too old to know what autotune is (hence his denial is plausible).
The reason I brought up Bublé is because it illustrates that autotune is sometimes used even when the singer denies it, even when it’s unnecessary. Please don’t lay on the hate.
Gross.
The kind of artists worth listening to wouldn't touch autotune with a 1000ft pole. They're less popular to begin with - but can still have tens of millions of fans globally (say, someone like Tom Waits).
Whoops, bad news! Tom Waits' mixing engineer is explicitly mentioned in Melodyne's brochure.[0] Melodyne 3, too, so he's been using it for quite a while. You may want to shorten the length of that 1000ft pole.
[0] https://musicmarketing.ca/DNET/rack/brochure_melodyne_3.pdf
Which is not Tom Waits. A guy worked with 200 clients, has done gigs with Waits, and puts the most famous client names as ("has worked with"), not necessarily the one they used Melodyne on (which he doesn't even claim).
An alternative you've also heard extremely often is a singer recording the same line 100 times, then producers going through each word (or each syllable) to cherry-pick the sample where it was most on-pitch and blend them together.
Neither one of these represents the singer's real ability (whatever that even means), but both can be unnoticeable by 99.99% of the population when applied skillfully.
That seems a bit too general. If we narrowed that down to a few popular music genres I would probably subscribe to it. Plenty of solo artists wouldn't be caught dead using autotune or melodyne or whatever (outside of using it for effect/intentional distortion). When being an exceptional singer is your entire brand, you don't want to show up with training wheels.
I mean, it's a joke, it's South Park. It's why I said 'jokes aside' in my next comment. However, if you're into jokes and such, I'd highly recommend that entire episode.
Saw a cool device recently (it's old) called Pocket Miku
>VOCALOID6 is an AI-based technology created by Yamaha to fully support the musical expressiveness of creators from all perspectives, offering an even more natural singing voice than ever before together with unprecedented freedom to express your vocal ideas. This product lets you express your ideas on the spot in vocal form while producing music.
What does this even mean? It reads like GPT-3 output.
“Software that sings for you” is a great explanation, which should have been right there at the top of the landing page instead of that endless stream of buzzwords.
Its just voice synthesizing