My Vocaloid song recommendation: Ungray Days by the producer Tsumiki. Tsumiki creates a sharp, aggressive sound that is disagreeable at first but really addictive. https://www.youtube.com/watch?v=UvF3Mwj5d4E
My Vocaloid song recommendation: Ungray Days by the producer Tsumiki. Tsumiki creates a sharp, aggressive sound that is disagreeable at first but really addictive. https://www.youtube.com/watch?v=UvF3Mwj5d4E
I'm being genuine when I say I'm interested to hear what about this moves you. I almost always get it even if I don't like it. This...I don't get it.
[1] King Gnu - Hakujitsu: https://www.youtube.com/watch?v=ony539T074w
[2] LiSa - Gurenge: https://youtu.be/MpYy6wwqxoo?t=45
IMHO, the best gateway drug is this album/series this is from, though that may just be my personal nostalgia:https://www.youtube.com/watch?v=s_lGrcOtzck
Another interesting vocaloid artist is SOOOO. However, anyone who has struggled with depression/self harm or has suffered abuse should not look them up if they think there is any possibility of being triggered by a mention of it (to the degree where I opened youtube in a private tab to find these links so that I wouldn't risk them being recommended while I'm in a bad headspace).
That being said, https://www.youtube.com/watch?v=RUIelJYMO4U, https://www.youtube.com/watch?v=0OOWWNFTguY, and https://www.youtube.com/watch?v=uGZ0I71Yawo encompass many of those feelings better than any other songs I've heard.
I personally like the chaotic quality that many vocaloid songs carry, and much like other somewhat experimental genre's you start to be able to follow more after listening to it over time.
Why?
Diversity is important, but it has the drawback that compatability suffers. Monoculture is no better, but the tendency to dismiss what others find unique is a recognition of the general (biological) strategy, to conserve some status quo to build community.
General advice, don't have strong feelings about what others like - if they enjoy similar things fine if not their opinions are not worth much in the first place. At the same time, don't be afraid to enjoy what you like or find your tastes change over time, it's natural and not necessarily wrong.
Worst thing you can do to a young person, be old, and tell them you love everything they like - either they will think you a foolish old person or be devastated they aren't as hip as they thought they were...
If you examine this stuff on the basis of harmonic structure, rhythm or arrangement you're basically going down the path of discounting most of electronic music, which is discretized into microgenres just on the basis of using a faster tempo, a different snare hit sound or an unusual mixing strategy. You have to really lean into timbre and texture to find what to appreciate.
Such an excellent and on-point description for 99% of Vocaloid content.
I'm also a musician, there's a lot I can enjoy with vocaloids and utaus.
Also fwiw, 'rhythms stolen from other genres' is a really weird comment for a musician to make.
EDIT: That being said Miku can also be amazing: https://www.nicovideo.jp/watch/sm23655091
The original in a more Miku style: https://www.youtube.com/watch?v=PqJNc9KVIZE
Or a more jazzy version but not as good as the first: https://www.youtube.com/watch?v=CwG9viczjhs
Compare to a human cover: https://www.youtube.com/watch?v=fEsyBaG-uNw
Here is a different example:
https://www.youtube.com/watch?v=gIBdpzporFs
Of course the original version which is true to Vocaloid style is like this:
https://www.youtube.com/watch?v=Mqps4anhz0Q
Even if Vocaloid is capable of much more than just Miku, Miku is immensely influential in the subculture.
Does it sound better live ( https://www.youtube.com/watch?v=K_xTet06SUo )? Or for a closer comparison with the first link https://www.youtube.com/watch?v=nepNc0Gk1E8 ? Sure, even though the Vocaloid style has its own charm. But you also do not need to be able to be able to sing in order to create a song with Vocaloid, so overall it's a great tool. If the song is good, someone will eventually cover it live.
Vocaloid allows many who want to compose but do not sing to participate in a remixing ecosystem. The collaborative nature of the community is an incredible strength.
CHO-DARI- - Hatsune Miku https://www.youtube.com/watch?v=DU1HjAPvHG8
Mum / 雄之助 feat. flower https://www.youtube.com/watch?v=IjAcngUNiZ8
Hana to Nare / Yunosuke feat. KAFU https://www.youtube.com/watch?v=XqKbuEDvaf8
IA - Conqueror https://www.youtube.com/watch?v=C3E5fb39xcs
Is there any examples of songs out there going the opposite way, trying to use something like Vocaloid to make the voice and singing as realistic and human-like as possible?
Most vocaloid definitely leans into the robotic quality.
It's also not particularly humanoid, just closer than most vocaloids I've heard. Going off of my memory of the song, it's more the intonation then the timbre that stuck out as being more realistic.
I'm afraid I'm not willing to listen through the song currently, I usually dredge up some stuff I'm not wanting to deal with right now when I listen to this artists music.
Twitter Land - STEAKA : https://www.youtube.com/watch?v=e_qQEU_uGjw
Chimera - DECO*27 : https://www.youtube.com/watch?v=c6HKcNVbByc
Start Up! - Nariyama Ryo : https://www.youtube.com/watch?v=LFOV9NbkiJM
My name is - yanagamiyuki : https://www.youtube.com/watch?v=1hj3BDehQGc
Dance with me - Osanzi : https://www.youtube.com/watch?v=n37kZTKbpSM
Highlight - KIRA : https://www.youtube.com/watch?v=AYUNaQaDfa8
Ghost city tokyo - ayase : https://www.youtube.com/watch?v=lWl5viCqGSc
Aqua illumination - PedestrianP : https://www.youtube.com/watch?v=F02fIei8gZU
These songs might be pedestrian (heh) to you, but there is so much niche and experimental all using the same voice--I find this highly fascinating.
The above songs I think give a pretty wide longitudinal view of Vocaloid music and the variety you can find in the fandom, from just Hatsune Miku as the vocal.
This one is incredible! It's like rap+vocoder. I listened to all of these and loved it, thanks!
Never thought I'd find something this interesting from a HN thread, thanks!
Luka(?) sounds pretty natural here, too, especially for 2009. I still remember thinking at first that it was sung by a human.
/jp/ themesong - anonymous ft. Luka, Len, Rin, Miku https://www.youtube.com/watch?v=SCSM4W8vk3Q
/jp/ themesong 2 - anonymous ft. Luka, Miku https://commons.wikimedia.org/wiki/File:Jp_themesong_2.webm
Edit: better audio on the first one: https://www.youtube.com/watch?v=UC2QrK4c3Qw
What I did find more interesting was the AI "sung" version of Joelene that was doing the rounds a few days ago, based on the voice of Holly Herndon:
https://youtu.be/kPAEMUzDxuo
Interested to see where that goes, although I've got to admit, I'm a purist, and any type of digital vocalist is going to make me go "meh" sooner or later when compared to even a half decent human singer.I even exploited this fact as a way of staying awake a couple times while taking long road trips, as a stand in for caffeine.
Having high energy music is OK, allowing it to disturb the peace is not, time to teach the lessons about manners and being considerate, I suppose (buy a pair of headphones for her, limit her volume so she doesn't suffer early onset hearing loss).
No guarantee she won't turn out to be obnoxious as an adult, but that's the genetic lottery, I'm afraid.
On the western side in a similar vein you've got hyper pop coming up from 100 gecs and laura les and what not. This kind of sound, hypertuned and almost as incomrehensible, sounds better to me. You do still get a vein of emotion. I love this sound.
In Japan, the term is "denpa" (電波ソング). Denpa music is intentionally strange as it is catchy, and hypnotic as it is awkward. There are many producers creating high-BPM electronic vocaloid music that is chaotic for effect. It is a bit more twee than the western sounds, as you mentioned, but it can be quite enjoyable if you're in the right mood.
More on denpa music: https://en.wikipedia.org/wiki/Denpa_song
Nanahira playlist, an example of a vocaloid character: https://www.youtube.com/watch?v=NHIyvhJadXM
Explaining Vocaloid in 3 minutes: https://www.youtube.com/watch?v=GODXMGAMpVc
Also, I think you'd enjoy the Song Exploder podcast. If you haven't heard it already, check out the episode where 100 gecs break down how Money Machine was created:
In Japanese, there is no distinction between syllable-final [n] and syllable-final [m]. But in English there is. Traditional romanizations of Japanese will transcribe this as "dempa", for the obvious reasons that (a) that is what the Japanese spelling says; and (b) that is also how the word is pronounced.
I often see English speakers get very confused over exotic modern transcriptions such as "denba" or "senpai", believing there must be a reason they are written that way. But I'm not sure what that reason is supposed to be.
Attempting to approximate pronunciation is a valid theory of transcription, but one which also ought to prescribe that 電気(でんき) be transcribed as dengki; English is not much less discerning of syllable-final [n] vs [ŋ] as it is vs [m]. This is not a position I've ever seen anyone defend in earnest, though.
(Romanization for anglophone is a bit of a lost cause anyway, since we're going to fuck up the vowels no matter what you do.)
That is blatantly incorrect. English converts syllable-final [n] to [ŋ] when followed by a velar exactly the same way Japanese does, and English spelling reflects that. Consider the English words "think", "clunky", or "handkerchief".
How would that suggest that it's reasonable to spell the Japanese word "dempa" as "denpa"?
For demonstrating lack of assimilation of /n/ to following bilabial, there are a couple distinct questions you might ask. It's very frequent for people to preserve the tongue gesture associated with /n/, because a bilabial stop doesn't use the tongue and so [n] is easily coarticulated. But that turns into /mp/ or /mb/ over time because the difference is not easy to hear. In contrast, for a word such as "impossible" where this process completed many hundreds of years ago, the tongue is not used at all in the pronunciation of /mp/. This is a kind of lack of assimilation.
You can also see lack of assimilation in the very people who go to special efforts to pronounce [n] in Japanese words where that is inappropriate.
Note that the English and Japanese phenomena you're talking about are very distinct. This is a fact about the historical development of sounds in English (and Latin...) that doesn't apply to current English, where a sequence like /ng/ will often be preserved across word boundaries. ("One ghost"; this is the only context in which such a sequence can occur at all.[1]) English maintains a robust distinction between /n/ and /m/ and a weaker one between /ŋ/ and the other two.[2]
In contrast, Japanese ん assimilates to whatever follows it, and in the case that nothing follows it it may (rarely) be realized as nothing more than nasalization of the preceding vowel. Word boundaries are not relevant. Japanese does not have a phonemic syllable-final /n/ or /m/ (or /ŋ/). It has a single sound (usually indicated /N/ by specialists, apparently, due to even more weirdnesses that it involves) that gets realized differently in different contexts.
So again - what would justify representing the Japanese sound as "n" regardless of context in languages where, unlike in Japanese, the distinction between "n" and "m" is meaningful?
[1] You say that most would-be /np/s are already spelled "mp", but this is false - the words that are spelled "mp" changed long ago, and do not represent attempts by modern speakers to pronounce an /np/ sequence. They represent attempts to pronounce an /mp/ sequence.
[2] Why weaker? /ŋ/ doesn't have the status the other two do; it cannot begin a syllable. And it makes for a less than perfect contrast with /n/ and /m/ because it has a fairly pronounced effect on the vowel that precedes it, which makes drawing a clean contrast difficult.
Consider "inpainting", "unbiased", and (as suggested earlier) "government", each of which is a synchronically transparent /n/ across a morpheme boundary, yet a cursory survey of recorded English speech suggests that it's pretty common for these tongue gesture associated with /n/ to be absent—infamously, the second syllable of the last routinely loses its coda altogether. This occurs across a transparent morpheme boundary, even with affixes productive in the modern language, even in learned usage.
English does have a lot more wrenches to throw in this, like producing nuclear nasals in a range of situations and not always assimilating across prosodic word boundaries—heck, it probably goes both ways in an utterance like "in my main menu". Words spelled "mp" are reliably [mp] in the modern language, but it's not a simple case as "mp" spelling /mp/ read [mp] and "np" spelling /np/ read [np]; English phonotactics also coerces the nasal in /np/ to a bilabial realization.
> because a bilabial stop doesn't use the tongue and so [n] is easily coarticulated
That doesn't sound quite right—this assimilation surely wouldn't be nearly as globally prevalent as it actually is if that were true.
Try it. While you'd think from the descriptions that a bilabial stop shouldn't care where the tongue goes, I think you'll find it quite challenging to coarticulate [n] with [b]—tongue positioning at lower teeth is pretty obligatory—and much easier to sequence them or produce [mb].
Clearly you can see the unnaturalness of lack of assimilation to call the attempt to do so "special effort"! So of course, the typical anglophone is not going to try to realize [n.p], they'll just see the <np> and read [mp] because that's what they would with any other internal /np/.
> So again - what would justify representing the Japanese sound as "n" regardless of context in languages where, unlike in Japanese, the distinction between "n" and "m" is meaningful?
Now, this gets to an entirely different issue: the purpose of the transcription. You seem convinced that the main goal of romanization is to provide a pronunciation guide for anglophones. But in the context of discussing a niche musical genre on the internet, that's not necessarily a high priority in the first place; you might care more about, say, searchability: we're looking for https://en.wikipedia.org/wiki/Denpa, not https://www.worldbank.org/en/programs/debt-toolkit/dempa.
And in a wider context, the principal users of romaji Japanese aren't anglophones; they're Japanese-speakers who for some or other reason need need to coerce Japanese text into an ~ASCII-subset representation, targeting primarily computer systems with that sort of limitation (most common case being keyboards via IME, hold that thought) and secondarily other people who can read Japanese; and naturally they make the distinctions Japanese makes and largely don't make the distinctions Japanese doesn't make. So unless backed by a marketing department, they tend to produce n (or nn as needed) for ん, because they have a tenuous grasp on how anglos spell [mp] in the first place and でmぱ is garbage that their IME won't convert into the right word, so why type that?
(This is also why pinyin can be the way it is, yet their IMEs have routinely have modes to ignore s-sh/n-ng/n-l distinctions.)
> That doesn't sound quite right—this assimilation surely wouldn't be nearly as globally prevalent as it actually is if that were true.
> Try it.
You know, I mentioned a specific theory here that you've completely ignored. The coarticulation is easy. But it is difficult for a listener to tell the difference between coarticulated [nb] and [mb]. If you're willing to let multiple generations pass, this means that /nb/ will become /mb/ regardless of how easy it is to pronounce.
You will also note that this theory of what's happening mostly cannot be disproved by recordings, which you appear to want to do. You'd want an X-ray or MRI study, something which shows you what the tongue is doing.
> I think you'll find it quite challenging to coarticulate [n] with [b]—tongue positioning at lower teeth is pretty obligatory
This is just obviously false. You have no problems producing [b] with your tongue positioned however you like. You can position it for [t], you can position it for [tʃ], you can position it for [k]. And of the three coarticulations I just mentioned, all of them are well attested, though only the middle one is attested in English ("pshaw", a scoffing sound).
> and much easier to sequence them
This is worthy of comment; there is a linguistic concept called "coarticulation", but all cases of coarticulated consonants seem to have a conventional sequence associated with them. I have no real knowledge or opinion on how real the conventional sequencing is, or how much sequencing is allowed before you stop calling the sounds coarticulated. I suspect that indeed it is easier to sequence two events than to coordinate them to occur at exactly the same time; this is true for all types of events, not just language-related ones. I don't think that the linguistic concept requires absolute synchronization of particular points in time; my understanding is that producing any given phoneme requires some motion and therefore takes place over a nonzero span of time, and "coarticulated" consonants are those for which the durations overlap, not necessarily those for which the durations perfectly coincide.
But I will note that while sequencing of /nb/ is obviously necessary in a way that is not true for /pt/, since /n/ must have nasal airflow and /b/ must not, there is no reason for "coarticulation" of /nb/ to be more difficult than it is in the attested coarticulation /tm/ (exactly the as /np/ for our purposes; /tm/ also features a voicing difference between /t/ and /m/).
> Consider "inpainting", "unbiased", and (as suggested earlier) "government", each of which is a synchronically transparent /n/ across a morpheme boundary
I don't think "government" is a valid example, and you should stop trying to lean on it. In my view, the pronunciation of "government" has as much to do with the morphemes suggested by its spelling as the pronunciation of "comfortable" does with the morphemes suggested by its spelling.
I have no problem with "unbiased"; that's a great example of what we're talking about.
> Clearly you can see the unnaturalness of lack of assimilation to call the attempt to do so "special effort"!
I don't agree with this. I claim that it is common for Anglophones pronouncing "unbiased" to make contact between the tip of their tongue and their alveolar ridge while they pass over the /n/ in the word. (And here, we're on firm ground saying that the internal phoneme is /n/ and not /m/, since it's part of a productive prefix un-.) I further believe that they make no special effort to do so. They may or may not allow a longer duration of nasal murmur than they do in other contexts, to make the /n/ clear; doing this would constitute a special effort. I believe that some speakers will do this and some won't bother. Of those who do, only a small amount of effort will be given to the task.
But the case of English speakers attempting to pronounce Japanese is different. They will go to great lengths to demonstrate that they want to comply with the bizarre textual representation they see. They are happy to produce highly unnatural speech in order to do so. (Which isn't really a problem; they don't really have an alternative to producing unnatural-sounding speech in early attempts to pronounce a foreign language. But this is something they shouldn't encounter problems with.)
> And in a wider context, the principal users of romaji Japanese aren't anglophones; they're Japanese-speakers who for some or other reason need need to coerce Japanese text into an ~ASCII-subset representation, targeting primarily computer systems with that sort of limitation (most common case being keyboards via IME, hold that thought)
> (This is also why pinyin can be the way it is, yet their IMEs have routinely have modes to ignore s-sh/n-ng/n-l distinctions.)
This isn't a flattering comparison for the all-n Japanese transcription system. The pinyin for 吕 is lü. Chinese people don't use German keyboards, which makes the pinyin impossible to type. So where ü contrasts with u, pinyin input methods require you to input V. And Chinese people have responded to this by adopting v-based spellings; it is common to see pseudo-pinyin like "lv" where that pinyin has been generated by an ordinary Chinese person for their own purposes, such as a sign over their business or an online username.
But the letter V is formally not a part of pinyin at all, which means that text generated by the government never uses it and neither do instructional texts.
It is true that this situation is the reverse of the one we're discussing - the Chinese are making a distinction that is required by their language but forbidden by their keyboard, and the fact that they are aware of the distinction makes it easy for them to know what to do. The Japanese are failing to make a distinction that doesn't exist in their language but does exist on their keyboard; this is precisely parallel to the pinyin IME settings you note that will allow the user to ignore phonemic distinctions that they don't make. Again we see that the system maintains the distinction and it's the job of the input method to interpret what the user wants to say.
Chinese IMEs also offer a "double pinyin" input method, in which you type one letter to indicate the onset of a syllable and a second letter to indicate the rime. All syllables are two input-letters long; this model matches the traditional Chinese view of their own phonology. You could just as easily base your system of English transcription on this: instead of "Xi Jinping", 习近平's name would be "Xi Jnp;". Instead of "Sun Yat-sen", we'd talk about "Sp Yixm".
That's what it looks like when you base spelling on what it's convenient for foreigners to type as an intermediate input to their own, different spelling. (As is the case with Japanese input methods.) There are zero people who believe it's a good idea. It's not a better idea in the Japanese case.
ななひら (nanahira) is probably the most well known denpa artist, but she mostly sings normal songs now I think (and has a lovely voice doing so).
ココ is my favourite denpa artist https://youtu.be/2wl8Ofce8TE
To me the most interesting part to vocaloid is the ability for a sole producer to make a complete song without any external help. The vocal parts have always been a barrier, and while emotionless and still lacking in some areas, vocaloids are “good enough” to support a well produced song.
We’ve seen creators rise through the ranks through vocaloid, get experience and exposure, to then move to full professional production with a staff and an actual singer (who’s voice will also be heavily processed, but they have a ton of tuning experience at that point)
I also agree with the parent comment that some creators do benefit from the “mechanical” part. Throwing more links, Giga works with both singers and vocaloids and is pretty good at extracting the best of boths: https://youtube.com/c/GigaVideos
An early example of this was the debut album of Boston which was mostly recorded in Scholz's basement with him on every instrument except drums, then the tapes were mailed to LA for Delp to record vocals. I think it's rather funny in particular that Rock and Roll Band was written and mostly recorded before the band even existed.
The only conceivable surprise is a crude chromatic key change to the minor version of the raised mediant.
You'd think the precision of those dynamic envelopes and timbral games would push the artist to venture out and explore that mediant relationship to create quicker and more jarring harmonic progressions and modulations. But no-- it turns out to be less inventive than the mediant chains emanating from, say, Joni Mitchell and her acoustic guitar over fifty years ago:
https://www.youtube.com/watch?v=3q2jiRUVLgI
(I find some of the lyrics apt, too.)
Compared to the cookie-cutter harmony and melody of the music you linked, even Mitchell's augmented triad in the melody at the end of the chorus sounds like the musical equivalent of solving fast homomorphic encryption.
It's the the audio tech that is on display in the music you linked, so every other musical consideration shifts to the background to illuminate that tech. I get that. But holy shit why does that baseline have to be stuck in the fucking 1650s? While I love the "electrified Vivaldi" hack that is heavy metal from the late 70s/early 80s (Master of Puppets et al), I question whether we really need more than one musical genre based on that parlor trick.
It would be like every stand up comedian ending their set with increasingly theatrical pyrotechnic pull-my-finger jokes. I could laugh my ass off at the absurdity for a year, maybe two. But forever?
My (possibly wrong) impression of your comment is that you seem to have made the mistake of associating complexity with quality in music which is extremely common in those who’ve just started looking into music theory.
Most music needs only the smallest dash of novelty to achieve the perfect mix of the new and familiar to its target audience. If you start attempting to evaluate popular music on what about it is inventive or new, you’re likely to find yourself unable to appreciate most of what people are enjoying and cut yourself off from loving a broad spectrum of musical expression.
You might also find yourself unable to express why you enjoy the music you do like in a way that doesn’t come across as if you’re arguing an objective scientific point——an approach which might undercut your argument by making you unintentionally come across as someone who has just learned a lot of fancy theory jargon and is eager for an excuse to wield it.
The track you reference sounds like chipmunks sped up 2x; it's not unpleasant to listen to, and fun, but I feel it could be made just like that (record at 80bpm, high pass filter, maybe transpose 1 octave, and speed up to 180), no "AI" involved.
It's a synthesizer. It's an alternative to human singers. I can imagine someone seeing a digital piano for the first time. "I'm not sure what it even does. I could just use an acoustic piano. It sounds the same."
This one if an official track for a popular vocaloid rythm game.
Also, at this point the “chipmunk” sound is part of the brand and will be kept to some extent for tracks labelled as vocaloids (it’s kind of a market on its own)
There's also plenty of music directly derivative of the vocaloid scene that maintains a similar aesthetic with 'organic' vocalists and dispenses with some of the awkwardness of vocaloid-oriented compositions. Example: https://www.youtube.com/watch?v=hjJMIWyl_l4
the audio is generated from a voicebank that is a database of prepared phonemes recorded from a voice actor. some packages come with multiple variants of voicebanks, like you could have a "soft" voice and a "vivid" voice.
Human Japanese singers, especially women, tend to operate in a higher octave range than what is common in the west. It's slightly culturally insensitive to take shots at vocal pitch when talking about J-Pop. Pitch is largely a social/cultural construct, and Japan generally leans into the idea of higher pitch -> polite or cute and lower pitch -> aggressive or rude. (e.g. you raise your pitch when talking to your boss, and drop it to express your disgust with someone.) Just putting that out there, not trying to be accusatory or anything. It's just always good to keep in mind that western cultural norms are hardly universal.
For the record, I was responding to the gp saying
> ... pushing the boundaries of pop music in a way that wouldn't be possible with a real singer
=> I felt it was possible to do what the example does by singing slowly and speeding it up afterwards.
https://www.youtube.com/watch?v=Y2k8EOBL75o
Not great, not terrible.
It's definitely not my thing, but Tsumiki's use of vocals is interesting and well executed.
He also has some funny parodical bits he does, like rapping about having a lot of money/jewels/etc and then the vocaloid characters rap about having a lot of RAM.
Nope, certainly stays disagreeable to me. I wonder what makes people enjoy weird stuff in so many different ways. I might not like this, but I enjoy white noise artist Merzbow [0] or breakcore from Drumcorps [1]
As a small form of resistance to the surveillance state I partake in, I have taught kids to ask any nearby personal assistants to play woodpecker #2 and they find it hilarious.