Animal Crossing’s fake language sounds different in Japanese
polygon.com
polygon.com
The problem is that English is a phonetically inconsistent language, with a massive number of rules required to even begin to approximate the mapping from text to phonemes (and zillions of exceptions). So this kind of really dumb TTS not intended to be actually intelligible doesn't work at all in English. And so it sounds like actual nonsense.
Or maybe it was and they tweaked it well enough to work. It's been a while, my memory could be off. I don't know if it would have been prohibitive on the GameCube to have audio of every line (there were a lot), but I wouldn't put it past them to have done so.
Besides, can you imagine voice actors dubbing this stuff in this kind of voice line by line? They'd go insane.
To me it does sound exactly like that too in English.
I'd wager it's more obvious in the Japanese version because Japanese is the exceptional language. There are only 44 syllables in Japanese (English has about 16,000) and one would probably still notice this in otherwise unintelligible distorted speech.
Spanish is about as phonetically consistent as Japanese. Most languages have simpler phonology than English. It's not about the number of possible syllables, it's about the rules to go from text to phonemes. The rules in English are immensely complicated and inconsistent. Meanwhile, I can accurately describe Spanish phonology, such that you'd be able to pronounce ~any Spanish word (English loanwords excluded) accurately including stress, in about one page. Written Japanese lacks pitch accent information, but otherwise works similarly.
(By the way, your stats are off; modern Japanese has about ~106 possible syllables (mora) by rough count).
Most TTS software out there will be better at English than any other language despite the more complex phonology. Then it's solely about distinguishing language characteristics and I think the amount of syllables would have an effect in that. Japanese has a lower information density than English (at the same speed) so with the same amount of distortion Japanese should be more recognizable.
Banjo-Kazooie [1] was the first game on the N64 to use what Animal Crossing terms "Animalese/Bebebese" [2], and their intention was never to build a TTS engine. The first Banjo was released in 1998, while Dobutsu no Mori (Animal Crossing) for N64 didn't come out until 2001.
Nintendo was definitely talking with Rareware at the time and they exchanged ideas and techniques on game engine design, platformer mechanics, etc. Interviews from the Rare side admitted this (I'll need to dig up some sources to include here).
I'm curious if Nintendo picked up the Animal Crossing Bebebese voices directly from Banjo-Kazooie.
[1] https://www.youtube.com/watch?v=9ZE5A3DbHDk (Actually a video of the year 2000 sequel, Banjo-Tooie, but this is a better example of the same voice engine using various voices.)
[2] https://animalcrossing.fandom.com/wiki/Language#Bebebese
The distortion effects were presumably added so people overseas wouldn't notice how jarring it is. It's such a cool effect it feels like something that should have been in the original release.
https://www.youtube.com/watch?v=-VsmF9m_Nt8&feature=emb_titl...
I get this in a lot of things. Human brains love finding patterns and everyone does this to an extent. I think mine does it so much more. For example if I ever see numbers anywhere I'm compulsively adding them to see if any nicer looking numbers result.
I wouldn't be surprised if this inclination to hear jibberish and try to parse it into language is a me thing.
[1] Taken from wikipedia
As you move the cursor over each letter, the character speaks out a phonetic sound similar to the letter you're on
It's called speaking "yogurt".
It was originally used because young French people wanted to sing the English songs that they heard, but didn't know the language, so they would make up sounds that looked like English.
https://en.wikipedia.org/wiki/Barbarian
> The Greeks used the term barbarian for all non-Greek-speaking peoples, including the Egyptians, Persians, Medes and Phoenicians, emphasizing their otherness. According to Greek writers, this was because the language they spoke sounded to Greeks like gibberish represented by the sounds "bar..bar..;"
As a native English speaker (UK), I'll be pedantic. That's also how Americans sound to me.
That would bring about a new boom (pun) of creativity by allowing indie devs to write complex stories with spoken dialogue without having to worry about hiring actors and immutable recording sessions.
I imagine it's mostly a licensing/cost thing (since a "voice" for a speech synthesis engine still requires hiring an actor and doing a recording session).
That's what I'm talking about. Surely we can do away with that if we try, just as we don't need real people to build 3D models -from- if we don't want them.
Just like how we invented an “automatic programming” system where the computer will do programming for you, and it turns out that once we’ve made automated programming we have more programmers and not fewer. The tools for making 3D models are getting better and easier to use, and as photogrammetry is being used more and more, we see larger teams of modelers, not smaller.
A 3D scanner is an artist’s tool. By using a 3D scanner, you aren’t getting rid of artists, you are just changing how artists do their jobs. Vocal synthesizers and vocal transformers, similarly, aren’t making it possible to press a button and get reasonable sounding voice in your games if you don’t have a voice actor making it possible. If you aren’t convinced, then just look at soundtracks. You can press a button and your iMac will spit out the sounds of the BBC Symphony Orchestra violin section. In spite of this, getting a symphonic score is still expensive. It’s expensive enough that TV shows (with sizable budgets) often skip out on the symphonic score and do something cheaper. We haven’t gotten rid of musicians, it’s just that musicians are much more likely to have computers.
There is, in theory, nothing stopping you from buying like $200 in software and making a symphony orchestra right now. The problem is that you have no idea how to write a symphony orchestra. For the same reason, if you have an iPhone 12 or something similar then you can start using the LIDAR features and making a 3D model using photogrammetry in moments—except for the fact that you have no idea how to make a 3D model.
I think there’s a trap that people fall into, thinking that technology is just around the corner that will get rid of job X, Y, or Z. Often what you end up with is MORE people doing job X, Y, and Z, it’s just that they use computers to do it, and have a different skill set.
It’s actually pretty rare the company that can do its own thing and survive at it.
Github: https://github.com/equalo-official/animalese-generator
Video Explainer: https://www.youtube.com/watch?v=RYnI_ZLj5ys
Edit: In the english version of Breath of the Wild.
https://www.google.com/amp/s/www.theverge.com/platform/amp/2...