Google Translate “Get well [Swedish firstname]” translates to “fuck you”
translate.google.se
translate.google.se
If I used Google Translate to talk to a shopkeeper, it would be roughly equivalent to saying "Hey, little buddy, how much for this?" as opposed to "Excuse me sir, what is the cost of this item?"
And this is all without considering all the weird mistranslations you can get because Korean is much more heavily context dependent than English. Korean speakers often leave out the subject or object if it can be understood from context (context that the translation tools are likely missing). So Google translate will insert pronouns (it, him, her...) to make the English flow better, but are not based on anything in the original Korean. So, if it guesses wrong, you could imagine the level of confusion that could ensue.
And then all the homonyms in Korean combined with the heavy context dependence makes for some weird translation. I once tried checking my Korean homework with Google translate, and before I knew it, I was drinking a car.
1) Make an android app and publish it
2) Write "Get well [firstname]" in sweden locale
3) Enjoy your ban because google uses google translate to look for inappropriate language in app descriptions
Deepl gets it right btw:
https://www.deepl.com/translator-mobile#sv/en/krya%20på%20di...
Very impressed with GPT-4 translation though -- especially the ability to steer it between "transliteration", "keep the meaning", "keep the tone", "use local idioms where appropriate", "explain different possible meanings/intentions", etc.
I've tried it when I had the time to compare (DeepL vs GPT4), and find them to be pretty equal.
But DeepL easily win on speed. 5 paragraphs would take just some seconds with DeepL and be almost 100% correct, while GPT4 would take almost a minute (sometimes more) while being about the same amount of correct.
> Very impressed with GPT-4 translation though -- especially the ability to steer it between "transliteration", "keep the meaning", "keep the tone", "use local idioms where appropriate", "explain different possible meanings/intentions", etc.
I've found that DeepL already does this well even thought it's not a LLM (as far as I know).
Yes. DeepL is very good! But with ChatGPT I can "tune" it more towards one way, whereas DeepL just only does whatever it does. DeepL has very very sensible defaults and the UI is great. But in fairness, DeepL basically will never just insert an appropriate idiom. Also GPT-3.5 is still worth comparing to DeepL as well.
- krya på dig Helga - take care of yourself Helga
- krya på dig Dave - screw you dave
- krya på dig Mary - come on Mary
- krya på dig Linnéa - brace yourself Linnéa
- krya på dig Mohammad - fuck you Mohammad
- krya på dig katt (cat) - fuck you cat
Björn is also bear in swedish.
Capitalization gives additional context in this case, if it were in the beginning of the sentence though, then one would hope it contains other clues as well
It’s a strange little circle of anomalies
krya på dig Cat - screw you Cat.
krya på dig balloon - get on you balloon.
krya på dig tacos - grab some tacos.
krya på dig applesauce - put on some applesauce.
krya på dig Ingrid - get over it Ingrid
- Krya på dig Björn - Get over it Björn
- Krya på dig Helga - Get over it Helga
- Krya på dig Dave - Get over it Dave
- Krya på dig Mary - Get over it Mary
- Krya på dig Linnéa - Come on Linnéa
- Krya på dig Mohammad - Get over it Mohammad
Sorry about language and the poor description; I seem to have an idea of the problem, but it's not my field and have no way of describing it in formal way.
https://www-eranda-jp.translate.goog/column/24550?_x_tr_sl=j...
They have the word "czarny" which means "black". "czarny kot" is "black cat", "on jest czarny" is "he is black". The latter is purely descriptive.
There is also the word "czarnuch" which is the VERY offensive word for Blacks, best translated by "negro".
Now, the last name "Czarnuch" is a normal last name, without any connotations to the color black (except probably in its etymology) and does not sound weird/offensive.
The translation of this capitalized word would naturally yield "Negro".
There might be a reference somewhere where it is used as a description of a person, but probably not in Croatian.
I did notice Google Translate hallucinates words, some very amusing, when I translate from Any -> Croatian (this happens automatically when reading Google Maps place reviews). There has been quite a lot of words that naturally map to Croatian but there's no text (outside of blogspam) that uses it on the Internet.
What's wrong at Google lately?
شعب يباد = (people are being exterminated)
To:
Iraqi People
input 'baiser' -> Google translates to 'kiss'
now add a female first name after 'baiser', 'kiss' will become 'fuck'
Can a Swedish speaker say if something similar is going here?
For those unfamiliar with LLM architecture: "tokens" are the smallest unit of lexical information available to the model. Common words often have their own token (e.g.: Every word in the phrase "The quick brown fox jumped over the lazy dog" has a dedicated token), but this is a coincidence of compression and not how the model understands language (e.g.: GPT-3 understands "defenestration" even though it's composed of 4 apparently unrelated tokens: "def", "en", "est", "ration").
The actual mechanism of understanding is in learned associations between tokens. In other words: the model understands the meaning of "def","en","est","ration" because it learns through training that this cluster of tokens has something to do with the literary concept of violently removing a human via window. When a model encounters unexpected arrangements of tokens ("en","ration","est","def"), it behaves much like a human might: it infers the meaning through context or otherwise voices confusion (e.g.: "I'm sorry, what's 'enrationestdef'?"). This is distinctly different from what the model does when it encounters a completely alien form of stimulation like the aforementioned "Glitch Tokens".
At the risk of anthropomorphizing, try imagining if you were having a conversation with a fellow human and they uttered the following sentence "Hey, did you catch the [MODEM NOISES]?". You've probably never before heard a human vocalize a 2400Hz tone during casual conversation -- much like GPT-3 has never before encountered the token "SolidGoldMagicarp". Not only is the stimulus unintelligble, it exists completely beyond the perceived realm of possible stimulus.
This is pretty analagous to what we'd call "undefined behavior" in more traditional programming. The model still has a strong preference for producing a convincingly human response, yet it doesn't have any pathways set up for categorizing the stimulus, so the model kind of just regurgitates a learned lowest-common-denominator response (insults are common).
This oddly aggressive stock response is interesting, because it's actually the exact same type of behavior that was coded into one of the first chatbots to (tenuously) pass a Turing test. I'm of course referring to the "MGonz" chatbot created in 1989[2]. The MGonz chatbot never truly engaged in conversation -- rather, it continuously piled on invective after invective whilst criticizing the human's intelligence and sex life. People seem predisposed to interpreting aggression as human, even when the underlying language is, at best, barely coherent.
[1]: https://www.youtube.com/watch?v=WO2X3oZEJOA [2]: https://timharford.com/2022/04/what-an-abusive-chatbot-teach...
[1] https://www.albany.edu/news/releases/2002/june2002/gallupstu...
I'm just saying this comment doesn't seem to be related to OP.
So no I was not offended, lol.