Help! Is This Arabic?
isthisarabic.com
isthisarabic.com
[1]: https://developer.mozilla.org/en-US/docs/Web/CSS/CSS_Logical...
I've usually found that the arguments for these properties revolve around, "well, what if I decide to use traditional Mongolian in the future?," which seems like the biggest case of YAGNI I can think of. I suspect their popularity is owed to tutorial authors showing off their CSS prowess.
(I'm also not convinced of the need to flip sizing values for RTL/LTR, but that's at least useful)
Forget my question on your users getting utility out of this. Have you ever seen a site, any site, that supports switching to top-to-bottom writing? Against what future are we proofing – a new language arising?
Other examples (e.g. vertical Japanese) aren't too hard to find: https://nishinokensetsukogyo.co.jp https://ok-maru.jp etc.
Almost all extant vertically-written languages are more commonly seen horizontally on the web. Vertical Japanese on the web (from my American understanding) is a design choice, and is notably absent from mainstays like https://www.yahoo.co.jp/
That is to say, it is unlikely that a site's language switcher would opt for top-to-bottom writing.
Edit: the Mongolian president's site does have an English version! ...but it's a completely different site :( https://president.mn/en/. Still, this is the closest to a use case for logical properties I've seen, so kudos.
Using logical properties would make it easier to have a common base of CSS that controls spacing, sizes, etc., that can be used by both language versions, rather than maintaining them completely separately.
> the arguments for these properties revolve around, "well, what if I decide to use traditional Mongolian in the future?," which seems like the biggest case of YAGNI I can think of
...And even this one close example doesn't use them!
Chinese/Korean/Japanese/Vietnamese originally, but computers have made horizontal writing more common on the internet. Modern Vietnamese is also written in a variation of the Latin script, of course.
Mongolian still uses the vertical Mongolian script, though Cyrillic script m has also been introduced in Mongolia during Soviet times and seems to be common inside Mongolia. However, the Mongolian government seems to be moving back to using the original Mongolian script. Furthermore, the Mongolian people inside China never took up the Cyrillic writing system.
Websites can and do use the vertical script: https://president.mn/mng/ http://khumuunbichig.montsame.mn/index.php?home&readnews=572
Google uses the Mongolian Cyrillic alphabet (https://www.google.com/?hl=mn) so even companies that seem to have a page in every single language don't seem to bother with supporting the original script. This is probably because the Mongolian speakers inside China can't make much use of Google anyway.
Funnily enough, Mongolian is one of the few known scripts not only written vertically but also left to right, unlike other Asian vertical scripts (which go right to left).
Contrast this with a form in English and Spanish, where you need to put the text all on the left and decide which language goes on top.
https://www.w3schools.com/tags/att_global_dir.asp
https://caniuse.com/mdn-html_global_attributes_dir
Edit: oh use case? Imagine a PDF form you'd like someone to fill out, or a page that may get printed. You could create two forms and let the person filling out the form pick their language, but the person processing the form may prefer to read the form in the other language.
Anything that doesn't require input - particularly printed ones where you can't just swap with a button. Signs, menus, etc.
What it really needs is a simple reference input string, examples of how it gets broken, and what to do to fix them. The middle case, where the sentence is correctly rendered RTL but the individual words are LTR (breaking the ligatures), is particularly common and insidious because it looks plausible to non-Arabic speakers.
> In general, the letter combination ال should be common.
So if your text has more than a few words, you should be able to look though your text and see that somewhere.
I can't read Arabic but I can recognize that pattern. I went to https://www.bbc.com/arabic and could find numerous occurrences.
It's a bit like saying "if you have a paragraph in 'English' and it has no e's in it, it's probably not English."
This is true. You should also be able to see the word "the" in several places in English text, and if it is rendered as "ehT" or something like that, the ligature code may have a bug -- similar with reversing the ال (read as "al") pattern.
> I can't read Arabic but I can recognize that pattern. I went to https://www.bbc.com/arabic and could find numerous occurrences.
Really? I only ever learned to read just enough Arabic to read parts of the Quor'an and am by no means fluent, but I couldn't see any mistakes on that website myself.
For the exception, see
Edit: obviously it is English but it's definitely not correct for a company website.
There's more than just corporate websites, and frankly, if a company of any meaningful size offers content in Arabic, I'd expect them to hire someone for that. Even part-time or freelance.
It isn't; "to maintain afloat" is not grammatical.
You could replace that with "stay afloat" and it would fix the grammatical error without introducing an E.
There's a similar unforced error in referring to a "group" of kookaburras rather than a "flock".
"Individuals distinguish from distinct sorts" is gibberish. I cannot tell what it's supposed to mean.
"All flap wings as a way to [stay] afloat" is, at best, very awkward; fluent English would require "flap their wings", but that would introduce an E.
"Flapping about a body of liquid" is a very odd thing to say unless the body of liquid happens to be suspended in midair, since midair is the only location where you can find birds flapping.
GP also said "probably". That's how heuristics work.
This is a bit like writing unit tests, debugging code, and so on.
It was a really smart way of saying you don't need to know anything at all to pattern match on this a spot a very common problem.
Also, if you haven't checked out the youtube video on that page I highly recommend it. It gives a great concise summary of the issue, and it's impressive how much I was able to learn to visually parse Arabic script with only a couple of mins into the video.
Thanks to this training, I can now identify that numerous Arabic strings on that site are backwards. He wasn't joking, IJ is everywhere.
Who can’t tell these two character sequences from each other? Genuine question.
I think the main point the author is making is, if you are including Arabic script somewhere, take a bit of time to either do it right or hire someone to do it right for you.
I really like this idea. Just have a standard set of strings covering all edge cases (even the sprawling labyrinth that is bidi) with a visual reference that shows how the correct rendering of each string would look like. Each entry would also have a description of the problem and suggested solutions.
Unlike the solutions in OP, this one is pragmatic and is actually actionable for the vast majority developers. I'm kinda surprised that something like this doesn't already exist given the substantial amount of material and visual examples already available that covers the bidi algorithm.
- https://www.w3.org/International/articles/inline-bidi-markup... - https://www.w3.org/International/articles/inline-bidi-markup... - https://www.w3.org/International/articles/inline-bidi-markup...
I don’t think that is an unreasonable assumption on the reader
I don't think it does. Just hire someone.
I guess really, the point Rami tries to make is that not a single person who reads Arabic was involved in the video game/advertising/website/etc. The errors are often so basic that a child could point them out.
It would be like if someone wrote English without any spaces between words. It's so painfully obviously wrong.
What the site is about and where hiring someone makes sense is anything "big budget", especially if your target market includes Arabic-speaking or Arabic-adjacent countries.
Who can't tell these apart? I know literally no Arabic - these characters look very much like latin ones I and J, and it's just an order thing.
It seems like an excellent quick test to me to see if there's ordering problems.
This article can't possibly be scoped for someone like that.
There's a side note here, that dyslexia can become apparent in very different written languages.[0]
In which case don't try and handle multilingual text pay someone else, even if that's on fiverr.
[0] https://blogs.scientificamerican.com/observations/its-all-ch...
But isn't that exactly what it's just telling you? The order is important and if you see this very simple pattern it's wrong.
If you read any latin character based language then you surely must be OK with the idea that glyphs can be distinct? Are there many people who exclusively read languages where the order of symbols is not important?
> There's a side note here, that dyslexia can become apparent in very different written languages.[0]
If the point is that dyslexia means some people can't see the difference then that's fair, I'd not come at it from that angle. I don't see any surprising pre-assumed knowledge in this IJ/JI distinction however.
That's a really weird comment. Just Ctrl+F ل ا in your supposedly Arabic text?
Most Arabic is written without small vowels (harakat). You could have script that is justified right-to-left and with letters correctly connected and it still be gibberish. And many of those 'two billion people' would be none the wiser.
I know Muslims who would be challenged to speak more than 3 words of Arabic, as well as Arabic speakers who are Christians or atheists.
Source: my grand aunts saying "arapreme" where "ora pro nobis" would go (the former kinda sounds like an Italian word). Also probably related, "hocus pocus" is the "magic" expression "hoc est corpus" ("this is the body [of Christ]").
I don't know if it's true but that does not seem unreasonable.
> that does not seem unreasonable.
My point is that what seems reasonable to someone that doesn't know anything about the subject means nothing, and could even end up being somewhat offensive.
Try to reverse the logic and realize how absurd it would look to you: say a guy in middle east wanted to roughly estimate the number of Latin speakers and made the assumption that it should roughly be the same as Catholics.
Wikipedia says there's ~2 billion muslims, and 400 millions Arabic-dialect speakers (native and non native). So on average 20% of muslims are able to understand basic arabic, I expect that not even half of those would be able to read and understand classical Arabic such as written in Quran.
With few exceptions, Islamic revelations do not state which Quranic verses or hadith have been abrogated, and Muslim exegetes and jurists have disagreed over which and how many hadith and verses of the Quran are recognized as abrogated,[6][7] with estimates varying from less than ten to over 500.[8][9]
Not a weird claim at all. Even without harakat, it is obvious when Arabic script is messed up (via alignment or letter reversal) to anyone who can read the script. It just makes whatever you are reading look janky and unprofessional.
[1] https://www.babbel.com/en/magazine/how-many-people-speak-ara...
FWIW, the usual estimate for worldwide adherents to Islam is ~1 billion.
Though if your numbers are correct, 2 billion sounds way too high
[1] https://www.pewresearch.org/religion/2011/01/27/the-future-o...
Now I know:
> 3. In general, the letter combination ال should be common. The combination ل ا cannot occur in Arabic script, as those characters should be connected.
I still cannot read a jot of Arabic and know enough to identify some cases where the dev/writer has gotten it wrong.
Alternatively, it may be referring to people who read a language written in Arabic script, but not necessarily the Arabic language. Languages that use some variety of Arabic script include Persian (Farsi and Dari), Urdu, Pashto, Western Punjabi, Uyghur and some other languages. Some of those languages use vowel diacritics, particularly Uyghur.
My understanding, as a non-Muslim who doesn't speak Arabic, is that the standard Arabic is pretty conservative such that MSA, as would be commonly heard on TV or read in the news or other more formal settings, is not very far off the Arabic used in the Quran. So I think most native Arabic speakers would understand the Quran well, but may not be able to fluently speak or write it.
And that many Arabic learners learn MSA along with a dialect, so many second language speakers would probably be able to productively read the Quran as well.
To most Arab speakers, both Classical Arabic and MSA are simply referred to as Fusha.
There are also many non-speakers of Arabic who understand many verses and can pick up on the gist of many verses in the Qur'an.
It won't fool the people who can speak the language, but I think the website is just designed to educate people so that it doesn't look like complete nonsense.
Facebook is surprisingly an offender here. It's common to mix both Arabic and Latin, say, begin with بسم الله and then write the text, sometimes in english, and it'll throw off the alignment of the text completely. You get the Latin words right aligned or the Arabic one left aligned.
Edit: I'm actually quite surprised how well HN handles this.
HN /could/ make it better by setting css auto directionality for all <p>s, but that would be antithetical to its goals as a English-written forum.
Not really, because
>there aren't two billion people with functional Arabic literacy.
isn’t something that the author claimed. “To some degree” is a phrase that explicitly states that the author isn’t talking about full functional literacy.
It would be a super weird claim if “some degree” and “full functional literacy” meant the same thing, but they don’t! You would almost have to intentionally ignore the meaning of the words the author used and invent a nonexistent overlap of meaning to become confused on this point!
They should at least be close if the author isn't trying to pump up the numbers in a misleading way. I definitely assumed that number would be close to the literate number, and not including people who can recognize a tiny fraction. Hell, I can read Arabic "to some degree" if we're being completely literal, but I think including me in a persuasive claim about people that use Arabic would not be appropriate at all.
> You would almost have to
No need to be rude.
I’m genuinely confused here. The author was pretty much crystal clear about how he defined the size of the group that he was talking about, I do not understand how a person could be confused let alone feel the need to accuse him of being intentionally misleading.
What exactly is the nefarious goal that the author was trying to sneak past you with his clever trick of speaking in plain english?
It's plain English, but I picked the wrong measure to support my claim.
Let me try a different analogy. "It's important for caterers in the US to provide a gluten-free meal choice. After all, the population is 332 million!" Without knowing the incidence of gluten sensitivity, it's a borderline-misleading statistic.
I also agree that the existence of a subset that wasn’t referred to at all in the article is completely irrelevant to the topic at hand!
And the parent comment was not talking about fluent people when writing "can read".
Posters have been able to pinpoint the number of fluent speakers, can you give me a ballpark to how many people matter and how many don’t matter?
This article about rendering text appropriately has taken such a fun turn into sorting folks into groups that “matter” and “don’t matter”?
If being rigid about these numbers is so important, how many people that don’t matter today might matter next year? How many people are learning arabic script? How many might want to look up something written in arabic without it being rendered in absolute nonsense?
I answered that in my first post! I can read a tiny tiny bit of Arabic. That puts me into the literal "some degree" group, but I am also definitely in the not-mattering group, because rendering mistakes with Arabic will not cause me any problems with reading.
> Posters have been able to pinpoint the number of fluent speakers, can you give me a ballpark to how many people matter and how many don’t matter?
Have they? But I don't have numbers, I'm just saying that "fluent" is too small and "some degree" is too big.
> This article about rendering text appropriately has taken such a fun turn into sorting folks into groups that “matter” and “don’t matter”?
Are you offended that I classify myself as not mattering in this very specific context? You don't have to make it sound like I'm saying people don't matter in general, jeez.
> If being rigid about these numbers is so important
If a number is worth busting out to make a point, it's worth being correct.
> how many people that don’t matter today might matter next year? How many people are learning arabic script?
What's your point? If the number changes, then use the new number. Don't use a wrong number because it might change later. Or if you have an expected future number, label it as such.
> How many might want to look up something written in arabic without it being rendered in absolute nonsense?
A lot of those people aren't even inside the "some degree" group, so now you're making a different argument. I'd rather not start any new tangents at this point, if you don't mind.
My point is that you’re trying to use some sort of odd pedantic mark trick to shift the conversation from your experience of “There is a group that I personally don’t care about” to “Math dictates that this is not actually a problem worth addressing.”
Your position that the important takeaway here is actually the importance of scrutinizing pointless minutiae rather than text rendering being fundamentally broken isn’t empirically based. Your entire argument is “look at how clever I am!”, which is fundamentally off-topic when talking about rendering text properly.
Like lol, how are people supposed to learn the script if their examples are all messed up? As a maths genious surely you could see the issue with how “impacted people” is somewhere between “fluent people” and “fluent people plus an unknown number of others.” What hard number did you land at when adding unknown variable x to the number of fluent speakers you googled?
I think this is a really uncharitable read of this conversation. This thread has been about the veracity and the relevance of the author's claim that "two billion people can read Arabic to some degree".
I don't think anyone is trying to refute the author's conclusion that Arabic text rendering is important. I also don't think anyone is trying to show off how clever they are.
Personally, I agree with the author's conclusion, and I thought the post was really neat! But I also think the 2 billion statistic weakened their argument -- it's better to omit a statistic than include the wrong one.
This is not really true. You tried to center your conclusion that your math was better than the author’s math while distracting from the topic of rendering text properly.
This thread has been about you insisting that people listen to your math and not discuss rendering text properly. lol this thread has been about how clever you are, _not_ rendering text properly.
And in these comments I'm assuming that the author has exactly the right number for the group they cited. Because it's really not about math. I have done no calculations and trust the number given. I just think they're citing the wrong statistic. That's why I'm also uninterested in the factors you mentioned that might influence the number up or down. The actual number doesn't matter for this criticism: even if the number in the article happens to match the right statistic, they're still citing the wrong statistic.
Literacy is ability to understand a writing system. I am literate in the Latin alphabet.
Fluency is ability to understand a language. I am fluent in English.
Neither implies the other. You can be fluent but illiterate (the default until modern universal education), and literate but not fluent (I am literate in the Latin alphabet but I am not fluent in Italian).
The claim that there's ~2 billion people who are literate in Arabic script and will laugh at you if you get it wrong, is more or less true. It's of course referring to the large number of people who can read from the Quran in Arabic but without understanding all the words.
The number of people who can vocalise Arabic without the vowel marks (e.g. a newspaper) is considerably less. But you don't need to be able to do that to notice any of the errors in the article.
That might be a small group though and probably outweighed by all the non-Muslim Arabic readers (for instance I work with 2 Egyptians, one is Coptic and one is ex-Muslim and both can easily read Arabic).
Is it misleading? You don't have to be anything like fluent to realise when text rendering is broken. The quantity that is actually relevant to the discussion is the number of people who, when they look at your UI, will know that the arabic text rendering is broken; not the number of people who are fluent in arabic.
So even though I know a single digit number of words, and you can count me in that two billion, nobody should care about getting it right on my behalf.
The people who can't read it, but who can see that it's broken, will form a lower opinion of your product. It's as if I went to a Polish website and the text was all right-aligned and in all caps. I can't read Polish at all, but I'd still form an opinion about the quality of the site.
If you screw up a language that has 0 readers, it matters far less.
The point of saying how many speakers there are was to increase the strength of that effect. Because of that, it's misleading if you pump up the number. For pumping it up to not matter, the number would have to not matter, and there wouldn't have been a reason to mention it in the first place.
You seem to be demanding that "people who know enough medical terminology and/or Latin to see through your fictional doctor" be nearly the same class as "people who are doctors," and what's more, implying that's some sort of deception.
> No need to be rude.
That just simply means you can't read arabic. You absolutely do not need to vowels to be able to read arabic, if you know the language, you know the vocabulary and know how a word should be pronounced. The vowels are there to aid in clarity. In most cases one single vowel could outright turn a word into its own antonym. And right now, 99% of arabic text is written without vowels apart from Quran and literary works.
I can muddle along somewhat in French, especially written French, but I have essentially no Polish despite working with more Polish people than French. Nevertheless, Polish is written in a Latin script (for about a thousand years) and so I "can read Polish to some degree". If you show me a Polish street address, and then some street signs, I can spot when the sign matches the address, because I understand that symbols which are slightly different just mean the same thing. If you give me Polish mirror writing, I know it's wrong - it's backwards, even though I don't understand it.
If I attempt this in China it won't work, because I don't understand the Han script, so I am not sure whether a symbol I'm seeing is the same symbol written more or less ornately or an entirely different symbol which just looks somewhat similar. Are the symbols backwards? Or maybe they're different symbols which just look backwards.
The layer beneath this by the way is recognition that intent exists. If you show me Chinese text I not only can't read it I'm not sure which symbols are "the same" and which are not, however I can immediately tell this is writing. The writer intended to convey meaning with these shapes, perhaps I can find somebody else to translate them for me. Whereas say, the pattern on my duvet is just a pretty pattern, it doesn't mean anything (yes, I have thought about this, no it isn't a secret messsage) and so I can't get that "translated".
No, reading and being able to decipher script is different. I can read portuguese, spanish or italian (even romanian) to some degree because the languages are close enough to each other but I can't read polish, basque, finish or magyar despite them using the same latin script
Just for fun:
Same symbol: 龙 龍
Completely different: 已 己
Photoshop since time immemorial had a Middle East edition (ME) which supported RTL and used to curb piracy
I don’t know how it’s with their new pseudo-sass but I think they integrated it into their main product
https://helpx.adobe.com/lv/photoshop/using/unified-text-engi...
http://news.bbc.co.uk/1/hi/7702913.stm
> When officials asked for the Welsh translation of a road sign, they thought the reply was what they needed.
> Unfortunately, the e-mail response to Swansea council said in Welsh: "I am not in the office at the moment. Send any work to be translated".
People often think text layout is "easy" because they don't consider how anything other than their common experience treats thing (the reality is there still isn't a good vertical text layout story on the web).
This is also demonstrated whenever people go "I'll handle text layout myself", because they think one key press = one character, and immediately break text entry for more than half the world.
English has cursive and I think most English speakers are at least aware of cursive. We have script that is designed for print as well, though, while Arabic script does not; it is always cursive.
For instance, the Simplified Arabic Alphabet was devised by Muhammad Shakeel as an alternative way to write Arabic. It is a non-cursive alphabetical script as opposed to the traditional cursive Arabic abjad. The letter shapes are based mostly on the early Arabic Jazm script. It is not connected to or inspired by Nasri Khattar's Unified Arabic script.
There were similar issues with typing JP/KR/CN glyphs on computers for a long time which thanks to technological progress has stopped being something that needed solving.
This led to an embarrassing mistake when Google OCRed older books that contained the word "suck."
The only kase in which "c" would be retained would be the "ch" formation, which will be dealt with later.
Year 2 might reform "w" spelling, so that "which" and "one" would take the same konsonant, wile Year 3 might well abolish "y" replasing it with "i" and iear 4 might fiks the "g/j" anomali wonse and for all.
Jenerally, then, the improvement would kontinue iear bai iear with iear 5 doing awai with useless double konsonants, and iears 6-12 or so modifaiing vowlz and the rimeining voist and unvoist konsonants.
Bai iear 15 or sou, it wud fainali bi posibl tu meik ius ov thi ridandant letez "c", "y" and "x" -- bai now jast a memori in the maindz ov ould doderez -- tu riplais "ch", "sh", and "th" rispektivli.
Fainali, xen, aafte sam 20 iers ov orxogrefkl riform, wi wud hev a lojikl, kohirnt speling in ius xrewawt xe Ingliy-spiking werld.
from history, sometimes attributed likely incorrectly to Mark Twain
Syriac script looks completely different to Arabic but has the same issues when rendered
But note that in some cases, eg https://en.wikipedia.org/wiki/Tajik_alphabet#Samples point number 2 won't be accurate. I'm not the best reader of Tajik written in (Perso-)Arabic script, but I don't believe "ال" appears in the samples there. (It does appear in the word "ALphabet" in the opening paragraph!)
Trying to read Farsi, it feels like I should know what's going on but am left with the feeling that I've forgotten all my Arabic. Then I'll see some of the bonus letters.
Wait till you discover Urdu, which confuses Farsi speakers even more. For extra fun try the Nastaliq script.
This is what Dutch sounds like to me as an English speaker - plenty of common sounds with English; it has a similar speed, rhythm, intonation to English. It feels like I’m hearing English but have lost my faculties to parse it
"Goedemorgen, ik hoop dat je bent goed".
He was in the war -- Hij was in de war (he was confused)
A stiff in the brook -- Een stijve in de broek (a boner in the pants)
Those are the two most famous examples I'm familiar with, but I'm sure there are a lot more.
One of her friends was half of a multi-nationality couple, I think it was French and Irish, and the punchline was their kid, at a beach, yelling, in a strong Irish accent "Look mummy! Phoques!"
It was certainly something close to what I wrote.
Also I wonder if there's a connection between trousers being "broek" in Dutch and "breeks" in Scots.
edit: wow ok I should've just went to wikipedia: https://en.wikipedia.org/wiki/Breeks
"Jah" "Oui" "Sí" "Ja" "Da"
Portuguese and Italian on the other hand are much closer, and if spoken slowly enough, somewhat understandable.
If you live in Europe and speak a language using a Latin script, you probably have come across most of the extensions other European languages add to the shared base in loanwords or foreign media. But then you look at something like Vietnamese and you are no longer sure how letters work.
Looking back on it, I remember feeling like I can't remember Arabic, but part of it is that this also happens during that time when I'm getting used to the script. There is always an adjustment period with every new font/handwriting that takes a sentence or two to sort out the style before I truly start reading.
3. The text is in the wrong Arabic-script language, for example Farsi in Egypt (Egyptians primarily speak an Egyptian Dialect of Arabic), or Modern Standard Arabic in Afghanistan (Afghans speak Dari, an Afghan Persian language). Comme l'alphabet latin, vous pouvez écrire différentes langues avec l'alphabet arabe.
The stick character is either a short "a' or "i", which you would know from the little tick marks above/below that you see in decorative script, but are generally left out in print where you get it from context.
The J looking character is an "L". So they make "al" like al-jabaar => the strong. It's 95% like "el" in Spanish or "il" in Italian, though it is genderless.
The o with dots attached or not to the word on the right indicates feminine gender.
There's a half dozen languages at least that are not Arabic but use Arabic script, such as Persian or Urdu. The typographic rules mentioned still apply though.
Actually, yes. Yes I do.
Try having to work in documents that trade text in multiple languages between Adobe Acrobat, and Microsoft Office. Add in some opinions from iOS, and I end up pasting text into a blank ASCII file just to get it back to basics so I can send it to the next program because neither Microsoft nor Adobe can reliably handle the macOS standard Command-Shift-V to paste text unformatted.
I can't imagine the disaster that would await me if i also has to do it in Arabic.
He replied with a plaint .txt file with a couple of Arabic lines in it, with a message saying "I translated some of those text from your excel sheet, please go ahead and use this text, I will send the rest by tomorrow". Nobody understood which translated to which.
For spoken Arabic, a decent placeholder is either MSA (standard Arabic) or a generally understood dialect (Egyptian or Levantine).
The latter is what a lot of the recent Arabic game dubs have been doing (e.g., Ubisoft). Another example: Pixar and Disney use Egyptian Arabic for their dubs.
Edit: Oh, and for written Arabic, you should almost always be using MSA. That is, you shouldn’t worry about the dialect of your target audience.
There's a faux-katakana font someone used for the titles on an otherwise amazing album of some FM synth Touhou remixes[0] that hardcore fucks with my brain.
What's the right way to handle this?
[1] http://andreasmhallberg.github.io/typing-arabic-in-vim/
[2] https://www.w3.org/International/questions/qa-html-language-...
And you can always double check in a different editor to be sure nothing gets messed up. I use VS Code which, likely because its browser based, seems to get it right as well.
I encountered a (sort of) similar problem while rendering Devanagari text and wouldn't have realised If I couldn't read the text https://stackoverflow.com/questions/44254171/devanagari-text...
Speaking about translations, I had. Laughable experience when I was experimenting with Google translate API a few years ago, when I picked some random text about solar system from Wikipedia that reads something like “Mercury is a planet of the solar system”, when translated to Arabic, it read “الزئبق هو أحد كواكب المجموعة الشمسية” , apparently, it was not be able to infer the meaning of “Mercury” from the context.
Some of them let you click on each word and see which it picked, but that doesn't always work.
I am interested in Lontara and Bugis resources in general, if someone has something they would like to share.
https://heistak.github.io/your-code-displays-japanese-wrong/
Now it displays Chinese wrong.
That is, the best way to deal with apps being too clever about rich text, is to never give them rich text in the first place.
See
Avoid "Just" and "Simply" in Documentation http://jackkelly.name/blog/archives/2019/09/20/avoid_just_an...
Don’t say “simply” in your documentation by Jim Fisher https://www.knowledgeowl.com/blog/posts/dont-say-simply-jim-...
Notepad is for multi-line text.
EDIT: While the address bar trick may be slightly more convenient than Win+R (and the easiest thing to do on Linux), keep in mind that whatever you paste there will likely end up recorded on someone else's servers. I bet this was a major driver for integrating previously separate address bar and search bar, and adding "search suggestions" function...
Simple!
> The text is rendered left-to-right, instead of right-to-left. siht ekil daer ot gnivah ekil s'tI!
But the text (the arabic one) is not rendered in different direction bad vs good version? So this is just alignment? I can understand this still is very annoying and triggering, but not as bad as needing to read the other direction?!
1. There are several languages which (sometimes) use the Arabic script, but are not Arabic, like Kurdish, Turkish up until a century ago, Javanese etc.
2. Farsi sort-uses the Arabic script, with some glyphs unique to it
... so things can be more complex than just deciding "Is it Arabic?" . But I like the website!
----
Right-to-Left languages with non-Arabic script are also often similarly mis-rendered. I'd given the more popular scripts as an example, like N'Ko and Adlam, but almost nobody in the West bothers to print those, plus I'm not familiar enough with them; so - Hebrew:
In the disconnected-form script, here's Hello:
שלום עולם
(sounds like: shalom)
do you see the squarish glyph at the end of each of these words? That's the final-form Mem consonant. But in the beginning or the middle of the word, its glyph is מ. So if you see:
םולש
that means someone printed it in left-to-right order of glyphs instead of right-to-left.
-----
These issues come up a lot in our LibreOffice RTL languages user group on Telegram:
https://t.me/+GY3UwBnlDN9mY2M8
and if you're particularly interested in Arabic, there's the Arabic-issues-specific group:
Weird.
And there are plenty of people other than this guy making contributions to it, you’ve only just read a singular post by a single person outlining his most common gripes with companies that do not consult native speakers for their translations. It’s not an exhaustive list of people working in the sphere. Imagine if you read a blog post by Linus Torvalds, and thought: ‘wow, it’s crazy that no one else is making contributions to this cool open source kernel?’ - that’s exactly what you’ve done in your comment.
how do I interpret that? something like "hundreds of millions of arabs constitute a minority of 2 billion"? or are you saying that vast majorities in the arab countries don't use arabic script to write things daily?
Many can't even get the spaces before and after punctuation right, even though the correct way involves no irregularities.
I think in software once there's one great, free implementation of a hard problem there's a tendency to stick with it. If that project is maintained by one person, then so be it.
1. https://onezero.medium.com/the-largely-untold-story-of-how-o...
If 1 person does the job very well others don't have to. That's the way we make progress in software.
Funny enough, Rami is Egyptian/Dutch and ل ا looks a bit like IJ/ij, which is a digraph in Dutch that even used to have its own codepoint once upon a time.
A digraph for which two existing letters in the ASCII code already worked perfectly had a code-point, and I'd argue that even historically almost no Dutch person has ever even used it because keyboards with it barely existed. The only place where you might find an IJ key is on ancient electronic typewriters, and the only reason I know is because I had touch typing lessons on one thirty years ago and the teacher made us use them, after which I had to unlearn that because no computer keyboard featured it.
And meanwhile Windows apparently can't even be bothered to copy/paste Arabic correctly (at least not in 2015 when the embedded video was recorded).
I am unsurprised that a website by Rami Ismail uses collective offense-taking as its frame, and uses it to try shaming companies into giving money to non-westerners or adjacent.
It's not automatically insulting to misrender arabic, any more than it is insulting to read hilariously mis/autotranslated "chinglish" in China, or see comically offensive English tshirt slogans in Japan. Depending on context it could even be an attempt at being _nice_ that just misfired.
A simpler explanation and frame is that text rendering is tough enough as it is, and a solution that works perfectly fine for one market won't scale to everywhere. Arabic in particular is very hard, not just because of ligatures, but because right-to-left text is actually bidirectional in practice.
They might've even asked an Arabic speaker to review, but it may have looked okay right until the very last step in the production pipeline, when it got mangled.
> for example Farsi in Egypt (Egyptians primarily speak an Egyptian Dialect of Arabic)
Nit picking here maybe, but neither Farsi (Persian) is a dialect of Arabic (not even in the same family of languages) nor Egyptians speak it typically. Although, they speak a dialect of Arabic.Update: My mistake, as pointed out here, I misunderstood the point initially.
Oddly though, it stores several copies of identical chinese characters. Traditional, simplified, and the Japanese loaned characters kanji.
Actually, all of those characters are usually unified into one code point. This is controversial enough on its own that there's a Wikipedia page on it: https://en.wikipedia.org/wiki/Han_unification
Traditional, simplified, Japanese, Korean. Some are identical, some different. I wonder why these rendering details were not left to fonts and you can type 雨 and ⾬ for example. Maybe stronger presence of those countries in whatever SDOs and tech integration has to do with it.
Ages ago, in 2005 or thereabouts, I was working on a website for a Jewish organisation. I don't know Hebrew, but Hebrew has many of these same issues, including the right-to-left direction, and I was quite surprised to see text from our own CMS, based on Java, with tons of XML pipelines (Cocoon! XSLT!) generating HTML to be viewed in a browser, just handled this correctly without any problems.
At least the text was right-to-left, which was the only thing I knew, but the customer presumably knew more and they were happy.
Though selecting text in a mixed left-to-right and right-to-left piece of text looks really weird. Not sure what we did with alignment there, but I vaguely recall that everything may have been centered. An ugly compromise perhaps, but centering text was still popular at the time.
(Arabic also has the connection challenge, which Hebrew does not, but usually renderers that do RTL right also handle the ligatures.
> 7. Progress bars fill in the same direction as content is read
In general, if it represents "start to finish", or it has "forwards or backwards", then it should be mirrored. Above site has good examples for how a bicycle icon should be mirrored (as facing left is going "forwards"), but a magnifying glass would not.
When reading the URL, I imagine an Arabic language called “Photoshop Arabic”, which came into being at the Photoshop team in Adobe XD
- Do you speak Arabic?
- Arabic is not just one language.
- Yeah, but do you speak an Arabic language?
- Yeah, I speak Photoshop Arabic.
How is a tacky font on some text that was never intended to be Arabic even approaching the badness of these other issues, let alone worse?
Or is the worry you've found a person that doesn't understand different languages exist at all?
I have definitely seen attempts at a "global-looking" interface where fake languages were put on screen using plausible-looking fonts.
Like, imagine a welcome page with "hello!" in lots of different languages, but the Arabic/Urdu/Armenian/Thai characters are all just literal gibberish that looks kind of like the alphabets in question.
If someone is using those fonts to pass text off as being in a different language, as in the example you mention, then sure, that's egregious.
If, however, the font is use to render some text in the style of another language, then I don't see a problem with it. We've done the same with Greek, Russian, Japanese, and many other languages that have a distinctive script. Sometimes it's tacky, but I wouldn't call it insulting.
My understanding of Arabic is that this would be _exceedingly_ hard to generalize into individual, discrete glyphs that could not only be broken out into their own segments but then repeated in a way that makes sense.
The only way I could personally think of is with matrix displays, which act more like normal displays.
This is because even though it is an alphabetic script, the blending between characters is non-trivial and highly dependent on the sequence of glyphs, which causes an explosion of possible combinations depending on what's being written.
That said, you can always design segmented displays that show _any_ shape or script, assuming you don't need to change them.
I'm going to guess the various shifts you allude to are simple transformations for a pen-plotter, perhaps accumulating some "context" offsets to the initial position or flipping some glyph etc. Not sure how to translate that to segments, unless you can electro-mechanically adjust the segments...?...
But it doesn't end there. Fonts have a crazy number of different lookup tables that can affect how glyphs are drawn. While not Turing complete they're certainly very close.
In Arabic, how the baseline blends affecting the base shape of the glyph, plus lots of kerning rules, etc. would make Arabic very difficult to replicate on a segmented display.
A 7 segment display would have nowhere close to enough resolution to properly render Arabic script. Digital signage would at least be using a low-resolution LED matrix but more likely you'd see LED/LCD displays in places that are affluent and plain analog signage where they are not.
I wonder how Arabic ended up with characters that differ only by dot positioning (like ب ت ث, and also ن whose combining form is especially similar to the combining forms of those). By contrast, the most similar Latin characters might be EF / CG / IJ / MN / UV which I think are clearly more different than the Arabic characters that differ only by dot quantity and positioning.
(The Il1| are also famously confusable in some Latin fonts, but one could say having to worry about these comes late in the history of Latin script writing. Although the scribal minim https://en.wikipedia.org/wiki/Minim_(palaeography) used at some points to write u, i, m, and n is every bit as confusing as any Arabic character.)
Much like latin characters, at some point a segmented display is just going to have to compromise on the tradeoff between correct an the number of parts. We see it work just fine casting the curvy latin letters to squares. So the real question I reckon is how much can you mangle the diacritics / curves / other parts and still read it? In other words, what is the bandwidth redundancy of this writing system?
Other commenter nailed the real technical hurdle. Context dependent spacing system. How do monotypes do it (or do those not exist)?
A 7-segment display works by turning on and off each of the segments in order to render a glyph that is supposed to represent English alphanumeric characters. It mostly accomplishes this, though just barely and honestly a lot of characters do not render well at all (like the lowercase "i"). How would you shape the segments in such a way that could accomplish at least that level of clarity in Arabic?
I personally think that what you'd end up with would very closely resemble a dot matrix.
https://patentimages.storage.googleapis.com/a5/dc/27/34d036c...
(I have no idea where this number comes from but it is vastly inflated.)
All that being said, this is a decent primer on common Arabic rendering mistakes.
Never hurts to support popular languages properly, such as Arabic.
Add in non-Muslim Arabs and you've hit 2 billion right quickly. The assumption is that a Muslim can recognize the Quran's written language somewhat.
I am deeply involved in the Muslim community, and can safely say that reciting holy works in Arabic offers zero useful knowledge of the modern language.
[0]: https://translate.google.com/?sl=en&tl=de&text=plant%20shape...
But besides, I doubt "plant shaped C4" contains enough information to guarantee a correct answer in that case. Trying partial sentences and/or individual words through (human or machine) translations don't go well.
The plant shaped was not a fragment, but the whole text, this has a screenshot [0]. It was about fact-checking the Seymour Hersh Nordstream 2 article.
[0] https://nitter.librenode.org/Samy_t42/status/162879312196521...
one trick is to check if the text looks the same as in google translate after you pasted it
This has led to some unfortunate tattoos: https://bleacherreport.com/articles/2431265-mario-mandzukics...
There are probably 350 million Arabic speakers.
It's possible that the article refers to the number of readers of Arabic script in any language, which I imagine would be a bigger number.
Because it's comparing the wrong numbers. The article specifies what the 2 million refers to:
"Here's a 18 minute video of I talk I gave at XOXO in 2015 that'll teach you more than enough Arabic to not embarrass yourself in front of everyone who can read the Arabic alphabet to some degree. That's about 2 billion people, or 28% of the human population, and they will absolutely notice that you don't know what you're doing and laugh at you in languages you cannot even read."
Oops.