Thai character ก็็็็็ (ก) gets rendered in a strange way
google.co.uk
google.co.uk
Thai belongs to a family of scripts known as abugidas. Abugidas include pretty much all South Asian and many Southeast Asian scripts, for example Burmese, Cambodian, Dai, Lao, Thai, etc. They all pretty much derive from Brahmi, which was the proto Indian script. You can see an example of Brahmi over here: http://en.wikipedia.org/wiki/Brahmi
Abugidas are based upon combining multiple glyphs in to syllables, often allowing glyphs above, below, to the left and to the right of the initial consonant, and often including a closing consonant. Most glyphs tend to be consonants, though some are vowels, and others can be special marks for indicating tone or other notions. Often shorter vowels are excluded (as in Modern Standard Arabic).
In old times, such scripts were handled with wacky font-hacks. However, with Unicode, there are some super complex algorithms that make glyphs combine both visually (when typesetting) and logically (when saving/searching/etc). You can actually type a character and a diacritic and it can sometimes automatically combine to form a single character, if such a beast exists, not just visually but when saving to disk.
What makes it even more confusing is that South Asian scripts in particular have mega-combo characters, where whole chunks of glyphs sort of fold in to flowing short-hand symbols. In the case of Sanskrit, I believe loads of these were used in history but few are used these days.
I think that's a fair pontification - corrections welcome!
Tangential tidbit: I sent a copy of The Cambodian System of Writing (http://pratyeka.org/csw/) to TPB's anakata while he was solitary confinement to help him stave off boredom. No idea if he ever read it, though his mother assures me it arrived.
Thai looks to be a pretty gr๏๏vy language, like many other natural languages - perhaps some programming languages will catch up in their enhanced use of lexical tokens one day, instead of just relying on grammar, long English names packed into name hierarchies, and multi-ASCII symbols.
The worst thing about "not using space" when it comes to computer is that it's nearly impossible to do word-breaking/line-breaking without relying on dictionary[1] which is very hard to convince software developers not using system text engine to add support for line breaking[2].
The suffer still continues to this date, Android, for instance, still lack a proper support (break by character instead of word) and few apps lack of Thai word-breaking support at all (Twitter, whose only do word break by space).
[1]: http://linux.thai.net/svn/software/libthai/trunk/data/ [2]: https://bugzilla.mozilla.org/show_bug.cgi?id=7969 it took Mozilla 9 years. Opera, on the other hand, still lacks proper Thai support.
Of course you also have multiple words that have been combined into one large compound word, by way of appropriate linguistic rules of combining sounds. This is similar to long compound words in German.
Ę̮̱͔͓ͯ͗ͫ̌̏ͫ͌́x̘̤͚̰̫̫̗̤̱̒̓ͨͯ͑̓ͥͫ̕å̰͚̓͒ͫm̛̤͕̫̳̺̩̄̓ͨͥ͜ͅp̰͉͗ͤl̵̖̗̫͍͓͋̍̐͌̐̒e̡̧͔̮̿͒͋̈́͡ ̸͉͔͗͐̍ͩͫ̀ͭz̨͎̱̟̘̓ä́͊̉̾͜͏̺̲̘l̛̥͇͖̹̻̜̈̀̀g̴̗̻͚͙̭͍̩̔̉̆ͦ͌͘oͬ̾͑̉̋҉̢͙̹̹̺̺ ̷̢͖̲͇̺̪̹̙̺̘͐̄ͬ̍͆t̶͔̣̜̟͌̀ͪ̅ͧ̒̒ͫ̚ȅ̠̪̻̄ͫ̋͝xͭ͆͝͏̮͔̜t̟̬̦̣̟͉͈̞̝ͣͫ͞,̡̼̭̘̙̜ͧ̆̀̔ͮ́ͯͯ ̢̮͎̦͙͇ͪͪ̈͌ͬ̄̓̐͞ḷ̹̺̙̜̇̉́͡o̢̻̪̠̬̍͐̉ͮͥ̑͊ͪt̢̘̬͓͕̬́ͪ̽́s̢̜̠̬̘͖̠͕ͫ͗̾͋͒̃͛̚͞ͅ ̝̣̥̳͇͎̭̾̔̀̀̔̽̕o͇ͮ̋̅͋͆̈́̔͗͟f̙̙͕̮̈ͪͯ̿̈͠ ̯͎̺͎̺̃̀͟͟d͍͍̺͂̂i̪̩̙̭̝͖ͥ͂̂̈̒̎r̥̜̃̏̃͋̓ͥ̃̉̄͘͢t̳̦̬͆͂ͬͧ̏ͬ̓y̵̮̗̟ͩ̃̾͐́ͩ ̣͍̘͈̫͓̊ͤ̚͡͝cͥͭ͐̎͆͘̕҉̫̞h̴̢̫̘͉̖ͪͩ̓ͪͯ̑͑̓̎͝a̧̢̖͔̗̬̘̯̟ͪ̐͌̍͂̊r̷̝͓̬͆̄̽̓̋ͬ̈̔͝͠ā̗͑ͬ̀c͒̎͌̔͛͘҉̘͖͖̖̯̖͖͙ṱ̶͇͚͎ͯ͋͢͝eͦ̽͆͏̟̭̠r̙̖͙̳̾ͯ̈̕ṣ͙̈͆̔͗̉ͥ̋̔̕
Although searching for the same text results in Google telling me this:
414. That's an error.
The requested URL /... is too large to process. That’s all we know.Why - seems to select perfectly well when I try it using Firefox on Windows 7.
I really don't like trying to read 1400px long lines of text.
media-fonts/arphicfonts
media-fonts/baekmuk-fonts
media-fonts/cardo
media-fonts/corefonts
media-fonts/dejavu
media-fonts/droid
media-fonts/font-bh-lucidatypewriter-100dpi
media-fonts/font-bh-lucidatypewriter-75dpi
media-fonts/font-bh-ttf
media-fonts/font-bh-type1
media-fonts/freefont
media-fonts/freefonts
media-fonts/inconsolata
media-fonts/intlfonts
media-fonts/kochi-substitute
media-fonts/symbola
media-fonts/terminus-font
media-fonts/ttf-bitstream-vera(to me it looks like some kind of particle accelerator experiment happening in your browser, where ก็็็็็ emits some kind of unknown radiation. After closer examination [zoom to +300%]: maybe it just shows the escaping life spirits of the toppled latin small letter «u» after being shot in right side.)
But it's not a "rendering" bug.
This is an interesting result. Even more interesting is the extra 2000 results that Google throws in my direction.
Both Opera and Chrome on Windows XP. On windows 7 I get the interestingly looking results with both Opera and Chrome.
ก ็ ็ ็ ็ ็ก ็ ็ ็ ็ ็ ็ ็ ็ ็ ็ ็ ็ ็ ็ ็ ็ ็ ็ ็ ็
$ echo ก็็็็็็็็็็็็็็็็็็็ | hexdump
0000000 e0 b8 81 e0 b9 87 e0 b9 87 e0 b9 87 e0 b9 87 e0
0000010 b9 87 e0 b9 87 e0 b9 87 e0 b9 87 e0 b9 87 e0 b9
0000020 87 e0 b9 87 e0 b9 87 e0 b9 87 e0 b9 87 e0 b9 87
0000030 e0 b9 87 e0 b9 87 e0 b9 87 e0 b9 87 e0 b9 87 0a
0000040
$ echo ก็็็็็็็็็็็็็็็็็็็ | wc
1 1 64That would be zalgo: http://eeemo.net/
glitchr's tweets may cause other twitter clients to crash too, eg this one (you have been warned!) https://twitter.com/joshlogan42/status/303975029698342912
HN is now not wrapping long lines, because your unbroken long line has widened the margins.
I come home and look at it on my Windows machine with Chrome, and now I see the big stack of diacritic marks that I assume everyone's making a fuss about. I assume it's something to do with the way that the system's installed font lays out the marks in question.
𝑀𝑎𝑛𝑦 𝑝𝑒𝑜𝑝𝑙𝑒 𝑤𝑜𝑛'𝑡 𝑏𝑒 𝑎𝑏𝑙𝑒 𝑡𝑜 𝑠𝑒𝑒 𝑡ℎ𝑖𝑠, 𝑜𝑟 𝑎𝑡 𝑙𝑒𝑎𝑠𝑡 𝑎𝑙𝑙 𝑜𝑓 𝑡ℎ𝑒 𝑐ℎ𝑎𝑟𝑎𝑐𝑡𝑒𝑟𝑠, 𝑠𝑖𝑛𝑐𝑒 𝐼 𝑡ℎ𝑖𝑛𝑘 𝑎 𝑈𝑛𝑖𝑐𝑜𝑑𝑒 6.0 𝑓𝑜𝑛𝑡 𝑖𝑠 𝑟𝑒𝑞𝑢𝑖𝑟𝑒𝑑.
𝔼𝕧𝕖𝕟 𝕗𝕖𝕨𝕖𝕣 𝕗𝕠𝕟𝕥𝕤 𝕙𝕒𝕧𝕖 𝕥𝕙𝕖 𝕗𝕦𝕝𝕝 𝕕𝕠𝕦𝕓𝕝𝕖-𝕤𝕥𝕣𝕦𝕔𝕜 𝕒𝕝𝕡𝕙𝕒𝕓𝕖𝕥, 𝕥𝕙𝕠𝕦𝕘𝕙 𝕚𝕥 𝕨𝕠𝕣𝕜 𝕗𝕚𝕟𝕖 𝕗𝕠𝕣 𝕞𝕖 𝕠𝕟 𝕆𝕊 𝕏.
Also, it seems to not work on all browsers and even then, FF and IE do slightly different things : http://i.imgur.com/hfWu5Bs.png
I'm on Win7.
Edit: I just noticed, on FF, the character spills out of the tab preview text and onto the chrome background as well.
I (OS X; crome) get little blobs over the n. That's wrong? But doesn't break the page?
Is it a "Like a Boss" character?
http://www.leer-leren.com/wp-content/uploads/2012/07/Like-a-...
.st, a {
display: inline-block;
overflow: hidden;
}
Hacker news should do the same thing but with .title, .comment and .comheadhttp://i.imgur.com/698dzMo.png
After trying with the css suggested here on Google inside Chrome, the problametic characters don't have the top going across too high like originally, but it gets capped on the top of the first line of the paragraph, instead of the line the character is actually at.
That said, I'm not sure how to solve this though.
But just for the technical challenge you could do something like this in Javascript to chop every individual special character:
string.split("").map(function(a){
return /[a-z0-9\s]/i.test(a) ? a : '<span class="s_char">' + a + '</span>';
}).join("");
And apply the mentioned CSS to the "s_char" class.That said, if those sorts of things bother an individual, they could run that on the page themselves, I suppose, so it's good for that :)
It seems like there isn't really any CSS only solution for this without wrapping every character in its own element like jQueryIsAwesome suggested.
If stacking diacritics are a legitimate part of a language and they change the meaning of words or characters, then it's more important that the content be correct than the "usability."
In fact, if a person can't read it or reads it incorrectly because parts of characters are hidden, then it's not very usable.
Hypothesizing about a Unicode character that fills the screen with black is a nonsensical straw man, because it makes no sense in "real world" written languages, so there would never be a Unicode character for it.
This kind of stuff is precisely the reason why we make sure that every filename we create is only using a subset of ASCII (and no space of course). In our source code, in our builds, in the desktop app we're serving, etc.
Unicode characters entered by users should go in one place: the DB.
I smiled the other day when I read about the build script for Chromium: it clearly specificied that the source directory must not contain any space in its name.
Of course it shouldn't. That's experience.