Unicode Text Converter
panix.com
panix.com
Suggestion (if you are author): There are a lot of chars that look like another char, often used on the web, so i think that there are more advanced versions to be made. I think i read that a lot of thai signs and cyrillic look like latin chars.
So they are using this trick (but in the opposite direction, latin `a` instead of cyrillic `а`) to avoid undesired competitors from entering those biddings and lowering the purchase prices (and not paying kickbacks, obviously). https://navalny-en.livejournal.com/52565.html
𝖕𝖚𝖇𝖑𝖎𝖈 𝖛𝖔𝖎𝖉[] 𝖒𝖆𝖎𝖓(𝖘𝖙𝖗𝖎𝖓𝖌[] 𝖆𝖗𝖌𝖘) {
𝕮𝖔𝖓𝖘𝖔𝖑𝖊.𝖂𝖗𝖎𝖙𝖊𝕷𝖎𝖓𝖊("𝕳𝖆𝖑𝖑𝖔 𝖂𝖊𝖑𝖙");
}
// 𝕽𝖊𝖈𝖍𝖊𝖓𝖒𝖆𝖘𝖈𝖍𝖎𝖓𝖊𝖓𝖘𝖕𝖗𝖆𝖈𝖍𝖊 "𝕱𝖗𝖆𝖐𝖙𝖚𝖗" 𝕰𝖎𝖓𝖘 𝕻𝖚𝖓𝖐𝖙 𝕹𝖚𝖑𝖑 𝕹𝖚𝖑𝖑 # 𝕲𝖊𝖒ä𝖘𝖘 𝕽𝖊𝖎𝖈𝖍𝖘𝖆𝖚𝖘𝖘𝖈𝖍𝖚𝖘𝖘 𝖋ü𝖗 𝕬𝖑𝖌𝖔𝖗𝖎𝖙𝖍𝖒𝖎𝖘𝖈𝖍𝖊 𝕬𝖗𝖇𝖊𝖎𝖙
𝖐𝖑𝖆𝖘𝖘𝖊 𝕭𝖊𝖌𝖗ü𝖘𝖘𝖚𝖓𝖌𝖘𝖆𝖓𝖟𝖊𝖎𝖌𝖊𝖇𝖊𝖉𝖎𝖊𝖓𝖒𝖊𝖈𝖍𝖆𝖓𝖎𝖘𝖒𝖚𝖘:
𝖉𝖊𝖋 __𝖆𝖓𝖋𝖆𝖓𝖌𝖊𝖓__(𝖘𝖊𝖑𝖇𝖘𝖙, 𝖁𝖔𝖗𝖓𝖆𝖒𝖊):
𝖘𝖊𝖑𝖇𝖘𝖙.𝖁𝖔𝖗𝖓𝖆𝖒𝖊 = 𝖁𝖔𝖗𝖓𝖆𝖒𝖊
𝖉𝖊𝖋 __𝖘𝖈𝖍𝖓𝖚𝖗__(𝖘𝖊𝖑𝖇𝖘𝖙):
𝖟𝖚𝖗ü𝖈𝖐𝖌𝖊𝖇𝖊𝖓 𝖘𝖊𝖑𝖇𝖘𝖙.𝖁𝖔𝖗𝖓𝖆𝖒𝖊
𝖉𝖊𝖋 𝖇𝖊𝖌𝖗ü𝖘𝖘𝖊𝖓(𝖘𝖊𝖑𝖇𝖘𝖙, 𝖁𝖔𝖗𝖓𝖆𝖒𝖊=𝕹𝖎𝖈𝖍𝖙𝖊𝖝𝖎𝖘𝖙𝖊𝖓𝖟):
𝖉𝖗𝖚𝖈𝖐𝖊𝖓("𝕲𝖚𝖙𝖊𝖓 𝕿𝖆𝖌, " + 𝖘𝖊𝖑𝖇𝖘𝖙.𝖁𝖔𝖗𝖓𝖆𝖒𝖊)
𝖟𝖚𝖗ü𝖈𝖐𝖌𝖊𝖇𝖊𝖓 𝖘𝖊𝖑𝖇𝖘𝖙
𝖇𝖊𝖌𝖗ü𝖘𝖘𝖊𝖗 = 𝕭𝖊𝖌𝖗ü𝖘𝖘𝖚𝖓𝖌𝖘𝖆𝖓𝖟𝖊𝖎𝖌𝖊𝖇𝖊𝖉𝖎𝖊𝖓𝖒𝖊𝖈𝖍𝖆𝖓𝖎𝖘𝖒𝖚𝖘("𝕳𝖆𝖓𝖘-𝕻𝖊𝖙𝖊𝖗 𝕯𝖊𝖚𝖙𝖘𝖈𝖍" )
𝖇𝖊𝖌𝖗ü𝖘𝖘𝖊𝖗.𝖇𝖊𝖌𝖗ü𝖘𝖘𝖊𝖓() 𝕹𝖊𝖚𝖎𝖌𝖐𝖊𝖎𝖙𝖘𝖇𝖑𝖆𝖙𝖙 𝖉𝖊𝖗 𝖚𝖓𝖐𝖔𝖓𝖛𝖊𝖓𝖙𝖎𝖔𝖓𝖊𝖑𝖑𝖊𝖓 𝕽𝖊𝖈𝖍𝖊𝖓𝖒𝖆𝖘𝖈𝖍𝖎𝖊𝖓𝖎𝖓𝖌𝖊𝖓𝖎𝖊𝖚𝖗𝖊 | 𝕹𝖊𝖚𝖊𝖘 | 𝕶𝖔𝖓𝖛𝖊𝖗𝖘𝖆𝖙𝖎𝖔𝖓𝖊𝖓 | 𝕶𝖔𝖒𝖒𝖊𝖓𝖙𝖆𝖗𝖊 | 𝕱𝖗𝖆𝖌𝖊𝖘𝖙𝖊𝖑𝖑𝖚𝖓𝖌 | 𝕭𝖊𝖗𝖚𝖋𝖘𝖋𝖎𝖓𝖉𝖚𝖓𝖌𝖘𝖆𝖇𝖙𝖊𝖎𝖑𝖚𝖓𝖌 | 𝕰𝖎𝖓𝖗𝖊𝖎𝖈𝖍𝖊Which is bullshit and just a parody on linguistic purism.
I could write more, but i have to configure the Zuwachssicherung of my Klapprechner over DFÜ.
𝕭𝖊𝖌𝖗𝖚̈𝖘𝖘𝖚𝖓𝖌𝖘𝖆𝖓𝖟𝖊𝖎𝖌𝖊𝖇𝖊𝖉𝖎𝖊𝖓𝖒𝖊𝖈𝖍𝖆𝖓𝖎𝖘𝖒𝖚𝖘
Right :). Though it's not quite centred for me.
𝔇𝔞𝔰 𝔠𝔬𝔪𝔭𝔲𝔱𝔢𝔯𝔪𝔞𝔠𝔥𝔦𝔫𝔢 𝔦𝔰𝔱 𝔫𝔦𝔠𝔥𝔱 𝔣𝔲𝔢𝔯 𝔤𝔢𝔣𝔦𝔫𝔤𝔢𝔯𝔭𝔬𝔨𝔢𝔫 𝔲𝔫𝔡 𝔪𝔦𝔱𝔱𝔢𝔫𝔤𝔯𝔞𝔟𝔟𝔢𝔫. ℑ𝔰𝔱 𝔢𝔞𝔰𝔶 𝔰𝔠𝔥𝔫𝔞𝔭𝔭𝔢𝔫 𝔡𝔢𝔯 𝔰𝔭𝔯𝔦𝔫𝔤𝔢𝔫𝔴𝔢𝔯𝔨, 𝔟𝔩𝔬𝔴𝔢𝔫𝔣𝔲𝔰𝔢𝔫 𝔲𝔫𝔡 𝔭𝔬𝔭𝔭𝔢𝔫𝔠𝔬𝔯𝔨𝔢𝔫 𝔪𝔦𝔱 𝔰𝔭𝔦𝔱𝔽𝔢𝔫𝔰𝔭𝔞𝔯𝔨𝔢𝔫. ℑ𝔰𝔱 𝔫𝔦𝔠𝔥𝔱 𝔣𝔲𝔢𝔯 𝔤𝔢𝔴𝔢𝔯𝔨𝔢𝔫 𝔟𝔢𝔦 𝔡𝔞𝔰 𝔡𝔲𝔪𝔭𝔨𝔬𝔭𝔣𝔢𝔫. 𝔇𝔞𝔰 𝔯𝔲𝔟𝔟𝔢𝔯𝔫𝔢𝔠𝔨𝔢𝔫 𝔰𝔦𝔠𝔥𝔱𝔰𝔢𝔢𝔯𝔢𝔫 𝔨𝔢𝔢𝔭𝔢𝔫 𝔡𝔞𝔰 𝔠𝔬𝔱𝔱𝔢𝔫-𝔭𝔦𝔠𝔨𝔢𝔫𝔢𝔫 𝔥𝔞𝔫𝔰 𝔦𝔫 𝔡𝔞𝔰 𝔭𝔬𝔠𝔨𝔢𝔱𝔰 𝔪𝔲𝔰𝔰; 𝔯𝔢𝔩𝔞𝔵𝔢𝔫 𝔲𝔫𝔡 𝔴𝔞𝔱𝔠𝔥𝔢𝔫 𝔡𝔞𝔰 𝔟𝔩𝔦𝔫𝔨𝔢𝔫𝔩𝔦𝔠𝔥𝔱𝔢𝔫.
This somehow reminded me of this one, in pseudo-Old Church Slavonic: http://lurkmore.so/images/d/d6/Pravoslavnii_koding.jpg
Sad thing is, Unicode still doesn't seem to properly support titlos and (not so sad, since personally I think Unicode shouldn't really do anything with fonts unless absolutely necessary) has no separate characters for Ustav and Poluustav scripts.
.𝖋𝖑𝖆𝖌,.𝖋𝖑𝖆𝖌:𝖇𝖊𝖋𝖔𝖗𝖊,.𝖋𝖑𝖆𝖌:𝖆𝖋𝖙𝖊𝖗{𝖈𝖔𝖓𝖙𝖊𝖓𝖙: ''; 𝖉𝖎𝖘𝖕𝖑𝖆𝖞: 𝖇𝖑𝖔𝖈𝖐; 𝖜𝖎𝖉𝖙𝖍:100𝖕𝖝; 𝖍𝖊𝖎𝖌𝖍𝖙: 20𝖕𝖝;}
.𝖋𝖑𝖆𝖌{𝖇𝖆𝖈𝖐𝖌𝖗𝖔𝖚𝖓𝖉: #000; 𝖕𝖆𝖉𝖉𝖎𝖓𝖌-𝖙𝖔𝖕: 20𝖕𝖝}
.𝖋𝖑𝖆𝖌:𝖇𝖊𝖋𝖔𝖗𝖊{𝖇𝖆𝖈𝖐𝖌𝖗𝖔𝖚𝖓𝖉: #𝖋00; }
.𝖋𝖑𝖆𝖌:𝖆𝖋𝖙𝖊𝖗{𝖇𝖆𝖈𝖐𝖌𝖗𝖔𝖚𝖓𝖉:#𝖋𝖋0}
(https://twitter.com/nickheer/status/535129309531635712) .𝖋𝖑𝖆𝖌,.𝖋𝖑𝖆𝖌:𝖇𝖊𝖋𝖔𝖗𝖊,.𝖋𝖑𝖆𝖌:𝖆𝖋𝖙𝖊𝖗{𝖈𝖔𝖓𝖙𝖊𝖓𝖙: ''; 𝖉𝖎𝖘𝖕𝖑𝖆𝖞: 𝖇𝖑𝖔𝖈𝖐; 𝖜𝖎𝖉𝖙𝖍:100𝖕𝖝; 𝖍𝖊𝖎𝖌𝖍𝖙: 20𝖕𝖝;}
.𝖋𝖑𝖆𝖌{𝖇𝖆𝖈𝖐𝖌𝖗𝖔𝖚𝖓𝖉: #000; 𝖕𝖆𝖉𝖉𝖎𝖓𝖌-𝖙𝖔𝖕: 20𝖕𝖝}
.𝖋𝖑𝖆𝖌:𝖇𝖊𝖋𝖔𝖗𝖊{𝖇𝖆𝖈𝖐𝖌𝖗𝖔𝖚𝖓𝖉: #𝖋𝖋𝖋; }
.𝖋𝖑𝖆𝖌:𝖆𝖋𝖙𝖊𝖗{𝖇𝖆𝖈𝖐𝖌𝖗𝖔𝖚𝖓𝖉:#𝖋00}
wouldn't this be more appropriate?[1] http://en.wikipedia.org/wiki/Question_mark#Greek_question_ma...
Russian is my first language, but English is my primary language, and I never had my chance to practice typing using the standard Russian keyboard layout, so I almost always use the "Phonetic" layout - where the latin c is the cyrillic ц. (Also, w is ш, and who the hell remembers what []\-= map to - always trial and error for me to find южэьъ.)
In [4]: class АnotherClass():
...: pass
File "<ipython-input-4-ad6e67ea5e19>", line 1
class АnotherClass():
^
SyntaxError: invalid syntax[1] PEP 0263: https://www.python.org/dev/peps/pep-0263/
⎧1 if n = 0;
F(n) ≡ ⎨1 if n = 1;
⎩F(n-1) + F(n-2) if n > 1.
⎛ ∇∙D⃑ = ρ ⎞
⎜ ∇∙B⃑ = 0 ⎟
⎜ ∇×E⃑ = -∂B⃑/∂t ⎟
⎝ ∇×H⃑ = J⃑ + ∂D⃑/∂t ⎠
⌠¹
π = 2⎮ √1̅̅-̅̅x̅̅²̅̅ dx
⌡₋₁
⎡1 0 1⎤ ⎡î⎤
⎢0 1 0⎥ ⎢ĵ⎥
⎣1 0 1⎦ ⎣k̂⎦
Γ ⊢ t:S S<:T
――――――――――――――― (T-Sub)
Γ ⊢ t:T
⎛ 1 ⎞ⁿ
ℯ = lim ⎜1+ ― ⎟
ⁿ→∞ ⎝ n ⎠𝑟₁ 𝑣
𝐷 → 𝑅 → 𝑉
𝛼 ↓ ↓ 𝜔
𝐷 → 𝑅 → 𝑉
𝑟₂ 𝑣
https://twitter.com/mxfh/status/532575085337792512
formula is from http://algebraicvis.net/
╔════════════════════╗
║ Yes. All manually ║
╙────────────────────╜ ⎛if ⎛> (+ a b)⎞ ⎛case x ⎞ ⎛cond ⎞⎞
⎜ ⎝ (- c d)⎠ ⎜ (1 'foo)⎟ ⎜ ((> y 2) 'quux) ⎟⎟
⎜ ⎜ (2 'bar)⎟ ⎝ (t 'error)⎠⎟
⎝ ⎝ (3 'baz)⎠ ⎠
...(hmmm. For some reason that looks better in my editor than on the webpage. Apparently a fixed width font isn't necessarily fixed when it comes to unicode).I'd say that the truncation algorithm operates on bytes and that it can't make sense of d8 35, but I'm not too sure how to fix that since graphemes can have arbitrary length (right?). Do you have to compute the width in advance?
There are libraries for doing it in Javascript: https://www.npmjs.org/package/grapheme-breaker (is that part of the Firefox UI done in Javascript? I've no idea)
This seems likely, as another notable weirdness is that even with full width tabs, where there's plenty of space for at least "𝑼𝒏𝒊𝒄𝒐𝒅𝒆 𝑻𝒆𝒙𝒕..." it still only shows "𝑼𝒏𝒊𝒄𝒐...".
An online version: http://www.pseudolocalize.com/
A library: http://code.google.com/p/pseudolocalization-tool/
By randomly mixing these Unicode letter and letterlike characters, you can simulate a cut-and-paste ransom-note. For example, an acquired company could announce changes to its privacy policy:
wE ℎåve yøuR ρrIvᴀçy ⅈn a ᴡiNdøwleSs ℞oøm,
& ℙℓaℕ τø ⅆo µnSρεaKᴀble †hiℕℊs t○ ⅈtThe cat should have stayed in a box, if this gains too much popularity, HN will read like MySpace back in the days.
And top HN news will be: "A browser plugin that translates Unicode back to ASCII".
The user though the oauth app was legit because it was the "same" as the company name, accepted the connection, and promptly had their account emptied: https://www.reddit.com/r/Bitcoin/comments/2lt76n/warning_coi...
Now, it's up for debate whether any (psuedo?) financial institution should offer full oauth access (at least without having a human review possible oauth connectors), but the point is, decorative hackernews submissions are the least malicious use of this trick.
We'll need a plugin to reverse this, anyone up for it?
On my windows box with chrome all i see are empty boxes.
I'm fairly sure this is no longer the case. Chrome is high-DPI aware on Windows now, and it uses DirectWrite for font rendering, the same as IE. It just can't display these characters for some reason.
Anyway, DirectWrite was horrible at high DPI, if I remember correctly.
I find Chrome better than IE, actually. IE ignores my DPI settings and scales pages to 250%, so everything looks too large. Chrome renders correctly at 200%.
𝒃𝒆𝒔𝒕 𝒗𝒊𝒆𝒘𝒆𝒅 𝒊𝒏 𝒊𝒏𝒕𝒆𝒓𝒏𝒆𝒕 𝒆𝒙𝒑𝒍𝒐𝒓𝒆𝒓 11
This reminds me of the 1990s. haha
http://gschoppe.com/fixing-unicode-support-in-google-chrome/
Fortunately I'd seen this story on my Ubuntu box before leaving home, so I wasn't totally out of the loop.
Works fine on chrome for mac, doesn't work on chrome for windows.
(the Fraktur variant is awesome btw, and is apparently in the valid unicode range for Java...)
Personally I find it annoying how mathematical notation seems so intractable today. Things that are easily understood in code for me are a mystery in math notation. But I guess there will never be an overhaul with a more intuitive typography...
I haven't finished the book (turns out I know less calculus than I thought), but the result is pretty effective. You're much less likely to get confused about which things are numbers and which are functions, and which of those functions operate on numbers and which ones operate on other functions, once you see the Scheme implementation of something.
http://mitpress.mit.edu/sites/default/files/titles/content/s...
According to my generalization of some advice from Knuth:[1] in a good math text, definitions of terms are presented as they go along, and they are explicit about what means what. Furthermore, one of the factors that determines the quality of mathematical writing is
- Did you use words, especially for logical connectives, whenever you could have used words (instead of symbols) to express something?
and
> Try to state things twice, in complementary ways, especially when giving a definition. This reinforces the reader’s understanding. [...] All variables must be defined, at least informally, when they are first introduced.
This is repeated:
> Be careful to define symbols before you use them (or at least to define them very near where you use them).
There are some cases where "the general mathematical community is expected to know what you mean," like when publishing papers in some specialized field, but if you're writing a book, these rules hold quite true. Books certainly should explain their notation, especially since the general consensus for certain notations is expected to change over the decades ...
[1] http://jmlr.csail.mit.edu/reviewing-papers/knuth_mathematica...
[1] http://en.wikipedia.org/wiki/L%C3%B6b's_theorem#Modal_Proof_...
[2] http://lesswrong.com/lw/l0d/a_proof_of_l%C3%B6bs_theorem_in_...
I'd love to see more of APL (and a "larger" set of APL functions, actually) in use. The idea of a notation we could run directly is/was awesome.
The art of typography and signage really only matured in the 20th century, and I'm certain some of the symbols would look very different if they were designed today. Anything that helps with teaching math and making it appear friendlier is a plus, imho.
Math symbols are more or less a universal language. Once you know how the symbol appeared, or get used to "reading it right" they are totally natural. I don't see ∂ as a "weird d," I read this as "partial." It wasn't natural at first, but I got used to it, just like I got used to English.
For the page, that's fairly obvious when you look at the pseudoalphabet converters.
Scribbling something resembling the latin capital letter A returns for example any of these codepoints: A𝘈ΑАÅ𝖠∆ДΔ𝐴𝟺дᎪߡ𝛢Å4𝛥ᴬᐃⵠ𐌀𝘼𝛬Λ△𝟦Ą𝜟𝓐⌓⧍ᗋ🜂Ⲇ🗻🍙ⲇѦᗩᗅ
http://shapecatcher.com/ (https://news.ycombinator.com/item?id=5150107)
Also the Unicode Consortium has some reports on security:
http://www.unicode.org/reports/tr36/
http://www.unicode.org/reports/tr39/
listing all kind of spoofing methods you haven even thought of.
I proposed that we should name him after the lack of unicode support in our browsers, and we ended up calling him "Box Boxbox" for a couple of months.
http://en.wikipedia.org/wiki/Mathematical_Alphanumeric_Symbo...
Don't be fool as I was! Had I manually transcribed a sentence into Google instead of copying + pasting the Unicode chars, I would have found hundreds of copies of the same article.
Note: The number of іllэБіъlэVаѓіаъlэИамэѕ [2] used in your production code is inversely proportional to the number of friends you'll make in the maintenance team.
[0] https://mathiasbynens.be/notes/javascript-identifiers
[1] https://mothereff.in/js-variables#h%C3%A1%C4%87%E1%B8%B1%C3%...
[2] http://www.panix.com/~eli/unicode/convert.cgi?text=illegible...
For instance, á into a, ñ into n, å into a, etc.
Had my hopes up when I saw the title.
Does anyone have any ideas or links to working scripts that I can turn into something useful? I need to "sanitize" a database of foreign documentaries before uploading to YouTube (their metadata input system chokes on extended chars). Thanks!
First is phonetic similarity. This is mostly just to allow users to be able to understand each other and to help automatically catch alternate latinizations so you find out "Hey, he already registered under a latinized-spelling name".
The second is glyph similarity. This is the security concern where you have two glyphs that are graphically similar but phonetically completely different, but can easily be mistaken for each other. These glyphs are used to trick and confuse users. The first kind of check won't catch these, but they're the reason we don't have unicode in domain names.
Probably a correct system would have a very liberal interpretation of glyph similarity and would treat strings as matched when they contain similar glyphs.
https://github.com/iki/unidecode
Originally Perl, there are ports for python, node, ruby,.Net, etc
Obviously it's imperfect and lossy, but it might be what you want
Free and ad-free, just a fun project:
https://itunes.apple.com/us/app/texting-upside-down-free/id4...
Having a drop-down for variables certainly isn't a solution, granted. Hopefully, there are some more sensible compromises - e.g. being able to specify a locale-dependent subset of unicode in your personal environment, appropriate use of metadata to describe the language of a file, etc.
(image is safe for work, though other stuff on imgur.com is likely not)
Also some of the menu of glyphs are only visual analogues, not 1-1 replacements.
Plenty of systems and indexers will not be sophisticated enough to cope.
This may be very frustrating if you have visual impairments and need a screen reader.
Did some automated or administrative process mutate the characters? Or is this just Firefox drifting, in choice of font?
Negative Circled
Squared
Negative Squared
Double-struck
Bold
Bold italic
Bold script
Fraktur
At least not with the fonts I have. Negative Circled
Squared
Negative SquaredThere's this great quote that anything that was fun when you were five is still fun when you're thirty five, and playing around with funky letters was certainly fun at the age of 5.
(And I had fun too!)
I wrote a similar tool that does this (http://lunicode.com). It's on Github if you want to use the code: https://github.com/combatwombat/Lunicode.js
When I paste from microsoft documents into putty, characters will often be transformed to weird versions. Example - emdash is a different character to '-'. It comes through as a weird tilda character instead of a dash. Mmm. Frustating.
Is there a robust program you can run on putty to catch such type and flatten it to ascii?
Which is unexpectedly common as MySQL's "utf8" can't handle codepoints outside the BMP and will just truncate text at the first astral codepoint[0]. You need MySQL 5.5.3 (because adding a whole new encoding in a minor version makes perfect sense) and "utf8mb4" (because why would a codec called "utf8" actually do UTF8?). And then the regex are probably broken because it's PHP and developers use neither UNICODE mode nor properties (PCRE's "\w" will not match all unicode letters, you need "\p{L}" for that, also note that e.g. "🆄" is a symbol not a letter, although "𝔹" is a letter)
This can have some strange effects if you try to use them like letters. Example: What’s the lowercase transform of 𝑼? 𝑼! Not 𝒖.
Someone submitting a path to an open-source program (in Ruby) with a NBSP somewhere that changes the program logic or something. (a<NBSP>or<NBSP>b, where earlier you did a<NBSP>or<NBSP>b=x, or something similar, is the first example that comes to mind.
Z̡̖̥̙̱͓A̶͚̬̺L̷͖͓Ģ͕O̳̮!̗
Same thing with the Borat DVD cover.
This doesn't feel like the future.
𝑼𝒏𝒊𝒄𝒐𝒅𝒆 𝑻𝒆𝒙𝒕 𝑪𝒐𝒏𝒗𝒆𝒓𝒕𝒆𝒓
comes in a fancy bold italic font in my HN list. I love this hack.
I still don't know what the sequence is though, any Unicode expert to explain? Apparently is d835 "invalid"?
http://www.charbase.com/d835-unicode-invalid-character
Edit: I see now emillon explains:
"U+1D48F whose UTF-16 BE encoding is d8 35 dc 8f."
That's:
"MATHEMATICAL BOLD ITALIC SMALL N"
Wonderful :/
üníto߀ɭs ìѕ ϻùcհ ƃettër!!
Unicode's name for 𝒉 explains it all, really.
I suspect that if you could find a couple of mainstream publishers in Taiwan or Japan that prefer to print the names of mainland Chinese using the same glyphs as are used on mainland China instead of the glypths used on Taiwan or in Japan, you might be able to reopen the discussion of han unification.
Now wouldn't that be fun: "When history textbooks coverthe civil war in 1927-50, they shall use traditional Chinese for the names of then KMT-held cities and simplified Chinese for the names of then communist-held cities."
inception