Things that vanish on a printout should not be in Unicode.
Remove them from Unicode.
Things that vanish on a printout should not be in Unicode.
Remove them from Unicode.
Unicode needs tab, space, form feed, and carriage return.
Unicode needs U+200E LEFT-TO-RIGHT MARK and U+200F RIGHT-TO-LEFT MARK to switch between left-to-right and right-to-left languages.
Unicode needs U+115F HANGUL CHOSEONG FILLER and U+1160 HANGUL JUNGSEONG FILLER to typeset Korean.
Unicode needs U+200C ZERO WIDTH NON-JOINER to encode that two characters should not be connected by a ligature.
Unicode needs U+200B ZERO WIDTH SPACE to indicate a word break opportunity without actually inserting a visible space.
Unicode needs MONGOLIAN FREE VARIATION SELECTORs to encode the traditional Mongolian alphabet.
I should be able to use Ü as a cursed smiley in text, and many more writing systems supported by Unicode support even more funny things. That's a good thing.
On the other hand, if technical and display file names (to GUI users) were separate, my need for crazy characters in file names, code bases and such are very limited. Lower ASCII for actual file names consumed by technical people is sufficient to me.
Sure, but more crazy stuff gets added all the time.
Rule of thumb: two Unicode sequences that look identical when printed should consist of the same code points.
And, for example, Greek words containing this letter should be encoded with a mix of Latin and Greek characters?
Yes. Unicode should not be about semantic meaning, it should be about the visual. Like text in a book.
> And, for example, Greek words containing this letter should be encoded with a mix of Latin and Greek characters?
Yup. Consider a printed book. How can you tell if a letter is a Greek letter or a Latin letter?
Those Unicode homonyms are a solution looking for a problem.
I can absolutely tell Cyrillic k from the lating к and latin u from the Cyrillic и.
>should not be about semantic meaning,
It's always better to be able to preserve more information in a text and not less.
They look visually distinct to me. I don't get your point.
> It's always better to be able to preserve more information in a text and not less.
Text should not lose information by printing it and then OCR'ing it.
And that's where it went off the rails into lala land. 'a' can have all kinds of distinct meanings. How are you going to make that work? It's hopeless.
Tell me what the problem is and what your proposed solution would be.
a) it's a bullet point
b) a+b means a is a variable
c) apple means a means the sound "aaaah"
d) ape means a means the sound "aye"
e) 0xa means a means "10"
f) "a" on my test paper means I did well on it
g) grade "a" means I bought the good bolts
h) "achtung" means it's a German "a"
I didn't need 8 different Unicode characters. And so on.Do you think 1, l and I should be encoded as the same character, or does this logic only extend to characters pesky foreigners use.
And what about the round-trip rule?
And ligatures? Aren't those a semantic distinction?
That's a problem with the fonts.
> And what about the round-trip rule?
Print Unicode on paper, then ocr it, and you'll get different Unicode. Oh, and normalization.
> ligatures
Generally an issue with rendering.
> semantic distinction
Unicode isn't about semantics (or shouldn't be). Consider 'a'. It's used for all kinds of meanings.
While at it we could also unify I, | and l. It's too confusing sometimes.
They render differently, so it's not a problem.
Some middle ground so that you can use greek letters in Julia might be nice as well.
But I don't see any purpose in using the Personal Use Areas (PUA) in programming.
The switch in text direction has resulted in malicious code injection attacks, as the reversed text becomes invisible. I had to change my compiler to reject those Unicode characters for that reason. It can be used in other cases to have hidden, malicious text.
Have you checked your SQL code for invisible backwards text that injects malware?
How would that work with Text-To-Speech output?
1. Tell the TTS program that the text is RTOL.
2. If the TTS program can speak Arabic, it can detect RTOL Arabic text.
The only purpose for RTOL English I can think of is to insert hidden text for malicious purposes.
Do you honestly think this is a workable solution?
Also this attack doesnt seem to use invisible characters just characters that dont have an assigned meaning.