How ASCII lost and unicode won
blog.goosoftware.co.uk
blog.goosoftware.co.uk
Unicode is the replacement, not the competitor, like 64-bit IP addresses are the replacement for 32-bit IP addresses. It was developed in the early 1990s when RAM got cheap enough that you could afford two-bytes per character.
Personally, I deal with data all the time and rarely encounter unicode. Of course, I'm in the US dealing with big files out of financial and marketing databases. In fact, I've seen more EBCDIC than UNICODE.
IPv6 addresses are 128-bit.
ASCII is an encoding which can represent a very small subset of Unicode.
UTF8 is an encoding which can represent all of Unicode and it is a superset of ASCII.
UTF32 can represent all of Unicode, but is not a superset of ASCII.
Therefore, something can talk Unicode via UTF32 without being able to talk ASCII.
Practically both character sets and encodings prove ample opportunities to mangle input, either by replacing unavailable characters by ? or �, or by mapping to the wrong characters because the wrong encoding was assumed (or by throwing errors because the input made no sense).
Unicode is not an extension of EDBIC or Latin1 because characters under them have different codepoints than they do under Unicode.
[Please correct me if I'm wrong. I'm not an expert but this is how it seems to me]
You're basically right, but the subtle point is that these comments are in response to lotsofcows who said, "Anything that talks Unicode can talk ASCII (although not vice versa)."
The implication of "talk" is that the numbers are being encoded. qznc was trying to clarify that lotsofcows's statement is only true if the Unicode system is using an encoding which is backwards compatible with 7-bit ASCII, such as UTF-8.
Extremely close but not quite right for some weird corner cases. There's either a cool feature or hideous bug in the unicode design, depends how you look at it, that lets you compose multiple unicode codes together into one glyph. So you can say, "Gimme the code for A with a bar over it" or "Gimme the code for A, now gimme the code for put a bar over the last glyph". Even worse you can stack compose characters, so you can write "A with circle on top and squiggly underneath" at least five different ways in binary that should be rendered visually identical.
There is the one true normalization technique that most people use to convert what they consider poor unicode grammar into a standard form. Some people don't use it of course. And anytime you give users freeform input who knows what kind of crud they'll feed you. So you can never really assume any unicode string is normalized unless you personally normalized it yourself. Even then theoretically two normalized strings, concatenated, might or might not be normalized anymore (although this is often not much of a problem). And this strikes substring manipulation too.
A fun source of buffer overflows is normalization can shrink OR EXPAND a unicode string, in theory. So if you use one of those languages without variable length strings, look out. Or if the language tries to take care of it for you, this leads to weird memory fragmentation. Maybe a DOS/DDOS opportunity?
Which leads to philosophical argument you must decide in your code, if two bitstreams don't match, but you get matching glyph renderings, from the program's point of view are they a match or not? Depends if what your program is trying to do I guess. Most languages have a library to handle this. Writing your own unicode handler functions is a "here be dragons" moment.
Composing characters are also a fun source of swear words if you're trying to count the number of characters in a file. Something that renders literally identically will have an identical number of characters but somewhat varying number of bytes.
Other fun philosophical arguments are if you read in a un-normalized string and output a normalized string have you changed the string? Well, both no, and yes. I'm sure there's crypto steagongraphy implications.
Google for "unicode normalization form nfc" and stuff like that.
The first page I found with a good faq was:
Character sets therefore could be modelled as a partial function c : Character ↛ ℕ, while the encoding is a partial function e : ℕ ↛ seq Byte. Something like that.
ASCII and Latin 1 are therefore both subsets of Unicode, even though only ASCII is a subset of UTF-8 while no comparable Unicode encoding exists for Latin 1 or any other encoding.
On the other hand software speaking UTF-32 isn't going to tolerate ASCII being shoved thru them. The output is going to be 1/4 the length you expected and probably a mass of random asian glyphs. You CAN write a shim that turns each 7 bit ascii char into a 32 bit UTF-32 char and it'll work 1:1 perfectly for all 128 characters in ASCII. But outta the box, no you have to write a shim.
Now if you REALLY want to confuse people, after they 100% understand ASCII, UTF-8, and UTF-32, then feed some UTF-16 into something designed to eat either UTF-32 or ASCII or UTF-8. If you understand what a byte order mark is, and why its important, and how to use it, then you're along the way to understanding UTF-16.
This was a common whine about unicode in the early days that all you've done is trade the agony of multiple extended ascii code pages for the agony of multiple formats to represent unicode... Should have just spec'd UTF-8 and no other encoding and been done with it... or maybe UTF-32.
I will say that a world where only UTF-32 exists would be a world with a lot more text compression of webserver responses and stuff like that. It wouldn't be an automatic end of the world.
What does that even mean? It doesn't mean DIP packaged DRAM because my dad was buying COTS Intel 1103's in 1971 or so before I was even born. And the first "I'm gonna store one bit of data in a capacitor" was done over the pond in the .uk during WWII at their code breaking plant.
"like 64-bit IP addresses are the replacement for 32-bit IP addresses."
Um...
I've worked on automated data submissions for banks in the UK, and the insistence on fixed-width, EBCDIC encoded data files for many regulatory filings (FSA, credit rating agencies) was annoying. On the other hand, it was so easy to automate in a VBA macro that I could have quite a bit of free time.
> What’s a byte?
I'm not sure what the target demographic is.
> A byte is a set of 8 bits.
Nitpicking, that's octet, not byte. I believe the distinction was lost and don't think byte has a proper definition anymore.
For all practical purposes EBCDIC is the "IBM Standard" as opposed to ASCII's "American Standard". The mood in the 70s/80s outside the business world was "In your face IBM!"
And putting "Information Interchange" in the name itself is another "In Your Face" posturing to the mainframe world. We're the future of data transmission and you'd best get used to it, IBM...
ASCII really was a rebellion in the olden times. One that won.
Another story was before ASCII note that teletype codes and such were usually modal, LTRS/FIGS to switch from 5-bit letters to 5-bit numbers. So there's that dramatic great circular wheel of IT where we've oscillated both before and after ASCII between simple encodings and modal encodings. This was an early whine against unicode, who cares about codepages, just embed it like, or in, what amounts to the MIME media type, and glyph-like Asian languages should just be drawn in gif files anyway. Or so the complaints went at that time.
Another design statement story: Kind of like the Uni- in unicode uniting all the extended code pages into one really huge space.
There are other dramatic stories not in the article. For example the Klingon in Unicode movement. Basically about 15 years ago they tried to get Klingon script into unicode, about a decade ago the unicode people (who?) said no, so the Klingon people squatted in the unicode equivalent of what networking people would call RFC1918 space, and its simmered since then. Will Klingon actually make it into unicode officially or not, who knows. You can add to the fire by pointing out that Unicode is already stuffed with scripts that no living human culture currently uses, and numerous glyph symbols (think of math stuff like + or %) On the other hand this would inevitably result in Tolkien Elvish script being included in unicode. And does this really matter one way or the other? And this maps perfectly into the wikipedia battle between the deletionist (expletive deleted) and the inclusionist saviors of humanity. Well maybe that last line was a little biased to my opinions...
There's tons of extra fun drama to tell about unicode and cutting off the story at the founding of ASCII misses some of the fun drama. Aside from some fun drama that wasn't mentioned in a post aimed at non-techs who we're told really like drama. You could probably turn the Unicode story into a trashy reality TV show somehow; Vampire Romance Fiction is going to be a harder translation although I'd love to see it.
[1] http://openlibrary.org/works/OL8019369W/Coded_Character_Sets
There is no technological reason EBCDIC couldn't have been the encoding for the whole desktop revolution, other than the dramatic central control vs local control, mainframe vs desktop thing.
There is some truth to the claim that whenever IBM reached a fork in the road they usually found a way to go both ways, at least for many years in the olden days.
Had the advantage that an open carrier (all zeros) mapped to NULL so you didn't waste paper (either tape or roll).
1. http://www.randomtechnicalstuff.blogspot.com.au/2009/05/unic...
In general it was a mistake to put variable-length encodings into the Unicode standard. A much better design would have been to use UTF-32 for the application-level interface to characters, and use a separate compression standard that is optimized for fixed alphabets when transporting or storing text. This has the advantage that the compression scheme can be dynamically updated to match the letter frequencies in the real-world text, and it logically separates the ideas of encoding and compression so that the compression container is easier to swap out. And, of course, an entire class of bugs would be eliminated from application code.
Edited first paragraph to clarify: Variable-length text encodings are the same.
If you want to get back the supposed benefits of UTF-32, you'll have to dynamically assign codepoints to grapheme clusters.
Unicode also would have been too incompatible with existing code that copied 8-bit character strings around. See http://www.cl.cam.ac.uk/~mgk25/ucs/utf-8-history.txt for some rationale behind UTF-8.
"A byte is a set of 8 bits. Computers typically move data around a byte at a time."
A byte being 8 bits is ok. Historically, a byte might have a different number of bits, but all modern architectures use 8 bits. Since this is an introductory article, this is fine. (more details: http://en.wikipedia.org/wiki/Byte)
Computers do not typically move data in byte chunks. You could say "a byte is the smallest unit of data a CPU can load or store". If you talk about moving data, the question is between what. Probably memory. However, there are caches nowadays, since bandwidth is cheap and latency is expensive. Data is moved in cache line chunks, which means 4-64 byte chunks depending on architecture and cache level. Bigger chunks in upcoming architectures.
Before that you had CDC machines with 6-bit bytes (or Nybbles if you owned an Apple ][ with a floppy disk)
If I had to pick a date, I'd pick 1977 as when it solidified. A few elements: ARPAnet standardized on octet-based protocols in the early 1970s, and was getting widespread by the mid 1970s; three popular microcomputers based on 8-bit architectures were introduced in 1977 (the Apple II, TRS-80, and Commodore PET); and DEC introduced the 32-bit VAX in 1977.
But sure enough, I was very wrong on (A): Apple II disk writes were done in "sets of 5-bit or, later, 6-bit nibbles" [1].
Thank you, chiph!
[0] http://www.nibblemagazine.com/ [1] https://en.wikipedia.org/wiki/Nibble#History
bit
nibble (4 bits)
nybble (6 bits)
byte (8 bits)
word (16 bits)
:A byte is the smallest thing that a pointer can point at. Something like MIPS (from memory) has 8-bit bytes, but a load instruction loads 4-bytes at a time (and although it has the unaligned load functions which can be convinced to load a single byte, they aren't general purpose)
So he thinks Americans are the only people to use the English language does he?
ASCII as we know it (which is essentially the 1967 version⁴) like the corresponding ECMA standard⁵ provided for overloading punctuation characters as diacritics ("/¨ ^/ˆ ~/˜ '/´ ‘/` ,/¸) to be overstruck in typewriter fashion; ECMA-35⁶ (1971⁷) defines further extension techniques using control and/or escape sequences.
So, yes, it's just a failed attempt at an anti-American cheap shot from someone who isn't familiar with the development of character set encodings.
¹ American Standard Code for Information Interchange, http://www.wps.com/projects/codes/X3.4-1963/index.html
² 7-bit Coded Character Set, http://www.ecma-international.org/publications/standards/Ecm...
³ http://www.ecma-international.org/default.htm
⁴ http://www.wps.com/J/codes/Revised-ASCII/index.html
⁵ 7-bit Input/Output Coded Character Set, 4th Edition is unfortunately the oldest available online; http://www.ecma-international.org/publications/files/ECMA-ST...
⁶ Character Code Structure and Extension Techniques, http://www.ecma-international.org/publications/standards/Ecm...
⁷ Extension of the 7-bit Coded Character Set, http://www.ecma-international.org/publications/files/ECMA-ST...
The ñ is right next to the "L" key, and the tilde (accent) is either right next to the ñ or on top of it.
This adds some sort of complexity, but in my experience, the average user simple expects the key to be "where it always was" and I've had to "fix the problem" (changing the input method) many many times.
To answer your question I think that the amount of users that actually take note of what happened and understand it enough to fix it again is minimal, the rest just expect it to work and ask for help when it doesn't.
There are also a lot of variations of spanish keyboards, so it just makes the matter more complicated... I use *unix, Windows and OSx almost interchangeably and know how to change the input language in most of them to ISO spanish spanish quite quickly, but I'm not representative in that regard.
Slightly offtopic but the spanish spanish keyboard layout is extremely comfortable for programming...
Is this really true? My impression was that UTF-32 is a fixed-length encoding which uses 32 bits to encode all of Unicode. It seems that this means that Unicode can never have more code points than could fit in 32 bits. Right?
The original limit was 6 bytes encoding 31 bits, but this was lowered to 4 bytes encoding 21 bits. If I'm not mistaken, 7 bytes encoding 36 bits (or 8 bytes encoding 42 bits if you make the 0 optional) should be the hard limit.
I.e.:
Oct.1 Oct.2 Oct.3
11111111 110DDDDD DDDDDDDD (7 octets more)
\_________/
10 bits
("D"s represent data bits)
This is what'd be necessary to use UTF-8 if we'd need to go beyond the current limits of 0x010FFFF upper bound.Also, how do you plan to handle 9 leading bits instead of the 10 from your example? That would make the second byte start with 10, the marker for continuation bytes.
In my mind that will only happen as a result of extreme carelessness.
If this were ever to become a problem (which I don't see happening any time soon), the transition from UCS-2 to UTF-16 is prior art on how to pull off the extension of a coding space.
Somewhat unrelated, but nevertheless worth mentioning as it's a common misconception: While UTF-32 is a fixed-length coding for Unicode characters, often, the more interesting unit is the grapheme cluster, effectively making UTF-32 into a variable-length coding.
a) compatibility to every pre-existing character set and
b) including historical scripts too¹
21 bits was what emerged from expanding UCS-2 to UTF-16 via surrogate pairs and Unicode was reörganised into 17 “planes”, the first of which, the BMP, containing all code points allocated so far. UTF-32 then just was a simple encoding scheme that allows to have one code point per code unit that is also efficient to process. 21- or 24-bit code units would be unwieldy on most architectures (especially regarding unaligned memory access).
___________
¹ Arguably the decision to include Emoji made a bigger dent in the code point space than hieroglyphs, Linear [AB], etc., though, but that came a litte later.
The more interesting question is if you're designing a new operating system would you pick UTF-8 or UTF-32 as the basis of your character system. Bearing in mind you need to normalise strings anyway for comparison purposes the general space efficiencies for UTF-8 for most systems seem tempting.
I don't think it's mere coincidence that the capital letters start at 65 and the lower case at 97 and the decimal digits at 48.
Chinese input methods typically require a sequence of key-presses and then a selection from a menu of matching characters, with the most common matches first. Multi-character sequences can be entered without making a choice until the end, in which case the most likely n-grams come first.
Most people in Taiwan learn phonetic (Zhuyin) in school, which is very easy, as long as you know standard pronunciation.
Other countries that make even less sense and use way too many special characters (from my blunt perspective) usually have different keyboards. Prime example that I know of is France where you have to use the shift key to make numbers.
To answer what you actually asked: Yes we are taught how to make those special characters. For me it's as logic as typing parenthesis or the euro sign, but I'm spending most of my waking hours typing one thing or another. Many people don't or barely know how to.
It's not what most of the time is spent on in those courses. We're being taught the asdfjkl; row just like everyone else. Or aoeuhtns, depending on your keyboard (in the Netherlands we only have qwerty though). Making accents is more of a side thing that's mentioned once or twice after learning everything else at an acceptable speed.
* We should|could|would|must|might|ought to|need to get rid of other overhead.
* You should|could|would|must|ought to|might get rid of other overhead.
* Those guys over there should|etc. get rid of other overhead.
* I can't type "ride". What does "...ride other overhead" mean? I dunno, but people write incoherent things all the time.
IOW, what you think of as "overhead" probably isn't, especially with natural language.
Right Alt key used as Hangul(Korean)/English toggle, right Ctrl key as Hanja.
When toggled to Hangul, only English characters are overridden by Hangul characters. All numbers, symbols are also same when you are in English typing mode.
Basically no additional key in there compared to QWERTY.
Is that complex and hard learning type in Hangul? Nope. Maybe 'Korean' is complex to learn, but 'Hangul' - I mean, script? character composition system? sort of that - is quite simple.[2]
Actually It's capable of implement more efficient input layout than English especially more restricted environment. Like basic cell phone key layout(E.161).[1]
There was a King, and He was really great hacker. Because he was a King, he grabbed bunch of smart guy all around country. ; ) Then push them working hard. (did I said he was King?) Therefore, invented many good one for country people. Today Korean has own quite good and expressive characters and he deserved quite good place.[3]
[0] http://i.imgur.com/j0Xk6oY.jpg
[2] http://blog.naver.com/PostView.nhn?blogId=neraijel&logNo=110...