Why didn't they simly replace the original bad one?
> nine hundred pages. Imagine tracking down a single character without a page reference
Not that hard to imagine, OCR existed back then?
Why didn't they simly replace the original bad one?
> nine hundred pages. Imagine tracking down a single character without a page reference
Not that hard to imagine, OCR existed back then?
Starting with version 2.1 they made it so that code points can't ever be unassigned. People were getting really upset with the removals, since they can result in old data being misinterpreted. (Er, nothing actually was removed in 2.1, but it probably didn't become official policy until 3.0… ?)
Besides, it's useful to have the bad characters in there for people who want to write posts like this one discussing them.
... And they can, since a code point was assigned for it; so what's the problem?
> was not added to JIS or Unicode until much later
then TFA is simply inaccurate; 𡚴 has been in Unicode since 2001. More importantly, replacing the incorrect character (instead of supplementing it) wouldn't realistically have made it possible to add it sooner. They didn't really know what they were doing back at the start.
How do you train OCR on a character that doesn't exist though? Even these days, OCR often confuses the 52 Latin characters, so I wouldn't expect as good of results with older technology and some 20k CJK characters.
I don't know if this is how OCR would work here, but they're made of components called radicals. It's kind of like how letters make words, but in a two dimensional layout, and one letter was wrong and a nonexistent word was made.
For example 彁 is made of parts of 引 and 歌 (the first that popped into my mind).
IIRC there's only like 200-something radicals, though some can be hard to spot due to how they overlap or nest inside each other.
[0]: Aside from the "Overview of National Administrative Districts" book mentioned in the article