Shapecatcher: Draw the Unicode character you want
shapecatcher.com
shapecatcher.com
But that aside, this looks like a neat idea. Not something I have any immediate use for myself, but could certainly be useful in some situations.
[1] http://www.fileformat.info/info/unicode/char/2603/index.htm
Major bug.
EDIT: As this is a deep topic, there are also books if that's more your style: http://www.amazon.com/dp/0123725380/?tag=stackoverfl08-20
or maybe you like Wikipedia (Ol' Trusty): http://en.wikipedia.org/wiki/Feature_detection_(computer_vis...
There is a whole chapter on shape contexts in it, which I use with shapecatcher, too.
Get the Japanese support in there and it will be amazing. What about using MS Mincho or MS Gothic for that? (It is free as in beer, but is the licensing off?)
http://i.imgur.com/NLkl75J.png
the unicode block containing "fried shrimp" 0x1f364 -- why does this exist???
http://www.unicode.org/charts/PDF/U1F300.pdf
Latest Abstruse Goose comic sums up my emotional response:
There's lot of other strange Unicode too. There's things like '' (U2062 INVISIBLE TIMES), ⓞⓓⓓ ©ⓗⓐⓡⓐ©ⓣⓔⓡ ⓢⓔⓣⓢ and sɹǝʇɔɐɹɐɥɔ uʍop ǝpısdn.
All of which can be used to bypass filters and generally cause browser-crashing havoc. For example, this address looks like Google, but it really links to hacker news.
Sounds like a Muse song title
I just discovered the OS X terminal renders the friend shrimp glyph in color!
edit: and it's now my $PS1.
(yes, you can copy the character between "" and paste in your OSX terminal)
http://i.imgur.com/RKTOBKe.png
edit: It appears (?) stock linux fonts don't include emoji.
(!) -- 8th suggestion, after three apparently identical "upside-down 'i'" (including an "upside-down capital 'I' with a dot underneath")
(@) -- 1st suggestion
(#) -- not suggested. Top suggestion is "capital 'H' with stroke"
($) -- not suggested, although the 14th is an indistinguishable glyph "Canadian syllabics carrier sh" 0x165a, a phoenetic symbol for representing a Canadian aboriginal language.
(%) -- 1st suggestion
(^) -- 10th suggestion, to be fair this one is impossible
(&) -- not suggested (fried shrimp lol)
(*) -- 3rd suggestion
(?) -- 1st suggestion
(∫) -- 1st suggestion
(∂) -- 4th suggestion (1st one is same thing in boldface)
----------
This is suggests a really easy way to greatly improve the results: weight them by a prior probability (i.e., frequency of occurrence in a letter count). The OCR itself seems pretty good. Common math symbols are more likely than silly shrimps. Glyphs from common languages are more probable than esoteric ones, and real languages more than constructed ones. Plain glyphs are more common than variants (bold/italic) -- and these should be grouped together anyway. fried shrimp horseshoe.
edit: just read your comment elsewhere on the page indicating that you do something like what I described.
A softbank-based Emoji set was imported (in an extended form) into Unicode 6.0: http://en.wikipedia.org/wiki/Emoji#Emoji_in_the_Unicode_stan...
Because Emoji originated in japan messaging, a number of them relate to japanese culture: foodstuff (not just fried shrimp but rice balls, dango, oden, fugu), cultural practices (kadomatsu, hinamatsuri, koinobori, Fūrin wind chimes) and other such things which may be present in other cultures but usually not as prominently (e.g. Unicode Love Hotel)
However, I now have a feeling that's not what that feature is for.
EDIT: forgot to link screenshot: http://imgur.com/LyaXy4h
I mean, I did found it funny, but I did not enjoyed seeing that in my work (specially because I work with kids stuff)
So far I haven't tried Shapecatcher a lot but I think that http://detexify.kirelabs.org/classify.html works much better. Detexify is of course only for LaTeX symbols and doesn't do unicode.
To be useful Shapecatcher needs to become better at recognition.
could be awesome for a character-based PacMan impl.
Are there not some CJK (or otherwise) fonts from, for example, Linux distributions that could have been used?
Or perhaps the emphasis could be on clarifying what is meant by "good" that deserves excluding such a large and useful character space for this type of application?
Edit, I just drew the ř larger and it recognised it correctly. Cool :)
I have learnt that Unicode contains even more weirdness than I thought before though, including 'Alchemical symbol for borax-3' (U+1f744), and 'doughnut' (U+1f369).
Clearly it's not impressed by my drawing skills.
Seriously, it even showed some very PI-like things, but not PI itself. This is a downer.
A version of this for kanji that is very accurate
http://i45.tinypic.com/3535j4x.png — screenshot of variant with three diamonds.
Also it seems to fail on badly-drawn birds http://i.imgur.com/IZgrRkq.png
On the upside, I learned that there is a Panda Face unicode character.
However how would stroke order work for a Unicode codepoint? If it exists, there has to be a lot of info in addition to the codepoint of, say, a kanji.
http://www.hnsearch.com/search#request/submissions&q=Sha...
Maybe the author can write a post detector next.