Show HN: I trained a neural network to write Kanji
otoro.net
otoro.net
This is quite 'clever'.
I don't understand the example shapes at the beginning. They're not correct strokes. How does that work?
The about page has some neat made up characters.
But after trying a few strokes, and doing so more carefully, it seems if you put in a clear radical, the character is well formed, kinda; if you put in a squiggle, all you get is a doodle... that makes sense.
Inputting 口 or 艹 for example, vs a random squiggle. Take care to make it reasonable accurate.
About characters, incase anyone doesn't know: A character is basically a 2x2 grid where 4 radicals get placed (there are (about) 201 radicals in modern Chinese, Japanese kanji too I guess?). Sometimes 'cells get merged' so the left column of 2 rows is merged to contain 1 radical, and the right contains 1 or 2 radicals. Or 'add a row' can happen at the top, for example adding a 艹 above the 2x2.
N.B.: I am not implying that it is incorrect - I am a dabbler in Japanese Brush Calligraphy (without any fluency in the language, only as an art) and I would like to read more about this so if you have links or books (preferably in English) I would be very happy to learn more.
The lines are meant more to help relative placement than to exactly divide characters, but many characters do divide into left-and-right or top-and-bottom portions -- see Jack Halpern's SKIP system for example:
https://en.wikipedia.org/wiki/Kodansha_Kanji_Learner%27s_Dic...
Note that this structure can also apply recursively, further subdividing an area.
A character is a combination of radicals, of which there aren't that many.
You're correct, obviously there are more than sticking a radical in a 2x2 grid, however the square framework is always there and emphasises symmetery.
It's "only" 9.3 GB compressed; 13.47 GB uncompressed.
https://drive.google.com/open?id=18I-5wU54CG1lty0udBOpptAG49...
Soon I'll write an article explaining how I made it, and then try experimenting with TensorFlow.
The dataset used to train the network had to be refined a bit as well to match how humans write on a tablet.
Some prev discussion a while back for the original non-interactive TensorFlow version:
Btw this is the dataset used for training. It is a part of a larger open source project used for educational purposes, that might help you learn:
I made https://pingtype.github.io - a program to break up sentences into words, pinyin and parallel translation, and typing characters by breaking them into glyphs.
I also just finished making a large dataset of glyph images of 52,000 characters from 1200 fonts - see my other comment for the download link.
It seems to just write random kanji based on the last stroke or something. But, writing random kanji has a certain coolness factor. Especially if I could copy-paste a kanji and get it to write it for me so that I know what the stroke order is supposed to be.
Some out-of-order inference would be cool. E.g. draw the bottom four dots (fire) of 煎, and have various top parts emerge. For that purpose, it would be good if there were a reference frame. That is to say: an underlying square box to serve as a target for the supplied input. If you draw something near the bottom of the empty box, then it's understood by the neural network to be a bottom part of the kanji requiring a top. I think the whole concept could really benefit from a precise agreement between the user and the neural net about the bounding box.
Can you please input things like 感覚的 and 意識 and 私は部屋 and 自殺 and 何私 to freak some people out?
TensorFlow.js only came out this year and the interactive sketch-rnn JavaScript browser demo that this was based off of is also quite recent.
There is some work in applying an NN to make the result look more like "realistic" code samples: http://www.cs.unm.edu/~eschulte/data/katz-saner-2018-preprin...
Aren't they allowed to obfuscate they're IP?
I think the reason that you believe the JS code is obfuscated is because the part of the code that contains the “weights” of the neural network, which contains 4-5 million floating point numbers of an LSTM recurrent neural network.
In fact I trained the neural network using the open source version of Sketch-RNN and encoded the weights using base64 to save you some bandwidth (https://github.com/tensorflow/magenta/tree/master/magenta/mo...)
Welcome to “Software 2.0”, I guess!