Bacon Ipsum, because Unicode is hard
geertvanderploeg.com
geertvanderploeg.com
Fullwidth conversion code in ruby:
"string".tr(' !-~', "\u3000" + (0xFF01...0xFF5f).to_a.pack('U*')) Nam
Nam
(By the way, they seem to break tab-indentation for pre formatted text in the markdown.)Since they are only used in CJK languages the majority of programmers are unaware of them. The seperate code points between half-width and full-width mean that you need a decent unicode library. Otherwise a user could spoof another user.
One cool feature is it lets you write monospaced even if you do not control the font. Since there are full-width codepoints for most programming symbols.
I'd rather see a website with fixed text that shows some often used Unicode characters, The one you can validate (i.e. shows exactly the same on your website). Adding a JPG or PNG picture of what it should look like would be a plus. By "often used" I mean: Latin with accents, Cyrillic, Greek, Hebrew, Arabic, Chinese, Japanese, ... etc. I'm sure a nice "Bacon Ipsum" with one paragraph from each of these would be quite usable. Although, some are written right-to-left, so maybe two different sets would be better: one for RTL and other for LTR languages.
You could even add options for "badly behaved" UTF8 (eg. overlong encodings, deliberate sync errors etc.) to stress test the other end.
Try it and see. Copy the output and paste it into a text document and save as UTF-8. In chrome, open that document and, under the View menu, choose a different encoding.
For my testing I also go to the wikipedia home page and copy the text in the middle of the page which lists how many articles there are in various languages. This is great because it uses a wide variety of code points, including ones greater than 0xffff.