Remind HN: Unicode hacks
[1] http://www.cs.tut.fi/~jkorpela/chars/spaces.html
[1] http://www.cs.tut.fi/~jkorpela/chars/spaces.html
The answer to these problems is whitelist filtering and neutralization; if a character isn't known-safe, substitute its HTML entity alternative. If you're writing blacklist filters that need to know what spaces are, you're already playing to lose.
With bind parameters you can pass data out of band, and the DB engine never tries to parse it as SQL.
Be careful.
In any case: don't use Windows on a server :-)
For those who are wondering, you can type Unicode codes directly from your keyboard (Ubuntu: Ctrl-Shift-u, other OS: http://en.wikipedia.org/wiki/Unicode_input)
http://en.wikipedia.org/wiki/List_of_precomposed_Latin_chara...
Unfortunately the amount of ligatures is small but it might come in handy.
Same with 0173, although mine seems to produce nothing, whereas yours is a line break (I think?)
>>> print("\ufeff#")
#
>>> print(len("\ufeff#".strip()))
2http://www.cs.tut.fi/~jkorpela/shy.html
While it makes sense to make usernames consist of only alphanumeric characters (something people are used to - thanks to email), I don't see why you'd need to "sanitize" Unicode characters.