Naughty Strings: A list of strings likely to cause issues as user-input data
github.com
github.com
I apologize for the lack of updates to the BLNS. (since I'm free today and this is on the HN front page, I'll do a cleanup pass).
Even though it's a GitHub repository with 12.3k stars, there's not much to say or improve on what is effectively a .txt file based around a good idea (I recently removed mentions of my maintainership of the BLNS from my resume for that reason, despite its crazy popularity).
I happened across it this afternoon and thought it was great!
Do you know of any automation around this? I was thinking of a script that grabbed your list and then hammered a given input filtering library would be awesome. It's not something you'd want to run all the time but pre-major release, it could useful.
https://github.com/mirrorer/afl/tree/master/dictionaries
Might be interesting to make an AFL dictionary with BLNS strings, or mine the AFL dictionaries to improve BLNS... :)
Basic starting point:
Rails.application.routes.routes.collect { |route|
route.path.spec.to_s if route.verb == "GET"
}.compactRecently used Burp Suite's Intruder which can take a text file and do the Fuzzing.
My Gmail username is the same as my HN username if you'd like to speak. Thanks.
Also it bugs me a bit that the "Scunthorpe problem" section seems to be in random order not alphabetic.
Turns out they were using a synchronous HTTP request with NO timeout, and their intrusion detection system was blackholing any request that contained 'cgi-bin' anywhere in the headers or body.
For example (I trust HN is suitably hardened :-) :
/dev/null; touch /tmp/blns.fail ; echo
1;DROP TABLE users
Edit: PS: "Feel free to send a pull request to add more strings, or additional sections."(Sure, people generally shouldn't use this test input outside of a discardable testing environment, but if we could rely on "People shouldn't..." clauses to govern behaviour then much of this list would be unnecessary anyway.)
> "Dropping a table is like checking a gun without bullets. It should not work, but just don't put it against your head while testing."
HN doesn't use a database.
It's written in Arc though, so it may take some effort to read.
https://www.obscure.org/javascript/archives/msg01369.html
https://www.cnet.com/news/yahoo-mail-puts-words-in-your-mout...
> I agree that another SQL injection should be included - not because the vulnerabilities exposed by this file should be tempered (as that would only be to assist a dangerous confusion of responsible practices), but because "DROP TABLES" is such a cliche in infosec that it's prone to be caught by extremely crude filters, naive to the degree that it's the only class of SQL injection they know to avoid.
It's a nice collection of text snippets to test against many systems
"# Human injection # # Strings which may cause human to reinterpret worldview
If you're reading this, you've been in a coma for almost 20 years now. We're trying a new technique. We don't know where this message will end up in your dream, but we hope it works. Please wake up, we miss you."
"Hey can you reset my Jira login. I can't get in. It says my account is locked. I am working from home so send it to dave@mydomain.com. Thanks Dave Smith"
I expect some nasty strings to contain newlines (I wonder how many bash scripts are sensitive to filenames with newline characters in them). It shouldn't be a problem with the json file though.
However, I also imagine how such a list could be misused to actually decrease the security of a system:
Imagine this list is handled the same way as virus signatures in so-called anti-virus software. Instead of properly handling user input, an application would check against this list and call itself "secure". Maybe with with partial and/or fuzzy comparison. If you demonstrate that this approach is deeply flawed by showing another unsafe input, they'd simply add that to the list and call themselves "secured" against this attack.
I'm pretty sure I've used it to "end" a hashtag early, like in this made-up example:
I've eaten two #banana<ZWS>s today!
In my language, the possessive form doesn't take an apostrophe ("Alices Adventures" instead of "Alice's"), so for hashtags and user names it can be desirable to use the ZWS as an invisible apostrophe. > Its quite common in my
> language
I'd love to hear more details on why?https://en.wikipedia.org/wiki/Zero-width_non-joiner
* It's entirely possible that the browser you're using isn't doing a very good job with ligatures which explains the strange look of my examples
ZW[N]J as a standalone character or at the beginning of a word is very unusual on a day-to-day basis, so it's understandable that Twitter fails to recognize this pattern.
> When a ZWJ is placed between two
> emoji characters, it can also result
> in a new form being shown, such as
> the family emoji, made up of two adult
> emoji and one or two child emoji
That makes a lot of sense too, and I hadn't put sufficient work into how that's implemented -- retrospectively that makes perfect sense.Same for ️ (male version of raise hand). On phones without the Emoji, it's just "male emoji" and "female raise hand emoji".
/e: oh, HN is stripping Emojis
> Strings which punish the fools who use cat/type on this file
https://github.com/minimaxir/big-list-of-naughty-strings/blo...
# Strings which may cause human to reinterpret worldview
https://github.com/minimaxir/big-list-of-naughty-strings/blo...Thank you @minimaxir, I hadn't seen this before, this looks very useful.
(For any who want to take the blue pill: https://github.com/minimaxir/big-list-of-naughty-strings/blo...)
Drives me up the wall, i didn't have time to go deep into this.
isn't that quite a blind spot?
Native Australians were angry, that FB blocked their real names, because they seemed fake to them.
They have last names like "Creepingbear" and such.
(https://github.com/minimaxir/big-list-of-naughty-strings/blo...)
Innocuous strings which may be blocked by profanity filters (https://en.wikipedia.org/wiki/Scunthorpe_problem)
("Country" is not obscene, but Shakespeare makes "country matters" into an obscene reference in Hamlet. There are a lot of innuendoes in the classics)
Some of the best teaching of Shakespeare I've seen used the actual lines mixed with a little extemporaneity to better get the intent across. "Nay, gentle Romeo, we must have you dance. Come on, stop being so emo! There are like a million other girls out there."
Wikipedia [2] attributes it to a Yahoo email filter "which automatically replaced Javascript-related strings with alternate versions, to prevent the possibility of cross-site scripting in HTML email".
[1] https://github.com/minimaxir/big-list-of-naughty-strings/com...
[0] https://www.newscientist.com//article/dn2546-email-security-...
"In February 2006, Linda Callahan, a resident of Ashfield, Massachusetts, was initially prevented from registering her name with Yahoo! as an e-mail address as it contained the substring allah. Yahoo! later reversed the ban."
[0] https://discussions.apple.com/thread/1491462?start=10&tstart...
evaluate
mocha
expression eval
expr
I got nothing for "mocha", though. Edit: apparently (from below) there was a Yahoo! mail filter that replaced "expresso" [sic] with "mocha"; but either the story was misreported or the mail filter was wildly misconfigured. So the entry should be "expresso" [sic], perhaps.are there any drawbacks to this that i can't think of ?
in terms of perfomance - i guess it could be somehow optimized (with dictionary and sorting algorithms etc etc)
edit: newlines
is it wise to just take this list "as is" as a black list for, say, valid usernames?
I interpret this as a list of input that you should accept, and it's test-data to verify that the input is correctly handled.After all, I imagine Linda Callahan would be upset if she couldn't use her name when registering, especially if she couldn't flip a table in comments afterwards. (╯°□°)╯︵ ┻━┻)
But the issues they might cause are not all malicious: some are people's names, added to the list because an over-zealous profanity or offensiveness filter once choked on them.
So my suggestion is that you shouldn't block any of the strings in this file, but should use the file to make sure that your code works successfully when any of the strings are given. Where "successful" is naturally dependent on context: you may have a policy in place that says that messages may not consist solely of whitespace, so the correct response to receiving any of the whitespace strings is to return the correct error to let the user know that, avoiding Twitter's example of an internal server error in that case.
Hah. Totally filtering for that one now.
On that note, can anyone suggest how one could efficiently test that an RTL unicode char doesn't "infect" the whole following content of a template?
What would the user expect from inputting "U+200B ZERO WIDTH SPACE" into a form, anyway?
Let's try it on Facebook. Here's what happens when you put only a blank or space into a post and try to submit: http://i.imgur.com/bNtgky8.png
Here's what happens when you put a zero width space and try to submit: http://i.imgur.com/NMgyZqc.png
But yeah like others said I would expect this to turn into some sort of validation message on the client and never show them the backend error.
Also, the first time that copypasta actually spooked me out ;-)
It's wrong because it de-emphasizes the importance of HTML-aware template languages, such as some that are available for golang, or SGML, the natural template language for HTML. There's no such thing as a collection of regexpes for sanitizing HTML; it all depends on the context into which strings are inserted.
I think it's also good in that while you may not know all the latest tricks, this can help you reveal what you don't know. It can get you really thinking about the possibilities of what a simple string can do to your code if not properly handled.
Also, explicit HTML (or SQL or whatever) string handling in normal application code is just a failure to separate concerns: you haven't distinguished the level at which HTML has an abstract syntax and the level at which HTML's abstract syntax is linearized into strings in one particular way.
I do know to consider things like sql injection and having js injected into the site. But I don't know what a special white space character from a Persian alphabet will do to my server. Until today I haven't actually thought about it. Not every language handles strings the same, as you pointed out.
I still think it's good to have around for helping you reveal what you don't know, about what you don't know.
---
Replying as an edit, because HN complains that “I'm submitting too fast”:
Sure, what you said applies to entire applications. But something relatively stable and small, like, um, the definitions of HTML, JSON, SQL, etc. (do they become larger every time your boss requests a new feature?) surely should have formal specifications.
Alas, I don't work at NASA where these formalities exist. I'm given a rough sketch that I'm expected to bring into life, throw away and recreate again on a whim.
Please note that I am not complaining, nor excusing. Only pointing out that our expectations, environments, and programming languages are different. Each can massively affect how the program should handle the input. Adding checks helps, but does not mitigate the need for a nice set of test data to help verify everything runs the way we expect it to behave.
Exactly. Security check lists become unnecessary when the program is designed to be correct right from the start.
> (first paragraph)
The real problem is that we do a very poor job of embedding languages inside each other. For example, HTML parsers must contain special provisions to handle that embedded JavaScript. </tag> might no longer be a terminator, because it could appear inside a JavaScript string literal. This is terrible design! I don't even like Lisp, but Lispers do have a point when they say using S-expressions would avoid all of these issues.
“But writing parsers was sooo boring in college, and who has to do this in real life?”
Yes, writing parsers is a lot easier than those examples. But so far society has always ruled that inputs that purposefully try to abuse flaws are not freed from responsibility just because the flaw shouldn't be there.
I'm not a malicious person. I don't purposefully abuse any system's flaws. But if anyone else does, my sympathies aren't with the designer of the flawed system.
P.D.: Appeals to authority won't help make your case.
These so-called “naughty strings” expose implementation errors in code that processes widely used formal languages such as numeric literals, URLs, HTML, JSON, SQL, etc. These languages are so widely used that it's criminal not to have formal specifications for them. And the mathematical techniques for constructing programs that meet their formal specifications are very well known.
(1) Design implementations (parsers, code generators, etc.) with the logical argument for their correctness in mind. That is, don't attempt to verify an existing possibly incorrect program - write it to be correct right from the beginning! This is greatly aided by designs that meet criterion (0).
These are corner cases in the concept of user input, not just corner cases of any specific parser. What if it's a number, what if it's not? What if it's the same alphabet as the code, what if it's not? What if it is valid code? What if it's empty, what if it's not? etc. Even if you've written the perfect parser in the perfect language, you still need to have unit tests for all of this stuff. They are traps caused by human definitions of "input" and "string", which cannot be formally verified.
No. I'm saying that programs have to be proven correct. Then you can use tests to rule out other pesky problems that have nothing to do with your design being incorrect. (For example, you could prove a program correct on paper, then transcribe it incorrectly to a computer. It has happened to me before.)
> These are corner cases in the concept of user input
“undefined” and “null” aren't special cases in the concept of user input - they're special cases in languages that happen to have “undefined” and “null”.
Octal numeric literals aren't special cases in the concept of number - they're special cases in languages where octal literals begin with the prefix “0”, rather than something more sensible like “0o”.
Failing to distinguish between escaped and unescaped strings is also a language problem - they should have different types!
The list goes on.