*NEVER* sanitize your inputs
blog.hackensplat.com
blog.hackensplat.com
My major complaint is that after correctly identifying the solution for SQL, he ends up with nothing to say about HTML. The right approach for rendering user input into HTML is with the Javascript createTextNode() function. That's how you tell the browser that it absolutely shouldn't interpret that content as HTML.
Ugh, eyeroll. Seriously, let's waste time arguing over what to call security vulnerabilities & ways to address them - instead of using consistent terminology that security-minded developers instantly recognize.
To quote the hilarious Mean Girls - "stop trying to make fetch happen".
And when they try to "clean it up", they enter the realm of Falsehoods Programmers Believe About X.
e.g. http://www.kalzumeus.com/2010/06/17/falsehoods-programmers-b...
"Doctor, it hurts when I do this." Don't do that!
My google search "Input sanitization" yielded these first 2 results
http://en.wikipedia.org/wiki/Secure_input_and_output_handlin...
2nd page (or more with a lesser screen), under "other solutions," this is the only line about parameterization: "In particular, to prevent SQL injection, parameterized queries (also known as prepared statements and bind variables) are excellent for improving security while also improving code clarity and performance." Everything else is about filtering, blacklisting, whitelisting, escaping.
http://www.esecurityplanet.com/browser-security/prevent-web-...
Discusses filtering as solution to HTML injection. Lastly discusses SQL injection, first recommending mysql_real_escape_string(), then in the second paragraph linking to another article about parameterization.
It's not, to an inexperienced developer (this is the web remember?), a clear-cut best practice from just "cursory research". It's a popular tech joke with obvious but non-optimal solutions.
What does the inexperienced developer learn from the new search terms?
I don't know why magically using different, non-standard words would prevent a developer from being inexperienced.
"I don't know why magically using different, non-standard words would prevent a developer from being inexperienced." It's really hard to give any response to this sort of flawless logic...
"HTML injection" does sound cool though. Since XSS nowadays is not necessarily about sending cookies to another site, perhaps we could adopt "HTML injection" as a more generic term.
Now of course, the problem we're trying to fix is someone who does:
$content = htmlspecialchars(mysql_real_escape_string(addslashes($content)));
before $content ever hits the database, without any understanding of what those functions really do. It's a surprisingly common cargo cult among newbie web devs. Just throw all the security-related functions together and you'll be safe!Same principle, but different method of exploitation. If we supply plain HTML tags in a vulnerable parameter, it's HTML injection. If we use JavaScript (via a script tag or whatnot), it's XSS.
A word is nothing if not bound by a context. Developers have already developed part of this context. Design patterns names are an example of those words defined within the context. Sanitizing input is just another.
You know the plain English meaning of "Sanitize". Clearly, you need to remove those single quote characters as they are unsanitary?
Problem, your theorem is dealing with discrete numerable infinity...
On the side note, English meaning of Sanitize is "Make clean and hygienic", nothing more. It says nothing about "removing". Other definitions are extensions based on CONTEXT, once again.
"If you want to create a horizontal line in HTML, you write <hr>"
See that? There is nothing "unclean" about it, hence you should not "clean" it. You just have to encode it if you output it embedded in HTML. That's why calling it "sanitizing" is misleading.
Encoding without proper context means "convert in a coded form". Hum that's not exactly what we want. So, let's add the "computing context", now we have, as an example, the ability to encode a WAVE file into a MP3. But wait, we lost information here! Bummer...
Sanitization in the context of computing does not specifically means that you have to "encode", or better, "transcode". It means that you have to take appropriate measure so that your input DATA cannot be interpreted as CODE by the receiver. Bonus point is taken if the measure you choose is lossless in term of information carried by your data.
But no, in a way, you are getting it all backwards, or at least a bit confusing.
This is how you should construct a system that processes user input:
First, the input format should be defined such that it can only describe things that make sense within the given context, in particular it should usually not be possible to represent in it instructions for programming language interpreters.
Second, whenever you have to represent user input in some context, you have to encode (well, transcode) it into the format of that context. This transcoding generally should only change representation and not change the meaning of the converted information.
This automatically implies that you can not "inject code". There isn't really anything magic about "code". That's what I think is a large part of the confusion around "sanitizing input". The input can not represent code, the conversion does not change the meaning, so if the input can not represent code, the transcoding obviously can not cause code to appear either, and thus you are safe - and not only are you safe, but your system also works as it should otherwise, which it potentially does not if you start "removing dangerous characters".
That is why you should not "sanitize", but only validate and encode/transcode/convert. Which you need to do anyway for your system to work properly. Lack of injection vulnerabilities will result automatically.
Both SQLi and XSS have the same cause: concatenating strings when you are working with active code of some sort.
They both have the same solution: you need to know the escaping rules for the active code you are assembling.
You shouldn't be solving XSS by stripping tags (that's a great way to build a discussion forum where no-one can talk about how to use HTML) - you should be escaping user input before assembling it in to HTML.
To protect against dumb mistakes (because it's really easy to screw up just once and have a huge security hole) you should use abstractions that do this for you. If you're working with Django the ORM will do this for SQLi and auto-escaping in the template language will do this for XSS (watch out for variables you are outputting in a script tag context though).
Escaping, not sanitizing, should be the message.
People love to consider sanitizing the inputs, but how you do so doesn't depend on the inputs but on the specific usage of it - more-or-less the output of your program.
Rather than trying to think of all the ways the inputs to your program could be abused to cause abuse, I find that it is safer to start at where the output occurs - database calls, system calls, etc. The most commonly used of these (database calls, shell commands, etc) tend to have a variety of encoding capabilities to ensure that when you want to stick a string in a particular place it does exactly that regardless of whether the string came from user input or elsewhere. For example, bind parameters for databases, or proper escaping functions.
If you think about it as sanitizing input it means you tend to misplace your attention to detail and only consider the entry to your application. A single input is often used to do multiple things through a program so you cannot properly handle sanitization at input.
The real push should be for proper output encoding, not input sanitization.
Encoding a "safe" value doesn't make things any less safe. Failure to encode it, however, leaves potential holes in your application. Something may bypass input validation and be given to the database as an unsafe, unvalidated value. Usage of the value may change (new functionality using it differently, changed storage in database, etc) and in the new usage the value may not be safe.
Input validation is obviously something you want to do, but it should never be relied upon for protecting from injection attacks.
Here's the chain:
1. Get raw input.
2. Validate it (number, not number, in range, not in range?)
3. Optionally format it to canonical format (i.e. trim whitespace etc.)
... later....
4. Encode it for where you want to use it (SQL, HTML etc.).
Sometimes steps 2 and 3 are done in the opposite order, or as an atomic single operation, but point is, we have perfectly reasonable words for all that: validating, formatting, encoding.
Stopped reading here as I assumed the rest of the article was satirical
People normally dumb web vulnerabilities together. Xss and sqli especially. Preventing xss you have to sanitize. Preventing sqli you used parameterized queries.
To prevent stored xss you sanitize what you put in the database. So really... You still need to sanitize.
I've also seen people make arguments about inexperienced web programmers and how this advice can cause them to write bad code. I think the argument is bad because so many resources exist to help them. There is real code on stack overflow, w3 schools, owasp, and other blogs that can be copied and pasted in to their projects.
Assume a user uploads a TIFF file to your web application. Browsers don't understand TIFF. So, in order to display it on a web page, you convert it into a PNG. You wouldn't call that "sanitizing it for PNG" either, would you? For the same reason, you shouldn't call it "sanitizing" when you convert plain text to HTML.
For the average web Dev my approach is plenty good enough.. It's funny because your approach still requires sanitizing
So, for example, you could have a data model of "plain text field", in that case you check that the input is a valid character string (so no undefined codepoints present and, for example, no syntax errors in the UTF-8 encoding if that is what you are using). Thus you can be sure that you have only characters strings in that column of your database. Then, if you want to output one of those strings to be displayed within an HTML page, you convert it from plain text to HTML (replacing "<" with "<", "&" with "&", and so on). That way there is no XSS possible, and also, any input the user makes is displayed back exactly as they entered it.
The idea is the same. You take user input and put it in a safe format. The programmers needs may be different.
Calvin & Hobbes may have played a part in popularizing the term? http://calvinandhobbes.wikia.com/wiki/Transmogrifier
I hope this is satire, Irish didn't "make up" the letter Ó, it was the standard historical form but was converted into O' when the names were anglicized.
Frankly his advise about sanitizers seems equally suspect, I've processed a lot of complex scientific abstracts using html5lib and Bleach without any mangling like he describes. He must be using very naive sanitizers.
The problem is that you are silently changing information, and that's an absolute no-go for reliable data processing, and the cause is that people think of, say, html, as "some kind of text/strings".
HTML is a serialization of a tree, similarly, SQL is a serialization of a syntax tree ... - and if you want to add plain-text user input to such a serialized tree, you have to _convert_ it from, say, "plain text" to "HTML character data". You have to think of them as two different data types, and so when you want to use a value presented as one of the types as the other type, you don't have to "sanitize" it, even calling it "escaping" is confusing - you have to _convert_ it. And if it happens that some input can not be represented in the target type, then you have to _validate_ the input and _reject_ broken input.
This prevents our computers from helping us, even if we're using an ivory tower type system with a whizz-bang IDE, since everything's just "String" so the compiler says OK.
Here's an example of the alternative http://blog.moertel.com/posts/2006-10-18-a-type-based-soluti... (remember that most of the code there is building the libraries; using them is simple and terse).
People started using the word "sanitize" exactly because it conveys that information that "you want to treat it differently, depending of where it comes from". We also use the words "dirty" (sometimes "tainted") and "clean" conserving their usual relations to "sanitize".
Now somebody wants throw away a very concise and expressive jargon just because some people are giving bad advice on the Internet?
That you are using "dirty/tainted" and "clean" only shows how deep the confusion is. There is some justification to use those terms when talking about before and after validation, but other than that it's probably an indication of confusion (which also seems to be the common usage).
Take, for example, a general plain text field for optional free-form text. There is essentially nothing that could be validated (other than maybe that it's a valid UTF-8 string). Now, you want to generate a plain text email using the user input - how would you "sanitize" it?
There is nothing "better"/"cleaner" about any particular encoding, be it plain text, HTML, SQL, or any other, they are simply different encodings, and you have to always use the correct one, not the "best one"/"cleanest one", and you have to always know what format the data that you are processing is in so that you can convert correctly.
This jargon is not at all concise, actually (some people mean "remove 'strange' stuff/clean it", others mean "escape it", ...), and it makes you think in ways that obscure the actual problem that you are solving: Conversion between data types/data representations.
What the author actually means is the removal of apostraphes to prevent SQL injections can affect your data integrity, so paramatise your queries.
Alternatively replace single apostraphes with double apostraphes in your queries also works, but paramatising queries is a much better practise to get into.
Such as, say, comment fields. It'd be terribly restrictive for your users if they can't write about <script> tags on a technical forum without munging it.
And you're still not safe. All the characters needed for an SQL injection attack, for example, commonly occur in normal English usage. All the characters needed for XSS commonly occur too, so you'd need more restrictive filtering.
And have fun when a bug that causes your filter to be more restrictive than it should now means data is unretrievable because you've just stored the sanitised output of your buggy filter.
Once you've dealt with that, you're still facing the issue of changing filtering requirements: What is safe for HTML may not be safe for your CSV export. What is safe for your PDF generation may not be safe for your HTML generation, and vice versa. Suddenly you're asked to pass data via an API, with different expectations of what a "safe" value contains. Boom.
In other words, if you believe that what is in your database is safe from causing security problems, you've lost. You need to treat every piece of data that may possibly contain user input as a potential cause of problems whenever you output it or pass it on anywhere, whether or not you've (attempted) to validate and restrict the input.
A typical example I used to have to deal with: Mail systems. HTML that is entirely safe when downloaded and rendered by a mail client that contains the HTML in a document that is just for that one e-mail, can leak data all over the place and compromise the users account if left unfiltered when rendered on the web server. You can't insert it pre-filtered into the database without inserting the raw content too because the user may want to download it.
And because the only reasonably safe filtering method is white-listing tags and CSS due to evolving standards, you will regularly have to revise the filters and add functionality and people will be very annoyed if their e-mails still don't render correctly after you've fixed the bugs (and if you have to tighten the filters again, you don't want to have to re-filter all the data).
> If you have a user text input that only requires alphanumeric characters, space, period, and comma, then strip out any character which is not one of those things. That field is now no longer a possible source of XSS.
Except when somebody puts such "safe" string in an unquoted HTML attribute... Seriously, thinking of data as "safe" (safe to be carelessly mishandled...) is a fragile approach.
> it depends on someone else remembering to include something in their code
Get a template engine that escapes everything everywhere by default, so you won't need to remember to escape (or "sanitize"!) each thing.
If you get invalid input, you have to reject it, not just silently make it fit by changing the input (the only exception is when you can be sure that changing the data will not under any circumstances change its meaning).
2. Select Anonymous from the auth options.
3. Click Submit.
4. Laugh your heart out.
> but paramatising queries is a much better practise to get into.
The thick, thick irony of a guy who can't even follow his own advice.
I was actively trying to talk about his script example and instead I had to second-guess his parser to get past the validator (I eventually resigned and replaced < and > with [ and ]).
If you want to support some tags, have your parser be an HTML-like DSL language with those tags supported. Don't disallow perfectly good input.
Now, I haven't tried it, but I suppose his form field expects HTML syntax? Have you tried entering your text in HTML syntax? Was that rejected?
:)