No matter how good your input santization is, you still wouldn't ever send an unescaped query to a database, right? That's because the query is an output.
In other words, for the love of god please do sanitize your inputs.
"Sanitize inputs" means modifying the input before you even know where it's going. It's fine for stuff like normalizing user input (eg: "strip leading and trailing spaces") but should not be used to combat things like SQL injection or XSS.
For issues like SQL injection and XSS you should escape on output. Outputting HTML? HTML escape, or better yet: use templating framework that does it by default. Outputting to SQL? SQL escape, or better yet use prepared statements and pass in your arguments using an API that escapes by default.
In the "sanitize inputs" approach to handling these situations you can't store "O'Hara <3 Sue" as a value, because you need to "sanitize" the apostrophe for SQL and the less-than for HTML. In the "escape outputs" approach, you have "O'"Hara <3 Sue" in your SQL, and "O'Hara <3 Sue" in your HTML, and the user's input is preserved.
Okay.
That's not how I've ever used that term or seen it used. Prepared statements are a form of input sanitation. HTML purifiers are a form of input sanitation. Maybe this lingo is specific to PHP-land?
In any case, "You need to know the semantics of the sink in order to know what to do with an untrusted source" seems like an obvious truism not worth writing about.
Given how often developers get it wrong, I don't think it's written about enough.
Also, you say "untrusted source" here. Whether you trust the source or not is irrelevant. You should still be escaping the output where you use data from it in order to make sure your outputs are safe - the source could be compromised, or broken, or sending something valid that you didn't expect. Maybe this isn't quite so obvious after all.
That's the terminology being used by the document under discussion.
Honestly, I think what causes a lot of people to get it wrong, is that they don't understand the distinction between input filtering and output escaping. They see them as the same thing, and so they use them interchangeably.
> Prepared statements are a form of input sanitation.
No. Input sanitization involves removing "bad" stuff from the input. For example, you remove the "'" in "O'Hara" so that it doesn't mess up your SQL, but you end up storing "OHara" in the DB.
Output escaping (which prepared statements fall under) removes nothing. Instead, characters that happen to be special are escaped so that they are treated as literal characters, and not as special characters. The DB gets the user's original input: "O'Hara"
> HTML purifiers are a form of input sanitation.
I assume you mean HTML sanitization (https://en.wikipedia.org/wiki/HTML_sanitization). In which case, usually, yes. Note that there's a difference here because you're removing part of the input, not doing a lossless transformation as with escaping.
Another way to think about the difference is whether you're doing type conversion or not. When escaping for SQL, you're converting from text/plain to SQL. When escaping for embedding in HTML, you're converting from text/plain to text/html.
When you do input sanitization instead, you aren't changing the type, you're just making certain values impossible. For HTML sanitization, this means turning stuff like "<em>safe</em> <script>unsafe()</script>" into "<em>safe</em> ". Both are texp/html, but the latter has been "sanitzed".
In this case, input sanitization makes sense, as long as you have a universal concept of what "safe" means, ans as long as your input was actually HTML.
The place where people mess up is in thinking that they need to "sanitize their inputs" in anticipation of something downstream using that same string as a different type. In the HTML exaple, this would be taking a text/plain string, like "I <3 HTML" and stripping out "bad" characters to turn it into "I 3 HTML".
> Maybe this lingo is specific to PHP-land?
I've never used PHP, so I wouldn't know.
> In any case, "You need to know the semantics of the sink in order to know what to do with an untrusted source" seems like an obvious truism not worth writing about.
In practice, that doesn't seem to be the case. Almost every time someone says "sanitize your inputs" in response to an XSS or SQL injection exploit, they're getting it wrong.
The meaning of "sane" depends on where you're sending it to. A backslash is a perfectly reasonable character, for instance. Put it in the wrong place in a SQL string and you have bad news. Put a ' in the wrong place in a shell command, sometimes nothing bad happens, other times you get pwned.
The right way to escape strange characters is different if you're sending it to an SQL engine, or writing it into a JSON string, or into some HTML, etc.
> Another example of this kind of thing is SQL injection, an attack that’s closely related to cross-site scripting. NaiveSite is powered by MySQL, and it finds users like so: > > $query = "SELECT FROM users WHERE name = '{$name}'"* > > When a boy named Robert'); DROP TABLE users; comes along, NaiveSite’s entire user database is deleted. Oops!
Also from the article:
> And of course use your SQL engine’s parameterized query features so it properly escapes variables when building SQL: > > $stmt = $db->prepare('SELECT FROM users WHERE name = ?'); > $stmt->bind_param('s', $name);
And more from the article:
> The parallel for SQL injection might be if you’re building a data charting tool that allows users to enter arbitrary SQL queries. You might want to allow them to enter SELECT queries but not data-modification queries. In these cases you’re best off using a proper SQL parser (like this one) to ensure it’s a well-formed SELECT query – but doing this correctly is not trivial, so be sure to get security review.
And the article links to https://stackoverflow.com/questions/129677/how-can-i-sanitiz..., which further explains:
> What you should do, to avoid problems, is quite simple: whenever you embed a string within foreign code, you must escape it, according to the rules of that language. For example, if you embed a string in some SQL targeting MySQL, you must escape the string with MySQL's function for this purpose (mysqli_real_escape_string). (Or, in case of databases, using prepared statements are a better approach, when possible.)
I hope this satisfactorily answers your question.
You don't seem to understand the distinction between running algorithms on data and taking programs as input, which is what the GP talked about.
Any sufficiently complex program can be viewed as an interpreter for its inputs. Input into a calculator program is code which programs an equation. Input into a word processor is code which programs a document. Input into a video game is code which programs a real time simulation. Input into a compiler is code which programs an executable. These are all different types of executable code sequences.
However, the fact that modern computers can be exploited due to architectural and engineering decisions (eg memory unsafety) does not mean a separation between code and data is not possible.
In fact, it is precisely a hot topic how to cheaply bend current practices back to that model given the rampant amount of vulnerabilities in the wild.
If you never treat data as code, you can only do uninteresting things. My example was the email address. The instance you look into the "black box" of the string, you are starting to treat the string as an executable structure. An email address' raison d'etre is provide that; an address to send an email to. You can not do that without looking into it.
Now, from here we can discuss safe and unsafe ways of doing that. You could use string splits or what not, or you could use a parser combinator library. Doing the latter will make it easy to see that parsing and executing a program is not that different from parsing an email into an AST, (user, hostname), and then treating that as a higher order program (ie. we need to specialize with a message before we can execute it as a "send email" program).
In another context, though, the function 'f(x) = x + x' might itself be data. Within a given context, however, the answer is generally unambiguous: data is what is acted upon by code.
(The breaking of this distinction is one of the reasons self modifying code is Bad.)
No, the fact that current archs encode data and code in the same memory space and there are vulnerabilities does not mean the separation is not possible.
def echo():
user_input = input('enter some text: ')
print('your text: {}'.format(user_input))
Treating input as code (VERY BAD, DON'T DO THIS): def echo():
user_input = input('enter some text: ')
command = "print('your text: {}')".format(user_input)
exec(command)
The second example allows the user to do all kinds unintended stuff: enter some text: '); print(10**2000) #
your text:
1000000000000000000000...
(abridged)That would print however many zeroes the user specified, and use a whole lot of memory. With some creativity it's possible to cause lots of havoc.