XSS war: a powerful Java HTML sanitizer
roberto.open-lab.com
roberto.open-lab.com
http://www.davidpashley.com/blog/computing/livejournal-mozil...
I also don't think you're the right thing when you DO find something that's not on the whitelist - you should be escaping it rather than stripping it (we couldn't have a discussion about your code using your system, since our XSS examples would be stripped).
I'd suggest re-engineering to use a whitelist for everything.
Even if you filter out "position: absolute", there's a chance people might figure out a way to do something similar using enormous padding values or negative margins.
Your general approach (tokenise the HTML and use a whitelist) is an OK start, but you should be white-listing attributes as well. You should also have an ENORMOUS set of unit tests.
You allow object and embed which is very worrying - the allowScriptAccess attribute can allow Flash to make JavaScript calls to the parent page, for example.
Also remember this: you're not dealing with valid HTML, you're dealing with malicious HTML that might be designed to evade your filters but still be handled by browser's built-in error correction code. Since the most widely used HTML engine is closed source, there's no telling what kind of weird constructs might be error-corrected and rendered by IE.
HTML cleansing is a mine-field.
http://www.feedparser.org/docs/html-sanitization.html
I'm not convinced it's possible to stop a browser executing code, because there are so many possible ways a browser can be given code to execute. Not only do you have to read all the specs for all the versions of the browsers, you've got to find all the bugs in them too.
Case in point: An earlier version of this code didn't remove javascript from CSS expressions, making it possible to get past it in IE6 and 7.
I gave up trying to sanitize HTML, and instead used a library to render it to plain text and stuck it in a <pre> element with usual HTML escaping. But I need to take a very paranoid approach in my app.
EDIT to add: I think this is probably one of the better java implementation, and has a good whitelisting approach to HTML. However, it's let down by taking a blacklist approach to CSS.
In many cases, you don't have to use HTML at all. Use something like MarkDown and it will save you a lot of time and frustration. http://daringfireball.net/projects/markdown/
But you're right, using HTML is very often overkill.
<s c r i p t>alert('xss')</ s c r i p t>
I think older versions of IE might execute this - the LiveJournal XSS filter has defence against this one. At the very least though it should be escaped rather than being allowed through on to the page.
I'm no security guru, but isn't the best practice to have a whitelist of accepted elements rather than a blacklist of prohibited elements?