Where is this data they are sanitizing coming from? Why would you want every browser treating this differently?
Where is this data they are sanitizing coming from? Why would you want every browser treating this differently?
The reason being that all browsers parse potential XSS (and mXSS) content slightly differently so that a server side sanitizer by definition will never parse content in exactly the same way as a specific version of a given browser, since it's not sharing exactly the same runtime.
This semantic gap between browser rendering quirks and server side approximations can be exploited by an attacker to slip content past the server side sanitizer.
And that's why the browser vendors are actively working with Cure53, the author of DOMPurify, in order to ship Sanitizer API.
I'm just surprised this has taken so long.
Aside from that, a subset of people using DOMPurify—maybe even a plurality—seem to treat it as a talisman and can't explain how or if they've configured it correctly to provide the kind of protections they need for their use case. Security is not a separable concern.
It is not clear to me what subtle things are wrong with the parent comment. But I'm pretty sure I can see the flaw in claiming that something is wrong without saying how. It's an impossible-to-disprove accusation which therefore does not advance the conversation.
Your comment is nothing but vague, unverified claims. Why do should you be given the benefit of the doubt but not the person you're replying to? If you have a real concern, share it so that it can be evaluated and everyone, including the original poster, can learn from it.
And then you proceed to imply developers who have UIs susceptible to XSS are to blame, even though this is a known attack vector where browsers parse adversarial trees in an exploitable way that is completely opaque to everyone?
This is silly. Hold your tongue or speak clearly, don’t jeer from a high horse
This doesn't make much sense to me, my app doesn't even rely on any servers, what do you suggest I do then?
> Where is this data they are sanitizing coming from?
It's untrusted, what does it matter where exactly it comes from?
> Why would you want every browser treating this differently?
Why do you say that each browser would tread this differently? There's a spec for how this should work: https://wicg.github.io/sanitizer-api/
Although I agree that there is nothing wrong with client-side sanitization, does this even apply to non-networked apps? Are you protecting your user from?
Or perhaps by "doesn't rely on any servers" you mean "serverless", which doesn't literally mean no servers. If your app's client hits a database-as-a-service somewhere then that is a server and it very well can sanitize for you.
You can have no back-end code but still can call external APIs hosted on servers you don't control or which can be compromised by malicious user generated content.
Kind of from themselves, the app can open arbitrary files, you can't have people nuke themselves by opening a random file they got from the internet, or by pasting in something they don't understand. Also a plugin may sort of escape from the boundaries put on it by XSS-ing its user by outputting some malicious stuff.
> Or perhaps by "doesn't rely on any servers" you mean "serverless", which doesn't literally mean no servers. If your app's client hits a database-as-a-service somewhere then that is a server and it very well can sanitize for you.
No I actually meant no servers at all, it's a 100% local Electron app.
Some pages render data from GET parameters, which allows XSS by giving people a link. There's also scenarios where data from localstorage, cookies or manually copy-pasted data is an attack vector. Imagine I get you to paste my document into your text editor and that allows me to extract all your data.
The above example is a vulnerability in the app just the same, it just requires the attacker to consider different avenues of delivery (such as sharing the malicious script on a forum) then in traditional server connected apps.
There is also a spec for how html should work.
And finally, 100% of the useful functionality of this ‘api’ could be implemented using a Javascript library or server side. Yes, browsers have slight disagreements over handling broken html. So the first step of server side sanitation is to remove the broken stuff. This also removes a lot of cross browser incompatibility. And then you select tags and only allow the tags you want. Basic xss sanitation like it’s 1999.
Nobody needs any built in api for this. This proposal is about solving a solved problem.
This isn't the sort of API that unlocks some previously impossible feature, but DOMPurify is incredibly slow and I would trust the browser's implementation more, obviously. Rewriting DOMPurify better has always been an option, but that's not an easy value proposition for most developers, hence the best we've got is still DOMPurify.
- It isn’t hard to do what this api does in Javascript as fast as the browser can do it. There already is api to convert strings to DOM and it isn’t exactly rocket science to iterate over the result and drop tags that aren’t on a list. I’m sure there can be a library that does it slowly but it doesn’t have to.
Right, so what?
> - It isn’t hard to do what this api does in Javascript as fast as the browser can do it. There already is api to convert strings to DOM and it isn’t exactly rocket science to iterate over the result and drop tags that aren’t on a list. I’m sure there can be a library that does it slowly but it doesn’t have to.
Of course it's DOMPurify's fault for being that slow, nobody is arguing that it must be rewritten in C++ to be fast.
It's not as simple as you put it though, unless you have an extremely strict list of tags containing no tags in which case it's trivial, but it's also of very little use.
Since your comment sounds overly arrogant to my ears I'd like to point out that by following the rules that you mentioned for making a sanitizer (DOMParser + drop non-whitelisted tags basically) you can only produce either a useless or a broken sanitizer, it's impossible to make a useful and working sanitizer that way. Proof-ish: either you drop <img> nodes, in which case you've made a pretty useless sanitizer in my book, or you leave it in, in which case you leave yourself open to XSS via stuff like the onerror attribute.
It's not trivial to write a sanitizer that both works and it's useful, almost no developers should need to learn all the nuances necessary for writing one, hence the platform itself should provide it.
So you allow img tags with the attributes you know and drop everything else.
And if it’s useful to have a prepared list of safe tags, why would that need to be provided by the browser?
> It's not trivial to write a sanitizer that both works and it's useful
While I would disagree with this claim, if it is true it is even more reason to not set this in stone by building it into browsers and instead provide a basic external library that makes the policy decisions that are useful to you.
And of course this is what is going to happen anyway, because you can’t count on this always being available, so you need to include a library that provides this functionality in JavaScript if the browser doesn’t. Which is trivial because this doesn’t involve anything JavaScript doesn’t have access to.
That's why you should be using an API that relies on the browser's parser.
Apart from performance this smells of not using a whitelist mechanism (I hope this is not the case).
What makes you think that? I just skimmed the draft and it seems to use a sensible whitelist as default. Developers can allow or deny additional elements/attributes as they like.
The browser knows if a certain piece of data will perform execution or not, as it is the software implementing the functionality. It is the correct app to ask, as it is the one being exploited.
Further, it's not a panacea. Defense-in-depth still applies, so server-side and other mitigations will still be appropriate. Build and use threat models to understand what is appropriate.