GitHub's post-CSP journey
githubengineering.com
githubengineering.com
I figured it out by the end of the first paragraph though.
We should prefix the security CSP with an m (for mundane) to differentiate it with the interesting CSP :)
That said, I disagree with his original point. I am a regular HN reader and did not know about content security policy.
It's true that the submitted blog post doesn't expand the CSP initialism -- this has been acknowledged as an oversight by the author and the post will be edited, as it's simply good practice as recommended in several style guides [1][2][3][4].
But the post's very first sentence links to a previous entry about the same subject; there, the topic is explained to readers who may not be familiar with it. There is simply no reason for a follow-up post, which this submission is, to repeat the explanations given by its predecessor. Meanwhile, a reader approaching the post with confidence about its topic will have realized rapidly what it is and isn't about.
So given the intense disagreement, what approach would have been preferable?
[1] http://amastyleinsider.com/2011/11/04/cheat-sheet-for-abbrev... [2] http://www.chicagomanualofstyle.org/16/ch10/ch10_toc.html [3] https://www.rfc-editor.org/materials/abbrev.expansion.txt [4] https://en.wikipedia.org/wiki/Wikipedia:Manual_of_Style
It's certainly the rule for academic writing.
No child comment on here seems to have pointed it out. Surely that is acceptable considering the level of detail they go into in their original post?
But I was also overwhelmed. There's a quip that security is a losing battle, but that wasn't my takeaway -- rather, the knowledge space required to develop and host a web application that accepts user-generated content in a way that won't leak info apparently everywhere is getting too much for generalist developers working alone or in small teams.
I don't understand why so much of this article is talking about dangling markup? While I highly recommend using the highest quality library you can get your hands on for this task, cleaning up user-supplied markup to at least be valid HTML is generally not that difficult. Cleaning up every last vector within that valid HTML is much harder and much more subtle; I have fought that fight myself, so I understand the bits about how awful the <plaintext> tag can be and so forth (and how how many other vectors there are for javascript, and how many vectors there are for loading content you didn't want loaded, and how many vectors there are for subtle leaks of information, etc., even within syntactically-valid HTML). But making syntactically-invalid HTML into syntactically-valid HTML that is at least not "dangling" is not that hard, and can be done reliably.
So I assume I'm missing some sort of context here about why they are having so much trouble with this? What's the context where they can't run this sort of syntax cleaner over the user input?
The solution to sql injection is parametrized query construction, instead of automatically filtering the ' character and then pasting the template and the user input together.
Similarly, to avoid cross site scripting you have to combine your templates and user input by escaping it properly, instead of just concatenating them and then hoping you can avoid the issues by input 'sanitizing'.
The root cause of XSS is improper user input sanitization. Usually it is in places that shouldn't allow HTML at all, where the developer forgot to properly HTML-escape the user-provided input. If not encoding it at all can slip by, then obviously not verifying that its valid HTML could too (plus, it doesn't make sense to verify valid HTML when you aren't even expecting HTML).
The focus on dangling markup is because a CSP policy that prevents unauthorized JavaScript from executing (no inline scripts, remote scripts from trusted hosts only) can resolve most of the XSS issues, but it doesn't resolve leakage of information over other channels (such as via images).
While CSP could be used to completely block images from external hosts (and thus solve the leakage issue), you sometimes do want to allow users to inline external images and so an alternative solution to prevent sensitive information from being leaked is required.
Edit: haven't seen ptoomey3's comment while I wrote mine. He explained it better :-)
I'm having great reservations towards CSP however. I think it breaks the web in a way that wouldn't have been necessary had we been a little bit more careful about HTML syntax rather than dismissing markup validation as an obsolete technique back when the vulgar "HTML 5 rocks" campaigns were in full swing.
CSP spec drafts have been around forever but were never finalized. CSP basically blocks execution of JavaScript in script tags in content (as opposed to script in the header), as well as in content handler attributes (onclick and co.) by disabling those alltogether on a page. This totally breaks page composibility where you assemble content at the markup stream level from multiple sources, like, say on every single news aggregation site. The removal of scoped CSS styles from HTML similarly breaks composition.
From Chrome's Content Security Policy page:
[Blocking inline script] does, however, require you to write your code with a clean separation between content and behavior (which you should of course do anyway, right?)
I think this comment is totally clueless wrt. what the Web is about. "Separation of concerns" is most certainly not a characteristic of the Web, and never has been.
I'm sorry, but rather than using kludges such as CSP to turning the lights off with a broad brush, how about fixing HTML and JavaScript in the first place?
(note my comment isn't addressed at github but at web standard comitees)
[1]: https://developer.chrome.com/extensions/contentSecurityPolic...
OK... how?
To be clear, I'm asking for an HN-comment level of detail, not a standards-body level of detail. I can't speak for everyone else on HN but I won't go over things with a fine-tooth comb, I'll only look at top-level issues.
But I will at least point out that being able to casually float third-party content into any site that has a weakness to XSS or man-in-the-middle or vulnerabilities in any other third-party content in the website is pretty fundamental. The fundamental composition power of the web is too powerful and it is going to have to be cut back. Some of the obvious solutions like "whitelisting hashes of valid content" have their own problems, like how a lot of the scripts being included out there are deliberately not constant and that's their whole point in the first place.
2. Defining and using safe CSS subsets (with countermeasures against click-jacking/-phishing, and hiding "nose print" etc.) though granted this is challenging; I had hoped a formal semantics for CSS came along (such as in [2]) but it didn't
3. Using HTML-aware template engines (such as [1] but there are maybe lighter approaches with hard-coded HTML rules as well; disclaimer: my project)
[1]: http://sgmljs.net/blog/blog1701.html
[2]: https://lmeyerov.github.io/projects/pbrowser/pubfiles/extend...
The "Origin" header is similar to the "Referer" header but never contains the path or query. Furthermore, CSRF protection requires it only for "POST" requests (i.e. "GET" requests are unaffected). So there is little incentive for an option disable it for privacy concerns.
[1] https://tools.ietf.org/html/rfc6454
[2] https://en.wikipedia.org/wiki/Cross-site_request_forgery
[3] https://bugzilla.mozilla.org/show_bug.cgi?id=446344
[4] https://developer.microsoft.com/en-us/microsoft-edge/platfor... (filed by me)
However, if someone has the ability to make malicious HTTP requests on my behalf using my browser can you really be sure that they don't have the ability to make malicious HTTP requests with altered headers through a malicious extension or a browser specific exploit or some other vector?
You still have to do all the other attack mitigation strategies in addition to checking the Origin header, and I'm not sure the extra complexity buys you anything in the long-term.
Nonces have the benefit of only relying on browsers preventing cross-domain reads.
When Flash is deprecated, and if a site wants to use CSP, then this might start looking like a better trade off.
ATM though, nonces can be automatically added to all same domain forms on your site with JavaScript and you can check it trivially on all POST requests, getting most of the non-CSP related benefits without waiting on browsers.
And even if browsers were to implement it, there is still a long tail of browsers out there that will take forever to update.
Seems Firefox will implement it anyway because it's still better than nothing.
https://www.igvita.com/2016/08/26/stop-cross-site-timing-att...
If you look at : https://cspvalidator.org/#url=https://github.com you'll see that the CSP policy directive defines the origins from which images can be loaded
'self' data: assets-cdn.github.com identicons.github.com collector.githubapp.com github-cloud.s3.amazonaws.com *.githubusercontent.com ;
Previously, images could have been loaded from additional domains (gravatar) and could have been used to leak CSRF tokens. https://www.gravatar.com/avatar/0?d=https%3A%2F%2Fsome-evil-site.com%2Fimages%2Favatar.jpg%2f
But GitHub is the one opening and closing the tag, probably in some kind of template: <img src="{gravatar_url}">
<p>secret</p>
Which should result in this: <img src="https://www.gravatar.com/avatar/0?d=https%3A%2F%2Fsome-evil-site.com%2Fimages%2Favatar.jpg%2f">
<p>secret</p>
and not this: <img src="https://www.gravatar.com/avatar/0?d=https%3A%2F%2Fsome-evil-site.com%2Fimages%2Favatar.jpg%2f
<p>secret</p>
Any idea why they are getting the latter?> In a relatively unique project, we asked Cure53 to assess what an attacker could do, assuming a content injection bug in GitHub.com
If you liked this, btw, worth reading some content from their pentester https://cure53.de/ which also had some interesting findings and links.
https://www.troyhunt.com/how-chromes-buggy-content-security-...
I had a very similar experience, and I'm still hunting for the specifics of the browser that caused it in my case.