Bypassing CSP using polyglot JPEGs
blog.portswigger.net
blog.portswigger.net
1. If the site has a CSP allowing JavaScript from "self", and
2. If the site has an upload feature hosted on the domain of "self", and
3. If the site gives someone the ability to inject a script tag pointing to an arbitrary target,
4. Then it's exploitable.
But conditions 1-3 effectively are the "you've already been rooted" case. This is a technically-interesting way to exploit the fact, but in itself is not the vulnerability; it's simply exploiting a vulnerability that was already there and wide open.
There's a whole host of issues associated with not doing this, including potentially unwanted exif data, and e.g. just cat'ing a jpeg with a rar file and using the image host as an arbitrary file host, etc.
> Placing shells in IDAT chunks has some big advantages and should bypass most data validation techniques where applications resize or re-encode uploaded images. You can even upload the above payloads as GIFs or JPEGs etc. as long as the final image is saved as a PNG.
An attacker will likely be able to figure out the exact re-encoding you apply, so unless you add some form of randomness, the attacker can work backwards to get the payload they want included.
Even with that, I'm sure this is an attack vector on a lot of sites that don't bother with these simple sanity checks. It's a really cool little example though.
This puts all of the recognizing/parsing code in the same location. It also verifies the entire input at the same time, before the results are passed back to the main program. You get clean valid/invalid check of the entire input.
For a very good discussion of why formal recognizers are important (and why, if possible, it's important to design transport formats and protocols that are deterministic context-free or simpler[2]), see Meredith and Sergey's talk[3] at 28c3.
[1] https://github.com/abiggerhammer/hammer/
[2] http://www.langsec.org/occupy/
[3] https://media.ccc.de/v/28c3-4763-en-the_science_of_insecurit...
(I mean, that's what this comes down to, right? Both formats support comments, and starting a comment in one is a valid start for the other, so you can interleave them and do whatever you like.)
JPEG comments exist for the same reason that EXIF tags exist – it's handy to store metadata alongside the image data, it gets copied around when the file gets copied, the tags can be transferred if the image gets re-encoded. There are enough error recovery mechanisms built into browsers that one could likely make a polyglot by just abusing the data segment, maybe even while crafting a legitimate standards-compliant JPEG.
Ultimately, bytes are bytes! Interpreting them with a variety of content types can give a variety of results, so keep it in mind.
[0] Resynchronisation / recovery from bit errors is one of the explicit motivations behind the design of Unicode encodings, so the browsers get a pass from me on this one. It's almost certainly possible to craft a suitable JPEG using legitimate code points anyway.