This JPEG is also a webpage
lcamtuf.coredump.cx
lcamtuf.coredump.cx
This is, at present, the most efficient way to pack demos on the web; a few characters of uncompressed bootstrap code, then the rest is deflated.
You can see the final packed .PNG results here: Crankwork Steamfist https://stianj.com/crankwork-steamfist/, Everything is Fashion https://stianj.com/fashion/, and Inakuwa Oasis http://arkt.is/inakuwa-oasis/.
The tool used for creating both the demos and the packed .PNG is made by us and available on GitHub here https://github.com/ninjadev/nin/.
$ curl -o squirrel.html http://lcamtuf.coredump.cx/squirrel/
$ file squirrel.html
squirrel.html: JPEG image data, JFIF standard 1.01, comment: "<html><body><style>body { visibility: hidden; } .n { visibilit"
Open the file in a browser and read the page. Then: $ mv squirrel.html squirrel.jpg
Open the renamed file in a browser and only the image appears.I'm not sure what the security implications are. I'm not creative or devious enough to think of anything offhand, but a lot of attack vectors start off with this sort of misdirection.
You can use this technique to phish signatures. Send someone a document that reads "X" in format A and "Y" in format B. The victim signs file.A thinking they are endorsing X but you can plausibly claim that they signed file.B (because it's the same file) and hence endorsed Y. This is why digital signature standards need to include meta-data, e.g.:
https://github.com/Spark-Innovations/SC4/blob/master/doc/fil...
Scroll down to "bundle files"
And anyone else can plausibly claim that you carefully forged a file to get a victim to sign it -- the signature will be of the whole file, not just a single view of it.
But that said, you shouldn't sign binary files unless you have a reasonable understanding of what is in it (or trust the party presenting it to you).
Yes, of course, but by the time someone realizes this the damage may already have been done.
> you shouldn't sign binary files
There are a lot of things that people shouldn't do that they do nonetheless.
1. https://blogs.msdn.microsoft.com/ie/2008/09/02/ie8-security-...
If user content is on a separate domain, they can't do that.
Also fishing is a lot easier when you're on the real domain...
https://github.com/blog/1452-new-github-pages-domain-github-...
[0]: http://binwalk.org/
What router is it?
The router itself is a BT Internet (UK) branded one. Not sure of the exact model but I'll try to find out...
But, false alarm anyway, nothing interesting is happening. The firmware had updated and reset parental control settings on the router. The domain is on some blacklist apparently so it was redirecting to a page to finalise parental control preferences.
Sorry it wasn't any more interesting than that.
Edit: the reason it took me a while to figure this out was that the settings page it was redirecting to was nothing to do with parental controls!
Lesson learned: just use a VPN.
Is that an ASCII representation of what I think it is? (a well known .cx site)
At least we know a little about the software issuing this redirect!
https://en.wikipedia.org/wiki/Golden-mantled_ground_squirrel
A chipmunk would have a stripe going across the eye.
(Today, you learned your first squirrel fact!)
Thank you.
And here I was impressed merely by the delivery of image data in the HTML stream. Little did I realize your page is practically an Encyclopedia Rodentia.
Kudos and hats off.
Combining with PDF is also on the easy end of things, because the PDF header just has to be somewhere vaguely near the start.
Compare that to JavaScript, which will happily fail if you use new syntax or a missing function, and thus web pages which rely on JS often show up as just a full screen of white when something goes wrong, which it frequently does. That's not to say JS should be as flexible as HTML is here, but it provides an interesting contrast.
Why push the complexity onto the user? Someone who just wants to make a working website doesn't care about your pedantry.
Do they want their page to fail to render entirely when PHP outputs a warning? Do they want their website to be completely broken because they forgot to convert some of their text from Latin-1 to UTF-8 before pasting it into the document? Should we really expect them to have to modify their blogging software to validate custom HTML snippets, lest the entire page become unusable? Will they be pleased when their style of code falls out of favour in future, is deprecated, and then their page doesn't work at all later?
Moreover, strictness can backfire when you have such a diversity of implementations.
> Which creates room for bugs, exploits, and unexpected situations like OP's post.
The OP is not so much an unexpected situation as a carefully engineered one that's completely within the constraints HTML sets.
Displaying warnings is fine. Invalid XHTML, less so.
> Do they want their website to be completely broken because they forgot to convert some of their text from Latin-1 to UTF-8 before pasting it into the document?
The encoding of the content has nothing to do with the document markup.
> Moreover, strictness can backfire when you have such a diversity of implementations.
On the contrary: this prevents subtle bugs in the interpretation of invalid data by different implementations.
Because this approach has historically resulted in people who "just wanted to make a working website" making websites that only work in specific browsers (or worse yet, specific versions of specific browsers on specific platforms). And then those sites stuck around and infrastructure got built around them that made them hard to fix.
We've spent the best part of 00s fixing that mess, and there are still some pockets that haven't been properly cleaned up. If that's not a lesson to learn from, I don't know what is.
Isn't that more due to failure to handle exceptions and display errors to users?
You could describe it that way.... or you could describe it as failing to have reasonable default logic for handling faults & gracefully degrade.
I often wonder how different the internet would be if Postel's prescription never gained traction and fail-fast behavior were the norm instead.
Instead, it felt like we had a bunch of people who though that big lofty standards were so obviously correct that everyone else would take care of those boring implementation details, and 99.9999% of web developers correctly realized that there was very little measurable downside to sticking with something which was known to work.
1. Simple examples: namespacing is a good idea but it leads to gratuitous toil in most tools – e.g. a valid XML document which has <foo> should just work if you write a selector for /foo, as present in the document, rather than requiring you to do kludgey things like have to lard up every parser registering the same namespaces which are also declared in the document and writing fully-qualified selectors like /mychosenprefix:foo or /{http://example.org/fooschema/1.0}foo for every tag, every time.
Similarly, getting XPath 2.0 support to actually ship in enough tools to be usable would have made one of the better selling points for using XML actually exist as far as the average working programmer is concerned.
By the way, there's SLAX, it's isomorphic to XSLT but with nicer syntax. Nice approach but anyways, the standards are horribly complex and stiff.
Maybe it would be so bad that somebody else would make a new, more permissive standard that took off instead.
Maybe all of that has already happened.
Maybe that line from Battlestar Galactica was right - All of this has happened before; all of this will happen again.
They started making browsers more lenient since there was so much poor malformed HTML being produced
In a world where invalid HTML documents aren't rendered at all we could have had the evolution of the format dictated by Microsoft because of their market position.
The decision to allow this was made early and the liberal accept/strict transmit paradigm has in general made the web a mess.
On the plus side, the consistent failure of browser vendors to apply strict controls to input means that as an application security person I will probably never be out of work :D Even though this pattern of behavior is starting to change, legacy support means that I will still be dealing with these issues well into my retirement!
How many security flaws are the result of malformed HTML?
To take this article as an example, according to the HTTP specification, the `Content-Type` header is supposed to have the final say in what media type is being served. Internet Explorer decided it would be better to use heuristics. I think the idea was that if a web host was misconfigured, rather than have the web developer fix their bug, it would try to guess its way out of the error.
Which kinda worked. The problem was, it opened it up to abuse. If you had a web host that allowed untrusted people to upload images (e.g. profile photos), you could construct an image that tricked Internet Explorer into thinking that it was an HTML document, even if the server explicitly told clients that it was an image. The main difference between images and HTML, of course, is that HTML can contain JavaScript, which would now execute in the security context of your web page.
So all of these web hosts, thinking they were only giving people the ability to upload images, were now letting people execute JavaScript on their domain – simply because Internet Explorer tried to be lenient.
The workaround ended up being forcing downloads with `Content-Disposition` headers instead of displaying inline. That's why, for example, visiting the URL of an image on Blogger directly triggers a download instead of showing the image.
Other examples that spring to mind:
Netscape interpreting certain Unicode characters as less than signs. People were correctly escaping `<` as `<` but the Unicode characters slipped through and caused XSS vulnerabilities in that browser.
Browsers ignoring newlines in pseudo-protocols. Want to strip `href="javascript:…"` out of comments? No problem… except some browsers also executed JavaScript when an attacker placed a newline anywhere within the `javascript` token.
Being lenient in what you accept has caused security vulnerabilities over and over again and there's no reason to think that it will stop now.
Which kinda worked. The problem was, it opened it up to abuse. If you had a web host that allowed untrusted people to upload images (e.g. profile photos), you could construct an image that tricked Internet Explorer into thinking that it was an HTML document, even if the server explicitly told clients that it was an image. The main difference between images and HTML, of course, is that HTML can contain JavaScript, which would now execute in the security context of your web page. So all of these web hosts, thinking they were only giving people the ability to upload images, were now letting people execute JavaScript on their domain – simply because Internet Explorer tried to be lenient.
This is an interesting example, though I think this is a fair bit different than being lenient on HTML interpretation.
The topic was strict HTML. Accepting malformed HTML doesn't seem to pose much of a problem. Blindly executing a non-executable file seems like a much different problem.
> Netscape interpreting certain Unicode characters as less than signs. People were correctly escaping `<` as `<` but the Unicode characters slipped through and caused XSS vulnerabilities in that browser.
This doesn't sound like being lenient. This just sounds like a bug.
> Browsers ignoring newlines in pseudo-protocols. Want to strip `href="javascript:…"` out of comments? No problem… except some browsers also executed JavaScript when an attacker placed a newline anywhere within the `javascript` token.
Huh? I don't understand the scenario being described here. It again sounds like a bug rather than lenient acceptance of data, though.
It's not. There are two areas where the leniency was a problem here. Firstly, the leniency in rendering one media type as a completely different media type because the browser heuristic thought it was being lenient. Secondly, the leniency in parsing HTML out of an image file – you can't do that with valid HTML.
> Accepting malformed HTML doesn't seem to pose much of a problem.
I've literally just given three specific examples of it causing security vulnerabilities.
> This doesn't sound like being lenient. This just sounds like a bug.
No, it was intentional. It was specifically Unicode characters that looked like less than and greater than signs, but weren't.
> I don't understand the scenario being described here.
Somebody noticed that href="java\nscript:…" wasn't being parsed as JavaScript, and it was causing some malformed pages to fail to work properly. Rather than let it fail, they tried to fix it by stripping out the whitespace, and caused a security vulnerability.
If these three examples aren't enough, take a look at OWASP's XSS filter evasion cheat sheet. There's plenty of examples in there of lenient parsing causing security problems:
https://www.owasp.org/index.php/XSS_Filter_Evasion_Cheat_She...
I think you can argue the first is a problem. You have an example demonstrating as much. Arguing that the second is a problem is much harder. Lenient HTML acceptance been hugely advantageous to the adoption of the web. There may have been some issues from this, but it's valuable enough that the effort to "fix" it was abandoned and the W3C and WHATWG returned to codifying what leniency should look like.
> I've literally just given three specific examples of it causing security vulnerabilities.
Well, at least one example. Coercing a file served as an image to HTML isn't an issue of accepting malformed HTML, nor would I agree that the JS example is a problem with leniency.
> No, it was intentional. It was specifically Unicode characters that looked like less than and greater than signs, but weren't.
Okay, I reread your last comment. I initially thought you were saying that Netscape was treating '<' as '<'. So Netscape decided to treat some random unicode chars that happen to look kind of like the less-than symbol (left angle bracket: ⟨, maybe?) as if they're the same as the less-than symbol? This seems amazingly short-sighted and pointless. How was this issue not seen, and was this even solving a problem for someone?
> Somebody noticed that href="java\nscript:…" wasn't being parsed as JavaScript, and it was causing some malformed pages to fail to work properly. Rather than let it fail, they tried to fix it by stripping out the whitespace, and caused a security vulnerability.
So the issue here is incompetent input sanitization. I don't think the browsers being lenient here is the issue.
> If these three examples aren't enough, take a look at OWASP's XSS filter evasion cheat sheet. There's plenty of examples in there of lenient parsing causing security problems: https://www.owasp.org/index.php/XSS_Filter_Evasion_Cheat_She...
A few of these are interesting in the context of browsers being lenient. e.g. This one requires lenience as well as poor filtering:
<IMG """><SCRIPT>alert("XSS")</SCRIPT>">
Most of these are just examples of incompetence in filtering, though, and a great example of why you 1) should not roll your own XSS filter if you can avoid it, and 2) why you should aggressively filter everything not explicitly acceptable instead of trying to filter out problematic text.Wait, that's a completely different point. The argument here is that it caused a security vulnerability, and it did. If the lenient HTML parser didn't try to salvage HTML out of what is most certainly not valid HTML, then it wouldn't be a security vulnerability.
> This seems amazingly short-sighted and pointless. How was this issue not seen, and was this even solving a problem for someone?
You could ask the same of most cases of lenient HTML parsing. It's amazing the lengths browser vendors have gone to to turn junk into something they can render.
> So the issue here is incompetent input sanitisation.
No, it isn't. That code should not execute JavaScript. The real issue is that sanitising code is an extremely error-prone endeavour because of browser leniency – because you don't just have to sanitise dangerous code, you also have to sanitise code that should be safe, but is actually dangerous because some browser somewhere is really keen to automatically adjust safe code into potentially dangerous code.
Take the Netscape less than sign handling. No sane developer would think to "sanitise" what is supposed to be a completely harmless Unicode character. It should make it through any whitelist you would put together. Even extremely thorough sanitisation routines that have been worked on for years would miss that. It became dangerous through an undocumented, crazy workaround some idiot at Netscape thought of because he wanted to be lenient and parse what must have been a very broken set of HTML documents.
This is not a problem with incompetent sanitisation. It's a problem with leniency.
Thanks for providing actual, concrete examples.
(Yes, the output should have been escaped, but that is sadly not always the case)
The cost of a malformatted HTML document rendering despite the errors is not that severe compared to the benefits it provides, as we have seen.
It's not as if many of the web's security holes are related to whether a page displays valid HTML markup or not.
Ultimately it's less about saving keystrokes and more about amateur enthusiasts having the opportunity to start with the browser rendering their unformatted document rather than an "Error at line 1" warning, and changes they introduce being considerably less likely to break the entire page
The real security due to leniency problems in the web platform revolve around the handling of data and how JavaScript works, which were not addressed by XHTML. In that sense it's not surprising it went nowhere.
data:text/html, <html><img src="http://lcamtuf.coredump.cx/squirrel/"></html>
Put that in the url of the browser.https://developer.mozilla.org/en-US/docs/Web/HTML/Element/xm...
It's similar to the pre tag but doesn't require the escaping. I guess you just have to make sure you don't have a closing xmp tag :)
<![CDATA[ here &entities; or <angle|<brackets>> will not interpreted ]]>
There is no need for special-casing xmp, when SGML and XML already define CDATA escapes."User agents must treat xmp elements in a manner equivalent to pre elements in terms of semantics and for purposes of rendering. (The parser has special behaviour for this element though.)" — https://html.spec.whatwg.org/multipage/obsolete.html#require...
How to send money to your email address? Not that I would send you some, but I wondered how you want to have that money received?
now specifically in this case, lcamtuf (at google security) is joking and doesn't want your money.
this hack is actually pretty crazy - an arbitrary HTML / jpeg polyglot file that fooled a browser could be used for js injection, say from a site that allowed jpeg file uploads, and validated mime type.
https://websec.io/2012/09/05/A-Silent-Threat-PHP-in-EXIF.htm...
https://blog.sucuri.net/2013/07/malware-hidden-inside-jpg-ex...
The way we protected ourselves against it at <earlier company> (since we allowed image uploads at a variety of locations) was to decode and recode the image before storing and strip out comments.
[0] http://img03.deviantart.net/ebd4/i/2015/166/9/6/radical_dude...
2. Ask him his Bitcoin address
3. Paypal to this address
:)
But seriously, no, it's alive and kicking... https://www.google.com/wallet/
Page Info shows the last modification was: Thu 31 Mar 2016 01:02:44 PM EDT
I don't personally use it much but I have an account to pay for my domain name though them that I use for Google Apps.
I doubt this. In request with Accept:"image/png,image/;q=0.8,/*;q=0.5" server souldn't respond with something with Content-Type:"text/html"
cd ~/tmp/squirrel
wget -O index.htm http://lcamtuf.coredump.cx/squirrel/
firefox index.htm
Alternately, using an actual webserver instead of opening a file://
python3 -m http.server &
firefox http://127.0.0.1:8000/
Also works.
Note that the actual img tag has src="#" so when you are looking at the file opened locally, the image is also from local disk, not from his server, so it's legit.
However, the fact that it needs to be identified as a HTML file by the server implies to me that the idea others had ITT of abusing image hosting sites using this trick probably won't work.
http://dev.exiv2.org/projects/exiv2/wiki/The_Metadata_in_JPE...
00000000 ff d8 ff e0 00 10 4a 46 49 46 00 01 01 01 01 2c |......JFIF.....,|
00000010 01 2c 00 00 ff fe 03 72 3c 68 74 6d 6c 3e 3c 62 |.,.....r<html><b|
$ file index.html index.html: JPEG image data, JFIF standard 1.01, resolution (DPI), density 300x300, segment length 16, comment: "<html><body><style>body { visibility: hidden; } .n { visibilit", baseline, precision 8, 1000x667, frames 3
I wonder what are the security implications of that.
At least any terminal escape sequence can be executed if you run `file` on a JPEG, it seems, since this:
curl -s 'http://www.imagemagick.org/image/fuzzy-magick.png' | convert - -set comment "$(printf 'asdf\x1b[1;31mTest?\x1b[0m hmm')" test2.jpg
file test2.jpg
Results in red text on my terminal for me.(It also results in file writing a 0xff 0xdb to the terminal, which the terminal turns into the unicode fallback character since it's not valid text…)
test2.jpg: JPEG image data, JFIF standard 1.01, aspect ratio, density 72x72, segment length 16, comment: "asdf\033[1;31mTest?\033[0m hmm", baseline, precision 8, 320x85, frames 3It will be stopped dead by metadata filters though. Stripping out the comment would be step #1 for those devices.
Perhaps caching or other schemes would break this, I'll try to knock up a PoC.
2. JPEG also allows for comments and other embedded metadata which won't show up in the displayed image.
3. Start the file with a JPEG header and metadata section, then switch between HTML and JPEG using the comment functionality mentioned above
Essentially!
"But wait, why is it shown as an image in one context and as a web page in another?"
The answer is in the question: Context. If you expect a JPEG you will get a JPEG, and same for HTML.
There's an important note in that the HTML is not at the true start of the JPEG file, it's slightly after. You can even view some of the JPEG format bytes if you view source of the HTML.
So if the browser ignores some of the JPEG file, why not most of the PNG file? Perhaps you run the risk of some random byte screwing up the HTML though.. not sure.
Whether JS would work or not depends on how much the browser tries to recover from errors there. I would guess not much, but what would you gain from a combined JS/JPEG file anyway?
Funky File Formats
So this is technically a legitimate use of RIFF because it's designed to support multi-purpose contents, not just PCM data.
I'm not sure why they did that instead of just sending FLACs, though...
Unfortunately, your HTML gets lost I try to save the picture on my phone.
iPhone > Safari > Save to Camera Roll
Copy off the picture; it's been re-encoded and has no JPG comment field.
Any solutions to this are welcome! The re-encoding makes patterns impossible to decode as well (they degrade after being shared a couple of times). See Cemetech's jsTified emulator for an example of a ROM file as a JPG - he uses B/W and it still requires the file to be synced from a computer, not saved from the browser.
mv squirrel.html squirrel.jpg sudo apt-get install steghide steghide embed -cf squirrel.jpg -ef secret.txt mv squirrel.jpg squirrel.html
And voila...
There's a talk called "Funky File Formats" but there's (fittingly) multiple versions of it so you're best off searching for it.
Worth a look if you'd like to make your own! http://stegosploit.info/
https://mega.nz/#!49MnjKCJ!7HShAESfmM2R450x4z-zLmtCDotOhsLze...
Some of the PDFs also happen to be valid images, audio files, zip archives, etc.
The HTML file could be one that admonishes the user for attempting to scrape the file, all the while the file they wanted is sitting right there. A modern day Purloined Letter.
They've messed this up in the past, see this legendary bug bounty report [1]
1. https://whitton.io/articles/xss-on-facebook-via-png-content-...
This is different: it's a single file which can be parsed as either a HTML page or a JPEG. Hence, when a program expects a HTML page (like a browser loading a Web page), it will be parsed and displayed as a HTML page. When a program expects a JPEG file (like a browser loading the "src" of an "img" element) it will be parsed and displayed as a JPEG.
The trick is to use each format's comment syntax to hide the other format. Not sure if the HTTP headers need to be set differently for each request or not.
$ curl -I http://lcamtuf.coredump.cx/squirrel/
HTTP/1.1 200 OK
Date: Thu, 11 Aug 2016 05:18:00 GMT
Server: Apache
Last-Modified: Mon, 19 Sep 2011 23:31:49 GMT
Accept-Ranges: bytes
Content-Length: 135938
Content-Type: text/htmlI guess browsers only forbid ignoring Content-Type for stuff like JS, then. For JPEG it's probably not a security concern.
_____________________________________________________________________
| |
| =================================================================== |
| |%/^\\%&%&%&%&%&%&%&%&{ Federal Reserve Note }%&%&%&%&%&%&%&%&//^\%| |
| |/inn\)===============------------------------===============(/inn\| |
| |\|UU/ { UNITED STATES OF AMERICA } \|UU/| |
| |&\-/ ~~~~~~~~ ~~~~~~~~~~=====~~~~~~~~~~~ P8188928246 \-/&| |
| |%//) ~~~_~~~~~ // ___ \\ (\\%| |
| |&(/ 13 /_\ // /_ _\ \\ ~~~~~~~~ 13 \)&| |
| |%\\ // \\ :| |/ ~ \| |: 3.21 /| /\ /\ //%| |
| |&\\\ ((iR$)> }:P ebp || |"- -"| || || |||| |||| ///&| |
| |%\\)) \\_// sge || (|e,e|? || || |||| |||| ((//%| |
| |&))/ \_/ :| `._^_,' |: || |||| |||| \((&| |
| |%//) \\ \\=// // || |||| |||| (\\%| |
| |&// R265402524K \\U/_/ // series || \/ \/ \\&| |
| |%/> 13 _\\___//_ 1932 13 <\%| |
| |&/^\ Treasurer ______{Franklin}________ Secretary /^\&| |
| |/inn\ ))--------------------(( /inn\| |
| |)|UU(================/ ONE HUNDERED DOLLARS \================)|UU(| |
| |{===}%&%&%&%&%&%&%&%&%a%a%a%a%a%a%a%a%a%a%a%a%&%&%&%&%&%&%&%&{===}| |
| ==================================================================== |
|______________________________________________________________________|
source: http://chris.com/ascii/index.php?art=objects/moneyBut ultimately I think it's for the greater good.
Sth. that one needs to get used to on HN I think... Randomly picking on arguably off-topic and arguably offending posts proactively. Also, I'm shocked seeing that the mod here has so low karma and so short history of HN usage.
http://naturemappingfoundation.org/natmap/facts/chipmunk_vs_...
http://www.differencebetween.info/difference-between-squirre...
Sadly, that was my most upvoted comment in ~9.5 years on this site.
http://pmdvod.nationalgeographic.com/NG_Video_DEV/935/311/le...
https://www.google.co.uk/search?q=palm+squirrel&source=lnms&...
Good edge case for browser tests!