Fuzzing FFmpeg for fun and profit
obe.tv
obe.tv
There were some fuzzing projects before afl, but even with afl's recent popularity we're still in the situation where you just have to grab a random application, point afl at it and get some basic crashes in a few minutes. With clang-analyzer, coverity, afl, and many other projects available for free, there's no reason this should be possible.
Then again, I'm still waiting for the time when people don't code with sql injection issues...
The fact that the code is managed helps prevent against some nasty exploits, but it doesn't necessarily mean the code always does what the programmer intended.
So when asked, "Why fuzz test?" I can answer, "remember when you wrote fifteen test cases and wondered to yourself whether you needed a sixteenth? But then you decided that it would take more work than it's worth to do that. -- Well, now we can tell the computer to just keep trying stuff until it finds unique cases."
Example in Clojure: https://github.com/clojure/test.check
Because most software projects are so lazy about doing fuzzing themselves, making and promoting easier to use tools is a good way to nudge people into picking the low hanging fruit (bugs?).
Fuzzing works for managed code, but since you are looking for logic bugs instead of memory corruption crashes, more effort is needed to verify the behaviour of your app in the face of fuzzed data. Things like TOCTOU bugs, data sanitation bugs like path traversal or SQL injection vulns, etc. Of course this finds non-crashy security bugs in C code too, like in the Heartbleed case...
What's wrong with ffmpeg crashing when you feed it invalid input? What's the alternative to it not crashing? Should it continue transcoding or should it exit quietly?
What's the problem here trying to be solved?
After finding a crash, a researcher will generally explore this to see if it allows for arbitrary code execution. If it does, it is possible for someone to create a weaponized exploit.
For a utility as widely-used as ffmpeg, any bugs pose a threat to many systems. Fuzzing it makes the Linux ecosystem safer for all of us (hopefully).
I wish someone would do a JVM version (combined with the lines of QuickCheck and typed generators).
[1] https://googleonlinesecurity.blogspot.com/2014/01/ffmpeg-and...
This mentions the fixes they found (NULL pointer dereferences, Invalid pointer arithmetic leading to SIGSEGV due to unmapped memory access, Out-of-bounds reads and writes to stack, heap and static-based arrays, Invalid free() calls, Double free() calls over the same pointer, Division errors, Assertion failures, Use of uninitialized memory.)
Some of those things could have been caught by static code analysis or stricter coding standards. Or, if it had been a C++ project you could replace some of those problems (NULL pointer dereferences) with using references in the first place instead of passing pointers around. That might be a bit of an oversimplification as it is a very complex project but for me it was a reminder to change my coding style in C++ instead of sticking with the C-style way of doing it.
Further, templates do not make C++ slow; it is simply compile-time polymorphism.
Does runtime polymorphism make software slow? Does casting to or from a void* in C suddenly make it slow? The entire emphasis in C++ for templates is that you pass the error checking to the compiler, not at runtime to guess that this incoming void* is the type that you want. reinterpret_cast is only used where absolutely necessary!! It is far safer to NOT cast. Avoid casting. Templates (and the STL) help out in this regard.
I would use the STL for projects, and encourage every else in C++ land to do the same. Containers make sense, as do the algorithms associated with them (and they save you reinventing the wheel).
But I would NOT advocate use of Boost just because it is popular or written in C++. That would be a stupid reason to encourage its use; I personally do not use Boost. Boost != STL. You appear to be confusing Boost with the STL ??
No doubt your unbridled hatred of C++ and libraries written in C++ (eg. Boost) has stopped you reading reference works on the STL. You can find a comprehensive reference here: http://en.cppreference.com/w/
There's no Boost on there, btw. That's at boost.org.
Personally, I was only making the comment regarding a reminder of my own coding style in light of the problems they found in their C codebase. I do not write C, but I can see that the pitfalls they fell into could easily be carried out by me in C++ if I stuck to unsafe coding styles.
Furthermore, many of their problems could be solved in C++ land very easily: NULL pointer dereferences (don't use pointers, use references - it'll never be NULL), Invalid pointer arithmetic leading to SIGSEGV due to unmapped memory access (use STL containers), Out-of-bounds reads and writes to stack, heap and static-based arrays (use STL containers, you may still get out-of-bounds problems but at least an exception is thrown), Invalid free() calls, Double free() calls over the same pointer (use move semantics, never see a new or delete ever again), Division errors (use a static_assert to check that a value is in range, if possible. Or use an assert in a defensive coding style), Assertion failures (use static_assert to catch at compile time if possible), Use of uninitialized memory (RAII - initialise before use).
Your (or anyone's) invalid use of C++ does not make the language invalid, in the same way that stabbing someone with a knife does not suddenly mean all knives must be banned.
For FFmpeg that is a concern, but I believe the replier to my original comment was disparaging C++ for no good reason, other than it wasn't C...
I wish those titles had a weight to drag them off the front page much quicker.