Pulling JPEGs out of thin air (2014)
lcamtuf.blogspot.com
lcamtuf.blogspot.com
It's not just a fuzzer, it guarantees to hit every part of the spec (subject to what "profile" you're implementing). It's not free, it's a product for sale to implementers of HEVC for verification purposes.
You could imagine a parser that deals correctly with every single one of the test streams and implements every single feature in the spec, yet also has an undetected exploitable vulnerability because it made an assumption about objects' sizes (which the spec permitted it to make, but which an attacker could take advantage of).
(On the other hand, maybe I don't understand enough about H.264 to appreciate a reason why this isn't possible in this specific context.)
Validating against a spec is almost the opposite, since first it has to check that there is a code path for each part of the spec.
One extreme might be if you had a backdoor where the presence of a particular byte sequence in the input intentionally triggers some kind of malicious activity. The test streams can't detect this because they presumably don't contain that exact byte sequence, whereas something like AFL can find it because it can (potentially, depending on the nature of the test that recognizes the backdoor sequence) deduce what input would trigger coverage of that code path.
Since then the bug list has grown impressively: http://lcamtuf.coredump.cx/afl/#bugs
My guess is that it was something like:
afl-fuzz: There is a bug in tmpfs
openbsd: Ok, who's responsible for fixing tmpfs.
everyone: ...
openbsd: Ok, no one is maintaining it. Let's go ahead and disable since there are alternatives, and no one is providing maintenance.
For example I setup a site where I require users to upload an SSH key (for access to a git repository), and figured I'd do what github, etc, do in the display - show the fingerprint.
Given an SSH key you can get a fingerprint like so:
deagol ~ $ ssh-keygen -l -f ~/.ssh/id_rsa
2048 4d:19:f2:de:ba:f6:06:31:98:af:9e:2a:di:ce:ca:b2 ~/.ssh/id_rsa.pub (RSA)
Can you imagine an SSH key causing ssh-keygen, or ssh to segfault? I found one over a weekend:https://blog.steve.fi/so_about_that_idea_of_using_ssh_keygen...
I found similar issues with other well-known tools, for example a program that would cause GNU awk to segfault.
Really I should do more..
> When fuzzing a format that uses checksums, comment out the checksum verification code, too.
def client_for(arbitrary_socket)
valid_requests_discovered = repeatedly_fuzz_probe(arbitrary_socket)
valid_request_formats = cluster_and_generalize(valid_requests_discovered)
peer_idl = formalize(valid_request_formats)
client_module_path = idl_codegen(peer_idl)
compile_tree(client_module_path)
require(client_module_path + "/client.so")
end
I imagine that if we're ever doing the Star Trek thing, gallivanting around in starships encountering random alien species with their own technology bases, this would be the key to anyone being able to meaningfully signal anyone else.However, assuming the above sci-fi scenario—a cooperative peer that wants to communicate, but shares no common signalling standards with you—you could assume that everyone is doing broadcasts on loop of their daemon binaries to let their peers analyze them.
This would require some bootstrap logic on each side, to find what signalling methods their peer is using to broadcast the bootstrap. This could be made an approachable problem if everyone assumes everyone else will use "lowest-common denominator" signalling technologies for their bootstrap transmissions (e.g. binary over AM radio.)
You'd also have to figure out the ISA of the recovered binaries, in order to begin to fuzz them. This isn't impossible to do automatically either; it basically requires an "inverted" fuzz process—fuzzing up a VM with various decoding strategies and opcode definitions, and then seeing what it does when fed the provided binary, attempting to find something that produces a "normal" program execution (maybe according to a neural network trained on what memory-states of running programs usually look like.)
---
To jump back out of the sci-fi context, though—if you assume an "opaque peer" that nevertheless has its source code published online somewhere (e.g. as FOSS on GitHub), then you can do something like this:
1. set up a web spider that downloads source code from the internet, compiles it, and then fuzzes the resulting binaries (presumably all in a sandbox) to create mappings between project URLs and fingerprints of behavior under fuzzing.
2. when trying to communicate with an arbitrary server, first fuzz it a bit to try to fingerprint the response using the index.
3. Take the results as a fingerprint; feed them to the index service; get a source URL; clone the relevant project.
4. Fuzz-analyze the source code of the project along with the responses of the live peer, comparing the two as you go to try to get your local "mental model" of the peer as aligned as possible with how it's really behaving. (E.g., use the live peer's responses to generate the config file for your local copy of the project.)
5. With a behavior-matched local copy, generate an IDL client as above. (Which can, of course, be a very easy process if it turns out the server obeys some previously-discovered protocol.)
Over the internet, this probably gets lost in the noise, but maybe with enough inputs, it wouldn't.
But your idea is very clever and seems like a potential new application of timing attacks! Someone should definitely try this.
./configure --disable-shared CC="afl-gcc"
It calls the AFL GCC compiler instead, which presumably injects instrumentation.In case you'd like to test persistent fuzzing (should be faster due to a few factors), you can try modifying https://github.com/google/honggfuzz/blob/master/examples/lib... - it will probably take creating main() func which will read data in the AFL_LOOP() loop and call LLVMFuzzerTestOneInput with that input.