Visualizing binaries with space-filling curves
corte.si
corte.si
First, using a color function that encodes local entropy to show how crypto keys and other high-entropy data can be picked straight out of a visualization:
http://corte.si/posts/visualisation/entropy/index.html
And next, using this in bulk on samples from a malware database, with what I think are beautiful and interesting results:
Hmm, I'm not sure if this could work, but perhaps if you use a stream based compression algorithm you could relatively precisely see how much compressed data it takes to represent up to a certain point in the file, with only a single pass (rather than having to compress a huge number of local windows). Of course this is probably going to weight earlier parts of the file heavier, simply because the compression won't be calibrated yet to efficiently encode. So you could also run it on a byte-reversed version of the file, and a "rotated" version of the file (i.e. file[n/2:n]+file[0:n/2]), and a rotated-byte reversed version, and combine all those metrics together in some way (maybe min(entropy1,entropy2,entropy3,entropy4)).
That way you could get an entropy measure which compensates for the sort of alphabet runs that fooled shannon entropy.
For a continuous compression function your idea of running on a rotated or reversed version of the file and then taking a minimum is a cunning one! Right now, I'm still working with sliding windows, but I'll keep it in mind if I turn back to compression.
The resulting 'knot ball' of traces could actually be discerned once you watched a few of them. You can pick out all sorts of things like when the video was refreshing, keys were being processed, etc.
Nice writeup(s) and interesting approach with the entropy calculations!
This allows people to recognize similarities between different files as the resulting patterns are revealed. Take this previous blog post by the same author:
http://corte.si/posts/code/sortvis-fruitsalad/index.html
This shows various sort methods operating on a set of unsorted data. If all you had visibility into was how the data was manipulated, by generating an image based on the data transformations, your pattern matching brain will quickly be able to see which sorting algorithm was used.