How to defeat naive image steganography
incoherency.co.uk
incoherency.co.uk
Steghide uses a graph-theoretic approach to steganography. You do not
need to know anything about graph theory to use steghide and you can
safely skip the rest of this paragraph if you are not interested in
the technical details. The embedding algorithm roughly works as fol‐
lows: At first, the secret data is compressed and encrypted. Then a
sequence of postions of pixels in the cover file is created based on
a pseudo-random number generator initialized with the passphrase (the
secret data will be embedded in the pixels at these positions). Of
these positions those that do not need to be changed (because they
already contain the correct value by chance) are sorted out. Then a
graph-theoretic matching algorithm finds pairs of positions such that
exchanging their values has the effect of embedding the corresponding
part of the secret data. If the algorithm cannot find any more such
pairs all exchanges are actually performed. The pixels at the
remaining positions (the positions that are not part of such a pair)
are also modified to contain the embedded data (but this is done by
overwriting them, not by exchanging them with other pixels). The
fact that (most of) the embedding is done by exchanging pixel values
implies that the first-order statistics (i.e. the number of times a
color occurs in the picture) is not changed. For audio files the
algorithm is the same, except that audio samples are used instead of
pixels.
I wonder how hard it would be to detect (not decode) content hidden by steghide.[1]https://www.cs.ox.ac.uk/teaching/courses/advsec/ [2]https://www.cs.ox.ac.uk/teaching/materials15-16/advsec/advse... p48 [3] ^ p58
My pet idea for doing steganography was to use nonstandard or suboptimal encoding sequences. For compressed data formats, including images, music, video, and general files, there is almost always a choice of various encodings which will decompress to the same result. Systematically varying a compression choice gives an invisible way to encode data.
I'd be interested to read this.
The problem was that most image hosts were smart enough to detect the sub-optimal encoding and "helpfully" fix it.
And, in general, people don't hide the image of some text inside another image. They hide the text. And that isn't so linear (blocks of 0s or 1s)
Could be combined with barcodes as well.
Definitely not naive, but an interesting hack.
Here's some messages hidden in audio 'noise' generated from stretching out all the black and white squares of a QRcode and generating some audio from it.
It can also be decoded back to the QR code and back to the message.
Neat.
EDIT: Fixed. It was because I changed it down to greyscale too late in the process.
[1]https://www.cs.ox.ac.uk/teaching/materials15-16/advsec/advse...
Are you able to exfiltrate this information?