Hijacking YouTube to transmit your data
banmeihack.wordpress.com
banmeihack.wordpress.com
There is no hole in anything. You're not violating anyone's privacy or stealing anything from anyone. Even the bandwith is given to you for free. It's how things are supposed to work.
You're just exercising your right to privacy by using such a thing.
One can tolerate someone re-inventing/re-discovering steno and making it sound like it's smth new... but not someone having no f idea whatsoever of what "security" means and what his "right to privacy" is... ffs!
edit: stegano, not steno
The hard problem is finding a way to encode data in video in a way that will survive recompression, resizing, or other video processing. The watermarking people have struggled with this for years. There are various spread-spectrum like schemes with good noise immunity that can do this.
YouTube has an ongoing battle between the copyright-infringement identification system and versions of audio and video modified to evade it.
You can significantly increase the bit rate. For example, overlay a QR code over each of four consecutive frames. You can do this without frame dropping. You only need to add about 6dB of the code for this to be recoverable. Similarly, if you know how the codec works, you can exploit that. (Your proposed method is actually pessimal for a modern B frame codec!)
Then there is hiding modem transmissions in techno music sound tracks ;-)
There is nothing insecure about it too.
You could even need commercial software for it, and could use TeX and friends [2] to achieve the same.
[0] https://blogs.adobe.com/insidepdf/2010/11/pdf-file-attachmen...
[1] https://wwwimages2.adobe.com/content/dam/Adobe/en/devnet/pdf...
[2] http://tex.stackexchange.com/questions/208012/attaching-file...
Youtube deliberately has the "upload a video" feature. It's not a mistake. It's not a security hole.
Also see this misguided and confused soul:
On the video side, you're dealing with at least VP9 and H264 which I'm assuming "destroy" your data somewhat in the encoding process. The audio side is Opus and AAC, with similar challenges.
But writing a data stream into a 2D still image in a way that can be decoded later is a solved problem, ie. 2D barcodes like DataMatrix, QR Code, and Microsoft Tag (which has up to 8 colors to further increase data density). These formats have built-in error correction that compensates for some missing blocks. However, we can tune the format to be closer to the video codec's internal structure, to make them play nicer together.
For example, we can set each barcode block to be within 50% to 100% of pixel size of the video's macroblock, to make it more likely that the video codec can reuse the macroblock with motion vectors in a P/B-frame, instead of having to put more bits to it, or have it accidentally mangle it.
Realistically, we can also increase our color palette, as we're not going to be scanning these barcodes in bad light conditions -- all we need to do is get the color mostly right. But the more we increase the palette, the less video codec can reuse blocks; so this is something we'll want to experiment with.
The biggest problem for the barcode approach comes from the addition of the 3rd, temporal dimension. We can have each frame form its own independently scannable barcode, but doing so, we'll want to build in some temporal redundancy, ie. have a chunk of data, or error correction for said chunk, be present in more than one frame -- to protect against occasional frame drops, very inconvenient frame drops (like when you lose an I-frame and the video is grey- on green-blocky for several more frames), and offer some extra protection against "normal" decoding errors.
By the way, there are existing implementations of this concept:
[2] Demo of above: https://www.youtube.com/watch?v=2_8GlFdlb0Y
[3] Same idea, some hackable code: https://github.com/Neohapsis/QRCode-Video-Data-Exfiltration
Would this be useful ?
1 you don't need a header for every frame but these barcodes do.
2 (for QR this is the worst one) there is a lot of space wasted to help detect and correct perspective distortion.
And if you're really feeling adventurous libquiet provides floating point output that can be put into any channel like video if you're willing to plumb it in there.
</plug>
I think this is supposed to be previous odd frame, given that 1 2 3 4 5 6 7 8 9 10 becomes 1 1 3 3 5 5 7 7 9 9.
The biggest downside I could think of would be the lag: data > render frames > encode frames > network > decode stream > render frames > scrape data
um. they already do that. They scan all the uploaded videos for copyrighted audio, fingerprinting and comparing the uploaded audio of a bajillion videos against 1/2 a bajillion songs.
It's definitely an interesting idea, but it's really nothing new. I remember a few years ago reading about people hiding compressed .zip archives inside of jpegs or something like that.
Does YouTube cut out audio frequencies that are beyond the hearing range?