YouTubeDrive: Store files as YouTube videos
github.com
github.com
The encoding scheme that YouTubeDrive uses is brain-dead simple: pack three bits into each pixel of a sequence of 64x36 images (I only use RGB values 0 and 255, nothing in between), and then blow up these images by a factor of 20 to make a 1280x720 video. These 20x20 colored squares are big enough to reliably survive YouTube's compression algorithm (or at least they were in 2016 -- the algorithms have probably changed since). You really do need something around that size, because I discovered that YouTube's video compression would sometimes flip the average color of a 10x10 square from 0 to 255, or vice versa.
Looking back now as a grad student, I realize that there are much cleverer approaches to this problem: a better encoding scheme (discrete Fourier/cosine/wavelet transforms) would let me pack bits in the frequency domain instead of the spatial domain, reducing the probability of bit-flip errors, and a good error-correcting code (Hamming, Reed-Solomon, etc.) would let me tolerate a few bit-flips here and there. In classic academic fashion, I'll leave it as an exercise to the reader to implement these extensions :)
So, feel free to have a look and have a laugh, but don't try to use YouTubeDrive for any serious purpose! This encoding scheme is so horrendously inefficient (on the order of 99% overhead) that the effective bandwidth to and from YouTube is something like one megabyte per minute.
Can see one in action here
16x16, 8x8, or 4x4 would be the way to go. You'd want each RGB block to map to a single H.264 macroblock.
Using non order of 2 numbers means that individual blocks don't line up with macroblocks. Having a single macroblock represent 1, 4, or 16 RGB pixels would be ideal.
In fact, I bet modifying the original code to use a scaling factor of 16 instead of 20 would produce some significant improvements.
It would be better to use YUV/YCbCr directly instead of RGB.
There's a bunch of other things too, like YUV420p and TV colour range: 16-235, so you only get 7.7bits / pixel.
If anything you would want to encode your data in some way that abuses the P and B frames, and the macro block size of 16x16.
Coding theory for the data output at your end is only one side of the coin, the VP9 codec stupidly good compression is a completely different game to wrangle.
And I kinda doubt you'll get much better than your estimate of 1% from the original scheme.
It worked a charm.
Second round? A year later, when the archive was still available from umpteen hosts.
For all I know, it still languishes on who knows how many old hard drives...
"Real men don’t use backups, they post their stuff on a public ftp server and let the rest of the world make copies." -Linus Torvalds
Funny how these things work since I'm pretty sure I remember running into it around 2008 (i'm a few years younger).
I think i just deleted it though since I was suspicious of most strange files back then; I was the nerd who didn't have friends so i used to troll forums for anything i could get my hands on.
By doing that, I'm sure I stumbled upon my fair share of sketchy stuff unintentionally, so it's not hard to imagine the same for others :)
A little less crazy and more straightforward (software on audio tape was super common after all): Radio stations and vinyl discs that transmitted programs to the microcomputers of the time (C64, TRS-80 etc.) have quite a long tradition. Some examples:
http://www.trs-80.org/basic-over-shortwave/ https://www.youtube.com/watch?v=6_CZpFqvDQo&t=2s
Incredibly bizarre idea. I'm not sure who I thought would benefit from this. I guess I got swept up in RFC1939 and needed to build... something.
After some time the limit on single file was removed, but daily limit was set up to 100Mb. The trick is that POP3 traffic wasn't accountable, so we continued to use our "service".
Of course it was still possible to browse the internet and visualize arbitrary text, so splitting the .exe into base64-encoded chunks and uploading them on GitHub from another computer was working perfectly fine... I briefly argued against these measures, given how unlikely they are to prevent any kind of threat, but they're probably still in place.
Actually caused some problems for email clients, as they usually assumed emails were small. I got a few of them to crash with 200 Mb "attachments" (although this was in the early 00s, 200Mb was bigger than it is today).
https://github.com/fangfufu/Converting-Arbitrary-Data-To-Vid...
https://github.com/rekcuFniarB/file2png
https://github.com/nzimm/png-stego
Would be neater (and much more efficient) to encode the data such that it's exactly untouched by the compression algorithm, e.g. by encoding the data in wavelets and possibly motion vectors that the algorithm is known to keep[1].
Of course that would also be a lot of work, and likely fall apart once the video is re-encoded.
[1] If that's what video encoding still does, I really have no idea, but you get the point.
For example, when I upload a 4K vid and then watch the 4K stream on my Mac vs my PC, I get different video files solely based on the browser settings that can tell what OS I'm running.
Handling this compression protection for so many different codecs is likely not feasible.
But (almost) nothing prevents YouTube from not serving that particular codec anymore. This still pretty much falls under the "re-encoding" case I mentioned which would make the whole thing brittle anyway.
But it's indeed cool to think about. 8)
To compress data an arbitrary sequence of bytes, for each byte, you produce an image that your ML model would convert to the corresponding anchor vector for that byte and add the image as a frame in a video. Once all the bytes have been converted to frames you then upload the video to YouTube.
To decompress the video you simply go frame by frame over the video and send it to your model. Your model produces a vector and you find which of your anchor vectors is the nearest match. Even though YouTube will have compressed the video in who knows what way, and even if YouTube's compression changes, the resultant images in the video should look similar, and if your anchors are well chosen and your model works well, you should be able to tell which anchor a given image is intended to correspond to.
What you need is going to frequency domain. From my own experiment in university times most significant image info lays in lowest frequencies. Cutting off frequencies higher than 10% of lowest leaves very comprehensible image with only wavey artifacts around objects. You have plenty of bandwidth to use even if you want to embed info in existing media.
Now here you have full bandwidth to use. Start with frequency domain, set expectations of lowest bandwidth you’ll allow and set the coefficients of harmonic components. Convert to spatial domain, upscale and you got your video to upload. This should leave you with data encoded in a way that should survive compression and resizing. You’ll just need to allow some room for that.
You could slap error correction codes on top.
If you think about it, you should consider video as - say - copper wire or radio. We’ve come quite far transmitting over these media without ML.
For the sake of this discussion, wavelets are pretty much exactly that: A bunch of frequencies where the "least important" (according to the algorithm) are cut out.
But that's pretty cool, seems like you've re-invented JPEG without knowing it, so your understanding is solid!
Even if you train your model to work best for all existing codecs (I assume that's the "ML" part of the ML model), the no free lunch theorem pretty much tells us that it can't always perform well for codecs it does not know about.
(And so does entropy. Reducing to absurd levels, if your codec results in only one pixel and the only color that pixel can have is blue, then you'll only be able to encode any information in the length of the video itself.)
Now studios are using motion-picture film to store data, since it's known to be stable for a century or more.
[a] That way, in the future, if there’s any improvements to the transcode process that makes smaller files (different codec or whatever), they still have the HQ source
I wonder how long it'd take for Google to crack down on the system abuse.
Is it really abuse if the videos are viewable / playable? Presumably the ToS either already forbids covert channel encoding or soon will.
The process of creating and using the files is prohibitively unusable and so many better solutions exist that YT doesn't need to worry about it
You can probably do things like add frames that can't be decoded and so are skipped by a decoder; that effectively allows arbitrary added hidden data. That's maybe cheating.
If you stipulate that you can't already have a copy of the unaltered file, and the data has to be extractable from a pixel copy of the rendered frames ... that becomes more interesting, I think.
You'll notice this if someone has just uploaded a video to Youtube and the only version available for playback is some 360p/480p version for a few hours until Youtube gets around to processing higher bitrates.
So whatever you're encoding has to survive that transcode process.
Extending that into video files and it would likely be pretty massive, although you'd have some interesting time with youtube's compression algorithms
"compression resistant watermark" turns up some good resources for it. QR codes are another good example of noise tolerant data transmission (fun fact - having logos in a QR code isn't part of the spec, you're literally covering the QR code but the error-correction can handle it).
The best way I can describe it is that humans can still read text in compressed videos. The worse the compression/noise the larger the text needs to be, but we can still read it.
If creators start encoding their source and material into their content Google would probably be fine with that because it gives them data but also gives them context for that data.
Edit: I meant like "director's commentary" and "notes about production" type stuff like you used to see added to DVDs back in the day. Not "using youtube as my personal file storage". Why is this such an unpopular opinion?
it'd depends, as I don't think people using YT to store files would watch a lot of adds
Not true at all, lol. Google has a paid file storage solution. YouTube is for streaming video and that's the activity they expect on that platform. I couldn't imagine any service designed for one format would "probably be fine" with users encoding other files inside of that format.
You would hit Record on a VCR and the computer data would be encoded as video data on the tape.
People are clever.
edit: Video from the 8-bit Guy on how this worked - https://www.youtube.com/watch?v=_9SM9lG47Ew
There are many fun details about VHS, its chroma resolution, and especially some weirdness around the PAL delay line, but they all don't really matter for this.
Wikipedia says there is about 3MHz of bandwidth, so ~200-300 "pixels" seems like a very good ballpark (just going by the fact that a normal PAL signal has about 6 MHz and is commonly digitized as 720x576, 3MHz about halves the horizontal resolution and taking some pixels off for various reasons makes sense).
Damn rain.
I still have a C64 and tape drive.
There was a magazine in the 80’s where you could scan in the code with a bar code scanner.
Edit: OK, I see where this is going. Lol
From the README.md:
> Since Snapchat imposes few restrictions on what data can be uploaded (i.e., not just images), I've taken to using it as a system to send files to myself and others.
> Snapchat FS is the tool that allows this. It provides a simple command line interface for uploading arbitrary files into Snapchat, managing them, and downloading them to any other computer with access to this package.
4096 x 2160 x 24 x 60 is your theoretical max in bits/second, 127 billion.
Assume that to counter YouTube's compression we need 16x16 blocks of no more than 256 colors and 15 keyframes/second; that reduces it to
256 * 135 * 8 * 15 = 4.1 million bits/sec.
That's not too awful. Ten minutes of this would get you about 300MB of data, which itself might be compressed.
Just like 2K consumer video is 1920x1080 and 2K Cinema video is 2048x1080
On YouTube, the video and the description are also linked. They exist on the same page always.
And even if the concern this solution is covering is what if the video is somehow shared without the description, away from YouTube, then the video could just as easily contain the description or URL or QR code pointing to the file.
This is a just horribly unusable QR code.
- FacebookDrive: "Store files as base64 facebook posts"
- TwitterDrive: "Store files as base64 tweets"
- SoundCloudDrive: "Store files as mp3 audio"
- WikipediaDrive: "Store files in wikipedia article histories"
I like to think you could unify all of these into a FUSE filesystem and just mount your transparent multi-cloud remote FS as usual.
It's inefficient, but free! So you can have as much space as you want. And it's potentially brittle, but free! So you can replicate/stripe the data across as many providers as you want.
And as it was just storing emails it was even using gmail for it's intended purpose so no TOS problems..
It's called `M-x spook'. It inserts random gibberish that NSA and the Echelon project would've supposedly picked up back in the 90s.
;; Created: May 1987
So the things that the NSA and ECHELON would have picked up on back in the 1980s, not the 1990s :)For people who don't read Chinese: it encodes data into ~10M blocks in PNG and then uploads (together with a metadata/index file as an entry point) to various Chinese social media sites that don't re-compress your images. I knew people have used it to store* TBs after TBs data on them already.
*Of course, it would be foolish to think your data is even remotely safe "storing" them this way. But it's a very good solution for sharing large files.
It even has a full CRUD API, no need for using libgit.
although for propganda use, shortwave / sat tv is a much much simpler way to distribute information to place like that, but I belive now its hard to get one SW radio for anyone.
Users were given quotas of 5Mb for their home directory. He discovered that filenames could be quite large, and the number of files was not limited by the quota, so he created a pseudo filesystem using that knowledge, with a command line tool for listing, storing and retrieving files from it. This was the early 90s
Given that there are the options for uncompressed, lossy compressed and lossless compressed, I'd say RAW files differ in the stage of the data processing where capture is being done and doesn't imply anything about the type of compression.
What is relevant is that the formats vary widely between manufacturers, camera lines and individual cameras, so unlike JPEG, it's really hard to create a storage service that compresses RAW files further after uploading in a meaningful way. So anything they do needs to losslessly compress the file.
https://www.bbc.com/future/article/20160225-the-quest-to-sol...
Random example from an issue of Byte:
https://archive.org/details/byte-magazine-1986-05/page/n432/...
I've been uploading 2-3 hours of content a day every day for the past few years. On the same account too.
I have fewer than 10 subscribers lol.
But I guarantee there is some clause in the ToS that this project violates.
It's just recordings of myself when I'm doing deep work. I use OBS to stream my computer screen and a video recording of myself (mostly me muttering to myself).
It helps me avoid getting distracted (I feel like I'm being watched lol) and it's also interested to check back if I want to see what I was working on 3 months ago.
All the videos are unlisted or private.
Also, any potential issues with Google having access to proprietary code? I know the chance of any human at Google interpreting your videos is near-zero but still
Some forms of DRM are already essentially this, compression - and even crappy camera recording from a theater - resistant DRM that is essentially stegonagraphy (you can't visually tell its there) exist.
EDIT: "compression resistant watermark" is a good search phrase if anyone is curious
something like this but far more mundane
Others in these comments have also suggested steganography in both the video and audio streams. The problem with that is that when you retrieve a video from YouTube, you never get the original version back. You only get a lossy re-encoded version, and the very definition of lossy encoding is to toss out details that humans can't (or wouldn't easily) perceive, including ultra-sonic audio.
Or better yet, the file could be one third the size if the human says the numbers 0 to 7.
in other words; if YTs compression was affecting it so badly that it prevented the data from being re-read, wouldn't that compression scheme render normal video-watching impossible?
The video that is created in the example in the README is https://www.youtube.com/watch?v=Fmm1AeYmbNU
We can see that data is encoded as "pixels" that are quite large, being made up of many actual pixels in the video file. I see quite bad compression artifacts, yet I can clearly make out the pixels that would need to be clear to read the data. It looks like the video was uploaded at 720p (1280x720), but the data is encoded as a 64x36 "pixel" image of 8 distinct colors. So lots of room for lossy compression before it's unreadable.
(1) I was a freshman in college at the time, and Mathematica is one of the first languages I learned. (My physics classes allowed us to use Mathematica to spare us from doing integrals by hand.)
(2) I intentionally chose a language that's a bit obtuse to use. I was afraid that I might attract unwanted attention from Google if YouTubeDrive were too easy for anybody to download and run.
Would be interesting if you could treat each service (Youtube, Docs, Reddit, Messenger, etc) as a “disk” and stripe your data across them.
[1] https://www.cnet.com/tech/services-and-software/napster-hack...
(plus using more than one tld)
Just did a google and saw it had evolved over the years, used only the 1.0 implementation back in the days. For those on another nostalgic trip : http://hugolyppens.com/VBS.html
Horses for courses… this is how we end up with pictures clogging transaction ledgers
Good technical experiment though!
There are even Word files I've found that have complete file path notation to ZIP files.
So many ideas are flying to mind. Really creative.
Otherwise each frame would have to have a ridiculous amount of encoded overhead.
Ahh, NM cant even see that working.
edit: Maybe a file table at built from from specified first N frames, that delivers frameset/file map ...
Still nothing like skipping spots in a video. That relies on key frames and time signatures.
Cool stuff nonetheless...
Each frame gets the same amount of the file, about a kilobyte. So each frame is basically a sector. You need to read in a few extra frames to undo the compression, but otherwise it's just like a normal filesystem. And reading in a batch of sectors at once is normal for real drives too.
Even if you did need the frames to be self-describing, you could just toss a counter/offset in the top left corner for less than 1% overhead.
I wonder how many random videos like this are floating around that are encoding some super secret data...
Bet google isn't happy with this idea and will definitely try to break it asap
I could seemingly never explain the concept to other developers in a meaningful way or cared myself to code these out.
Anyway my quick summary in this is just think of a dialup modem. You connect to a phone line and you get like a 56k connection. That sucks today, sure, but actually it’s kind of mind blowing for how data transfer speeds worked at the time.
You know how else you can send data via a phone line without a modem? Just literally call someone and speak the data over the phone. You could even speak in binary or base64 to transfer data. It’s slow, but it still “works,” assuming the receiving party can accurately record the information and hear you.
That seems to be what this main topic is. Using a fast medium (video player) to slowly send data over the connection, like physically speaking the contents of other data. But there could be some problems with this approach.
Mainly, YouTube will always recompress your video. For this method, that means your colors or other literal video data could be off. This limits the range of values you can use in an already limited “speaking” medium.
if this wasn’t the case, we would like to use a modem connection. Just literally send the data and pretend it’s a video. However, where I left off on this idea, we appear to be hard blocked due to that YouTube compression.
We can write data to whatever we want and label it any other file type. (As a side note, Videos also are containers like zip that could be abused to just hold other files)
But YouTube is an unknown wildcard that changes our compression and thus our data which seems to invalidate all of this.
If we somehow convert an exe to an avi, The YouTube compression seems to just hard block this from working like we want. If we didn’t have that barrier, I think we could otherwise just use essentially corrupted videos to become other file types if we can download the raw file directly.
(steganography is a potential work around I haven’t explored yet)
Without these, we’re left to just speak the data over a phone which compresses our voice quality and in theory could make some sounds hard to tell apart. This leaves us in the battle of what language is best to speak to avoid compression limiting our communication. Is English best? Or is Japanese? What about German? Which language is least likely to cause confusion when speaking but also is fast and expressive?
This translates into what’s the best compression method for text or otherwise pixels in a video where data doesn’t get lost due to compression? Is literal English characters best? What about base64? Or binary? What if we zip it first and then base64? What if we convert binary code into hex colors? Does that use less frames in a video? Will the video be able to clearly save all the hex values after YouTube compression?
https://www.youtube.com/watch?v=VcBY6PMH0Kg
SGI IRIX also had something conceptually similar to this "YouTubeDrive" called HFS, the hierarchical filesystem, whose storage was backed by tape rather than disk, but to the OS it was just a regular filesystem like any other: applications like ls(1), cp(1), rm(1) or any other saw no difference, but the latency was high of course.
In the age of $5000 10 MB hard drives, this was the only sensible way to work with the 600+ MB of data needed to master a compact disc.
That's also where the ubiquitous 44.1 kHz sample rate comes from. It was the fastest data rate could be reliably encoded into both NTSC and PAL broadcast signals. (For NTSC: 3 samples per scan line, 245 scan lines per frame, 60 frames per second = 44100 samples per second.)
Dedicated controller could pack a lot more data, as in hobo tape storage system: https://en.wikipedia.org/wiki/ArVid
This was at a time when 3.5" floppy disks were expensive (and hard to come by), and hard drives were between 40 - 60 MB, so 130 MB was quite practical. The floppy drive in the Amiga read and wrote at 11 KB / s.
And yes, this was a DAC and an ADC in software, with added Reed-Solomon error correction encoding and CRC32. The goal was to be economical. The end price was everything; it had to be as cheap as possible.
This reminds me of the Danmere Backer.
"The entire hardware fit into a DB 25 parallel port connector and was easily made by oneself with a soldering iron and a few cheap parts."
This reminds me of the DIY versions of the Covox Speech Thing: https://hackaday.com/2014/09/29/the-lpt-dac/