Infinite-Storage-Glitch – Use YouTube as cloud storage for any files
github.com
github.com
A portion of the signal would be used for timing, metadata and error correction, so the program could tell you if the data was sufficiently damaged upon restore.
LGR has a video on the PC version from Danmere: https://youtu.be/TUS0Zv2APjU
Here's a video example of the Amiga industry's take on the idea: https://youtu.be/VcBY6PMH0Kg?t=573
Sony even did this in 1980 to record CD-quality PCM audio onto VHS tape. https://youtu.be/bnZFLzBO3yc
In a weird closing of the circle, I now store the internal sounds backup of my vintage Juno 60 synthesizer as a WAV file recorded from that tape backup output.
So the digital info of the internal synthesizers gets converted to analog audio in the synth, then passed as audio to my modern computer’s audio interface, which converts it to a digital representation of the analog audio.
And vice versa to restore the backup into the synthesizer’s memory.
Incidentally those backups are more reliable now than when using analog tape decks, since one doesn’t encounter physical tape degradation or a cassette deck “eating” the tape.
I haven’t done any testing with compressed audio formats, but I would expect even lossy formats to perform well, if one keeps the lossiness within certain bounds, so that the highest frequencies in the audio file are preserved.
MIDI as a compression format for that kind of audio data would be a lossy way to encode such an audio stream, and it certainly would perform well, so yes, such lossy formats do exist.
Most research in audio compression has been done on compressors that exploit the limits of human perception, though, so of the shelf lossy compressors may not do very well.
Yeah you're transmitting information about the source but to call it lossy is an understatement.
Modern synths still do this. The Korg Volca has a library for converting audio into white noise that reprograms/adds more samples.
This site describes the format, which was basically a header tone, a sync tone, data bits, and then a checksum (not described there but other sites say it was just an XOR). When we got a Disk ][ (5-1/4" floppy drive) all those issues went away.
http://www.applevault.com/hardware/apple/apple2/apple2casset...
Certain things had really distinctive sounds, like loading screens. I could recognise Manic Miner loading from just about any ten second chunk.
It was pretty finicky, though, and very slow.
Black and white dots in a strip on a card. Swipe the cards to load the games.
Here, for example, is a ruggedized S-VHS data recorder built for military and aerospace applications:
These days you get four channels of 96kHz 20-bit audio in a wee box the size of a Betamax tape, with hours and hours of recording on an SD card. That physical size is mostly a function of needing half a dozen XLR connectors on it and a big enough screen to see what you're doing.
(IMO there is not enough of these posts, and getting less over time.)
A refreshing "actual hacker" project that makes me look anew at the tools I always use...
So, my coffee maker is sending data to the net - maybe I can use that for backup, and have it replicated both in the fridge and in the living room lights...
But how would I retrieve that? Hmm. I assume that both Alexa and Google assistant are tracking everything that goes through my IoT devices. I'll ask GPT how to hack my Nest device to pull back data on demand, that oughta work, surely?! :D
Tangentially related and discussed in the past on HN: File transfer via color barcodes and a phone camera
[0] https://news.ycombinator.com/item?id=25459501
I did get a kick out of this from the OP: > Binary: Born from YouTube compression being absolutely brutal. RGB mode is very sensitive to compression as a change in even one point of one of the colors of one of the pixels dooms the file to corruption.
It's more than youtube compression -- video compression in general wreaks absolute havoc on our meticulously arranged (and sometimes colored) pixels. It's actually pretty fun/instructive to step through the transition between (what you want to be) two distinct frames when you're trying to (ab)use video for this sort of use case -- there are segments of the frames that get correlated and "flip" together first, resulting in in-between frames that are gibberish even with a modest amount of ECC in play.
0. Linux's SystemV Filesystem Support Being Orphaned https://news.ycombinator.com/item?id=34818040 by rbanffy 3 days ago, 70 points, 73 comments
1. TabFS – a browser extension that mounts the browser tabs as a filesystem https://news.ycombinator.com/item?id=34847611 by pps 1 day ago, 961 points, 185 comments
2. Vramfs – GPU VRAM based file system for Linux https://news.ycombinator.com/item?id=34855134 by pabs3 1 day ago, 226 points, 71 comments
I remember way back in the day someone came up with a clever way of using Gmail attachments to build a cloud storage drive mounted to your filesystem. Then Google themselves released Drive soon after.
Obviously this is so difficult to use that most people would rather pay $10/month to get 1TB of storage that can be very easily accessed. Even if someone has 100TB of data and wants to back them up, I don't they would do conversion to and from YouTube videos.
An interesting idea, but probably won't get much real world use.
https://en.wikipedia.org/wiki/AV1#Filters https://norkin.org/pdf/DCC_2018_AV1_film_grain.pdf https://waveletbeam.com/index.php/news/48-netflix-film-grain...
Google has been known to close accounts and "related" accounts for abuse (as defined by them). So even if you create another account, don't expect your main account to survive if there's any possible link between them.
They are the judge, jury and executor, so eff around at your own peril.
Ditto if someone gets hold of your phone and changes the login on your account, or they decide to not let you in because something "looks suspicious".
You are brave. I hope, for your sake, you have a local backup.
[1]: https://9to5google.com/2022/08/22/google-locked-account-medi...
https://www.zdnet.com/article/what-happens-to-your-g-suite-u...
The average person could buy a $100 external drive and replace it every five years, and that would be enough.
$240 annual x 75 years = $18000
Almost free huh?
$12000 a year x 75 = $90000
If I could pay that in and lock it in for the duration, maybe I'd consider that, but no one is going to let you do that.
Y'all got some funny notions on "Free".
Then there's the whole issue of "What if Google gets bored?"
https://workspace.google.com/pricing.html
The enterprise plan.
Enterprise is $20 (for me at least)
Price per TB appears to have fallen below $8. So that's $640 worth of storage. Basically, if you were to buy your own hard drives it works out to about $20/mo over two years..
A comparable Cloud Storage account on GCP with Coldline storage would be $320/month ($0.004 GB/month) or just $96/month for archival ($.0012/month).
The actual cost to Google is probably < $80/month for this 80TB ( most of the data is going to be in stored in a version of archival given the standard restrictions of 10TB on export.
80TB is also an heavy outlier, given the typical available bandwidth today and disk sizes commercially available for most users it will take a lot of dedicated investment of effort and time to upload this amount of data into the cloud.
Also Google's personal storage pricing is not competitive for pure storage, Backblaze is only $7/month for example. The higher price and value is derived from able to integrate into other Google products and provide storage for those like Gmail, Photos etc.
https://www.backblaze.com/blog/backblaze-drive-stats-for-202...
A well chosen model has an AFR of well below 1%. To get about say, 100TB, you'd need a dozen drives or so with ZFS and a nice enclosure. It is unlikely even one of them will fail in a given year and you will not experience data loss.
Here is a $100 case: https://ja.aliexpress.com/item/1005003125774264.html
Here is some YouTuber shoving 100TB into it: https://www.youtube.com/watch?v=boKmZKTKXHc
What a brave new world.
The 4x size increase is my biggest concern...too bloaty.
I want to expand this in into a fully modular service that you write payloads and scripts for various services, so when you upload a file its spread out across many different providers. When you're downloading, you just go down the list check what still exists, and verify the checksum. This should be stable for many years.
I plan to take a look into facebook and see what can/cant be accessed there. I had this exact thought with youtube and thought about using a pixel reader to exact out data. Same idea for different image hosting services like imgur.
Maybe you could join forces.
My favorite example of this was people storing files in "secret" subreddits by using posts and comments to store bytes. When they were later discovered by other users, the seemingly random strings sparked a huge conspiracy about their possible meaning.
However, you always have the problem that your unwilling host may remove your "files". I sometimes wonder about file storage using a textual output format that can't be distinguished from normal user interactions.
That was the first time I can across such a thing.
Someone even made an extension for Windows XP that allowed you to mount GMail as a storage volume.
> GMail Drive is a Shell Namespace Extension that creates a virtual filesystem around your Google Mail account, allowing you to use Gmail as a storage medium.
You could use a reproducible LM (for instance using Bellard's NNCP as basis), and encode one bit in one word by taking the {first, second} most probable next word.
https://blog.benjojo.co.uk/post/dns-filesystem-true-cloud-st...
https://news.ycombinator.com/item?id=16134041 (36 comments)
"Ask HN: What are these strange random strings spamming my blog?"
I guess it depends on what noise-to-signal density you’re after.
With a a long enough ChatGPT generated output, no one would question a few out of place characters or even an emoji. With 3000+ different emojis to choose from that encodes an entire byte of data.
Another idea is using “they’re”, “their”, “there” as bits.
It's a plot point in a Patriot Games, a 1987 Tom Clancy novel that introduced the term "canary trap" for this trick. He says he invented the term, but not the technique, which was already in use.
In a spat over the plot of Star Trek III (so, early 1980s), Harve Bennett distributed slightly different versions of the script, allowing him to track a leak back to Gene Roddenberry.
The book SpyCatcher says it was in routine use at MI-5, and you can find variations of it in lots of fiction too.
Makes me wonder if numbers stations are actually just the worlds slowest modems
If you encrypt the data and include a checksum or other identifying bytes in the ciphertext you can even have unwitting human participants in the discussions and if their posts are context your embedded data will be credible replies. You just have to be sure that threading behavior doesn't make it impossible to give the decoder identical context.
Well, with Chat GPT that's almost trival. POC https://imgur.com/fQvMh9S
And it’s simple: camera uploads automatically via FTP, inotifywait script uploads to google!
Alternatively, if they're allowed to use the footage to train some AI that will help them take over the world, then maybe they want all your random footage for free.
the modern qr readers are so fast and easy to use, its unbelievable
This guy extended the idea using fountain codes, which allows you to miss arbitrary frames and still recover the full message without waiting for the missed frames to re-appear:
teacher: 0b00010010101001, school: ...
and then the website can encode the data as a sentence and just text to speech it and the receiver can use whisper to speech to text and decode
will be the most creepy thing because it can be very steganographic and sound like a real sentence
Data Block: c3828abe
c5cfe61f4e61c9eda05e39903df580566859708a52957754e06fd18feaceca5ec0cdcac4b24b0f9ac8d9f212301916ea9ebcb2e291e2e950e0118f150c8cde02 34e770773cb93d6f1b757098890475cb00bef5ca4275c51021118ac1f01b71db3604063fd945480afc6b6b5b8d125129f7a9813a4997bdea27bbe5f6c17abfeb f46309c93430f78d37d23c0ef646cf7796e6de2b072d771b35b832a5b5328d1c09c5d32eaf6309b3119e8468ed02f62cd4b25c6785792ec82edc72667da8e36e 3b7b0d22fd708f5a3ff4787bf9474f84dff52fe33a38f4b4fee6759498b38d2c3af01db8d3dc5b1bb1cf6d203f24a4f6016caf42ad5cac76d1b0a0bf01a435b0 54a288c7cf9859dde401af51685eef23661ff0102a94caab2df9bf298c07538885baec81576513b9a7591d429db24b221c071cf0d929308243b0af4535810052
https://jis-eurasipjournals.springeropen.com/articles/10.118...
> "Moreover, most video-sharing channels transmit the steganographic video in a lossy way to reduce transmission bandwidth or storage space, such as YouTube and Twitter. . . Robust video steganography aims to send secret messages to the receiver through lossy channels without arousing any suspicions from the observer. Thus, the robustness against lossy channels, the security against steganalysis, and the embedding capacity are equally important."
I suppose in this project, the blocks of pixels are large enough to avoid data loss due to compression?
This is even easier, because jpg's ignore additional data past the end of the file. Post a low-res ~200kb jpg that has an additional ~20mb of data appended. It'll still render perfectly fine.
The other consideration is that Tumblr was always very “creator” oriented and while they might produce thumbnails of various sizes the original image is still available and not mangled by resizing algorithms. Other free image hosts are going to crush that image down the maximum amount tolerable to the human eye. Google even does that for paid photo hosting.
there should be some error correction in a system like this, though.
QR codes have 4 levels of correction you can use depending on how robust you wish them to be. CDs and DVDs use two chained, fixed, levels to keep the decoders simple. CDs have 25% overhead, but their correction is very strong: they can correct 4000 bits in a row.
There are two encoding modes, RGB and B/W. It uses a pixel-to-data width of 2x2, but says YouTube's compression algorithm is brutal, and one corrupted pixel already renders the whole thing corrupted.
Now, if you're using a 4:4:4 format to do this, then you should be able to use smaller chroma metapixels (I still wouldn't use the full chroma resolution, though, unless you're using a high bitrate or a lossless codec). However, that risks data corruption if passed through a pipeline that downsamples the chroma.
1. https://en.wikipedia.org/wiki/Delay-line_memory
Encode the data inside audio, preferrably outside human audible range, and then use a nice video of singing birds, or whales talking, and use the "hidden" frequencies to hide the data.
I don't know if Youtube has any filters that cut out frequencies, but this way they can't ban you, since you've uploaded a really nice personal video of your singing birds, instead of the conspicuous looking QR-like codes as in the OP ;-)
With any lossy audio compression algorithm, everything outside the human audible range is filtered away completely as a first step. That's compression 101.
Also there's much less bandwidth in the audio channel than the video channel, and then far less again if you're trying to hide a signal in another signal.
ADAT for example.
Does YouTube let you store unlimited video content (real video like screen recordings etc of our own work - no shady or sneaky stuff, nor any copyrighted stuff etc)
With all videos marked private ...so they are just "storage" by account owner and no other users can access them and youtube cannot monetize it ?
I wouldn’t use it as my ONLY backup of course.
I'd it's anything vital, as in your paycheck depends on it, I'd have multiple backups.
I had a good look into these sorts of technologies but the host almost always changes the file so it makes it impossible to retrieve the data hidden in the file.
You need a file hosting platform that guarantees not to change the uploaded file.
How does this avoid such problems ?
The assumption is that video compression won’t mess up those blocks beyond recognition, so you should retain the information as long as the rendered resolution and bitrate don’t drop too low.
Maybe this could be improved by e.g. using 32 colors instead of 2, and bumping the block size to 3x3 (for safety) which should yield ca 144KB per frame.
Seeing it come to life has just scratched a long forgotten itch and damn it feels great.
Yeah, I really like this stuff. Awesome project.
I chuckled because of my own thought that seek (FS call) can be implemented via youtube video seeking
pack-5d55e1f4809a8dae84591bc04b019b2ae8137f77.pack 146MI do love these kinds of hacks, but I hate these kind of weaselly cop-out statements. You made the tool, own it!
If anything, the author should be more clear about what happens if youtube gets mad: you might lose your google account along with access to mail, drive, photos etc
A true hacker spirit worthy of Captain Crunch whistle and its application toward free payphone calls.
jesus christ enough with the jerking
nice end of transmission simulator to boot!