[1] https://en.wikipedia.org/wiki/KGB_Archiver [2] https://en.wikipedia.org/wiki/PAQ6
I created a theoretical compressor which I haven't yet been able to implement which uses the fact that every sequence of bits appears in pi somewhere, so my compressor would just return the digit offset and the length of data. I keep looking for a source for all the digits of pi though, have yet to find it.
Alternatively, you could release a program that comes with an already long sequence of the decimal section of PI and connects regularly to the internet to download more digits (you could even run the computation & API on a google compute / amazon ec2 instance to keep costs low)
P.s. As the author states in issue #2, "the release date of pifs is an important part of understanding what's going on" ;)
https://blort.org/~kgasso/images/how-to-catch-script-kiddies...
I don't think that's completely true yet, it's just incredibly likely. I don't mean you can't build the thing, you just can't quite prove it accepts all input; it could work in practice before it works in theory.
Re: what komon said, if the offset dwarfs the input, perhaps you could find the offset of the offset, so on and so forth, keeping track of the # of indirections (until that number exceeds the input).
If you want further lossy compression for existing media files, look at the newest algorithms supported by ffmpeg.
If you want generic lossless compression, you can do a bit better than the usual suspects (gzip et al), but only if you're willing to put up with very slow compression times.
If you have some other type of specific data (e.g. sparse files) then you could do something custom, but I guess this is unlikely to be your situation.
Or the software linked via here: http://prize.hutter1.net/
As other people say, what you're asking for probably isn't possible.