The Stick of Jan Sloot (2004)
spronck.net
spronck.net
From my (rather limited) interactions with him, I'd say he was far from genius, I feel he just got sucked in and had no way out of his network of lies without losing face.
From wikipedia
There is a book on it with lots of details..
In any case, it's one of those crazy stories.
You'd give everybody DVDs of digits of pi (or they could calculate them themselves) , and then transfer files faster by just sending them the offset into pi.
At the time I thought it could work with a big enough bank of digits of PI on both sides. If transfer was expensive, and calculating digits was cheap then you could give everyone an infinite supply of digits of pi and have a nearly infinite compression system.
I discovered that often the offset into pi is much larger than the data you are sending. Turns out it's an expensive way to sent things.
Also, it turns out that this area was already well understood. There are no free lunches with entropy.
But it was a fun idea to kick around.
Although a flat memory space / full file approach would be nice, I wonder if a paging scheme would be possible?
Where you would only need all possible finite strings of length N, where N would ideally be the max page size of RAM, but could be less if the substring could not be found within a reasonable offset...
Ignoring compression speed, I suspect it'd have a practical limitation where using a maximum offset would mean that some strings wouldn't be as big as the cell needed to describe them. That's usually the gotcha in schemes like this...
In order to entertain this, I think all possible 12 byte strings (64bit offset + 32bit size) would need to be found within whatever a reasonable amount of storage is.
I suppose the computational challenge is "How big an offset is required to have all possible strings of length N?"
If backwards time travel were possible then your "compression algorithm" could simply be deleting the file. To recover the file go back in time to before it was deleted and make a copy.
With this you can "compress" giant files down to just a short description of a place and time where the file was on your computer.
Aside, I wonder if "Pied Piper" from the Silicon Valley series is a reference to him. Probably not but fun to entertain.
Was This Lost Computer Code Worth Billions? (Jan Sloot Digital Coding) (2020) [video] - https://news.ycombinator.com/item?id=36499676 - June 2023 (2 comments)
The Stick of Jan Sloot (2004) - https://news.ycombinator.com/item?id=29623524 - Dec 2021 (22 comments)
Ask HN: What was the secret that Jan Sloot took with him to the grave? - https://news.ycombinator.com/item?id=13443135 - Jan 2017 (4 comments)
The Stick of Jan Sloot (2004) - https://news.ycombinator.com/item?id=8699058 - Dec 2014 (18 comments)
I wonder if one was scammed and bought a Blu-ray with everything, only to find out they got novels rather than movies. After the initial disappointment they might realize well written novels are better than movies and never watch another.
How about we get GPT to turn good movies into great books? Like a large inverse prompt engineering problem.
Could we then even feed that text as input to make a better, though arbitrarily different movie?
GPT has trouble making great sentences, much less great books from a data type completely different from its training. GPT, at the core, is a "likely next word" generator. No great literature came as a process of "likely next word" in a vacuum. Likely next word from complex experiences of a decade and plot designed from those complex experiences, sure. But not the algorithm that GPT is.
> Could we then even feed that text as input to make a better, though arbitrarily different movie?
Nothing about a feedback loop of mediocrity describes a practical way to improve quality. This concept is as sound as the alien encoding stick and the sloot encoding scheme itself.
Deconstruct a movie into a set of 3D models, textures, voices, parameters for those 3D models' movement, procedural generation of background objects like trees, grass, sea, sand, weather conditions, etc, etc. Like the data that forms the content of modern, near-photorealistic games.
A game engine-like rendering system would then render a scene, tweak parameters until rendered scene / frames match the original movie closely, and work through the rest of the movie the same way.
When done, one could produce derivations of such movie by changing actors' voices, have them move differently or speak in another language, swap out buildings or other objects, reduce polygon count or resolution of textures, have a cornfield show a little taller stalks, put the sun in a different angle, etc, etc.
Of course this is way beyond current compute capabilities. Not to mention software frameworks to do this job. But in theory this should work. For a 2min trailer it would probably be pointless. But for a 2..3h movie, maybe not.
And yes, of course this would be lossy 'compression'. Just more high-level than current video compression schemes.
For an audio analogy: compare mp3 compression with MIDI + quality sound banks for every instrument under the sun + parameters like how hard a piano key was struck, etc. Vary such high-level parameters until rendered output matches the original music.
I don't think it is totally inconceivable. It is definitely possible that there are hacks such as compression that would allow a simulator to run at above 100% speed. It also might not need to be 100% accurate -- we might find that running below that threshold still produces an output that is practically indistinguishable from the real events, yet allows it to run.
You could create a kind of "reverse" grandfather paradox by preventing future events predicted by the simulation from happening. Perhaps it could run fine up to the point that it has to simulate itself, then it would slow down to the point of becoming useless.
If the universe is based on some solid deterministic rules with a fixed seed value, then it would work.
Quantum events appear non-deterministic, but that might just be because we don't have all the answers right now.
However, interestingly, by scanning other universes you may find similar movies with slightly different plot turns.
What's encoded is not the audio - even the best audio codecs need on the order of 4kbits/sec to encode legible speech - but the actual semantics on how voice is produced.
Suppose you want to do this for movies. 8 kilobytes doesn't sound like enough, though the script of the movie could easily be compressed to that. But it's posible to imagine a system into which all the skill of the cinematographer, director, script writer, actors, etc. are built, with from 8 kilobytes of instructions could create, perhaps not the original movie, but something comparable.
Is this doable? Does some random crank have the remotest chance of pulling this off with the technology of the day? No. But that's likely the line of reasoning employed.
The cheating in this context would be to study the thing the file represents rather than the representation.
A dumb example would be to make a 10 hour movie from a single image that doesn't move. There is no reason for the file to be larger than the original jpg.
> I can not only get Orson Welles' Citizen Kane. I can get Citizen Kane in colour!
https://www.youtube.com/watch?v=JI5qy9Zoh_0
To argue that this was not the method Sloot used is missing the point. The question is: How to do it, not how to imitate someone else.
In his demo Sloot was playing 16 full movies simultaneously on a 1995 laptop at any speed. A high end computer had 32 MB memory, 133 MHz cpu, PCI video cards had 4 MB ram, 66 MHz, 560 MB HDD
If it was not what he said it was why didn't he just sell what he had? Without the extraordinary claims the demo already requires cartoon physics. He drives the truck into the match box, making an U turn inside doesn't at all seem necessary???
https://sound.media.mit.edu/resources/mpeg4/sa-tools.html
It was a growing idea in the mid-2000s but AFAIK it has gone absolutely nowhere. Essentially, instead of somehow encoding the audio, you encode a description of how to generate the audio.
Given the notch = (A+B) / 10
You can only recover A + B. You can't recover A or B individually.
Whereas Arithmetic encoding is actually practical, extensively used, and a direct analogue to the stick.