Memories – 256 bytes demo winner of Revision 2020
sizecoding.org
sizecoding.org
When trying to make my significant other to understand what was happening I wanted to run it myself. I was amazed how simple that was!
- Install the assembler[2]
- Install dosbox[3]
- Get the source[4] and put it into c:\temp\demo\memories.asm
- Start nasm and enter:
cd c:\temp\demo
nasm.exe memories.asm -fbin -o memories.com
- Start dosbox and enter: mount d c:\temp\demo
d:
dir
memories
- Press [ALT][ENTER] for fullscreenThe dosbox config is not optimized but it runs with sound with the default settings!
For me this is somehow much more impressive than simply watching the video.
[1] https://www.youtube.com/watch?v=Imquk_3oFf4
[2] https://nasm.us/
[4] http://www.sizecoding.org/wiki/Memories#Original_release_cod...
brew install dosbox nasm
nasm memories.asm -fbin -o memories.com
dosbox
mount D ~/Development/memories (or whatever)
D:
memories
So, almost the same!Hit FN+Ctrl+F12 to speed it up (it's time-independent, for smoother animations, hit that combination quite a few times).
It didn't output any audio for me, but that's probably fixable.
It will be fun to replace "sleep 5 && echo ok" with an automated version of such writeup :)
There is a breakdown of this demo on youtube, it's roughly ~2 hours. They explained how they made this demo. Really interesting to watch.
Demo: https://www.youtube.com/watch?v=QhqT0DhV9yE
Breakdown: https://www.youtube.com/watch?v=hFIyj5Yv440
It's interesting that the demo scene is very Windows/DOS focused, unlike other hacker scenes. Linux or Mac demos are basically not a thing. You're far more likely to see C64 or Amiga demos.
Back in the day we wanted to have the cool demos at the parties to show off how we managed to do something that others deemed impossible.
Part of the challenge is to match what others did and out win them, without having access to how they achieved it in first place.
Older Macs were never a thing in Europe given how expensive their were, Comodore, Atari and Sinclair machines ruled Europe.
While UNIX and demos never were a thing, those were more the domain of cracking competitions. Trying to gain access to some random university server.
Řrřola's trick. A bit like https://en.wikipedia.org/wiki/Fast_inverse_square_root
More clearly: DI = (y * 320) + x
Multiply by 0xCCCD => (y * 0x1000040) + (x * 0xcccd)
Take top byte is equivalent to divide by 0x1000000. So that gives you Y. The next lower (third) byte is then (x * 0xcccd / 0x10000) == (x * 52429 / 65536) =~ (x * 256/320). And the lower two bytes are noise.
That said, those demos are truly impressive.
The entire video is here [2], if you just want to watch the final product.
EDIT: link fixed.
Also, your Pouët link is wrong. I think you actually meant https://www.pouet.net/prod.php?which=85227 instead.
For reference, the original was https://www.pouet.net/prod.php?which=78044 which appears to be a 64-byte demo showing off just the raytracer from 85227.
Seems like this would be perfect for videos that are just e.g. gameplay of games made of tiles+sprites: the video could just store one copy of the assets, and the frames could just be tile maps + sprite position information.
It would also work well for “videos” that are really just a single static image. Or videos that are visualizations of the audio stream: the VM could actually take the audio frames as input and output the respective video frame.
An application state-data format can only be decoded by the original application, because necessary context—in this case, the game engine that translates user input to game-state and then to displayed frames, and also the library of visual assets the game uses to render those frames—is in the application, rather than in the video.
A video format is self-contained, and usually not domain-specific. Many encoders and many decoders can be written to target a video format, and the decoders should not have to ship with an asset library (let alone a game engine) in order to properly render specific videos.
A format like I'm talking about—one that doesn't know anything about application state, but does understand that it's compositing and placing a set of embedded assets each frame, rather than only knowing about pixels/gradels—seems like something generically useful to me. (Heck, we're close to support for such a format already, since many video players already understand the idea of compositing arbitrary stuff with placement instructions on the screen each frame, care of support for the https://en.wikipedia.org/wiki/SubStation_Alpha subtitle format. That format is exactly the kind of "vector video" I'm talking about, except the only primitives it can position and style are text elements. Add RGBA-textured rectangles as another primitive type to it, and you'd get a video format!)
And yes, I'm basically talking about the visual equivalent of a https://en.wikipedia.org/wiki/Module_file (embedded samples/synth patches + sequencing information); or, if you prefer another analogy, "what Flash movies are if you exclude the ability to execute ActionScript."
https://www.youtube.com/watch?v=_YWMGuh15nE ("Elevated", 4k)
https://www.youtube.com/watch?v=fp0t2jCMGZE ("One of those days", 8k)
Also discussed previously at https://news.ycombinator.com/item?id=14164907
- the graphics "driver" reads values out of certain registers (AL and AH?) at a set interrupt (maybe every X clock cycles?) and writes one pixel to the screen of whatever color those registers had in them
- by writing values into those registers and aligning the number of operations the program does with the frequency of the interrupts, you can get animation?
Even achieving any sort of flow control so you can switch between the effects is mind-boggling to me.
This is sixteen-bit assembly, so you have the famous 640kb of RAM available to the user and a 64k bit of RAM beyond that (see "0xa000" in the program). The graphics hardware is continuously rendering frames out of there at 320x200, one pixel per byte, using the default system palette.
The rendering is rather like a pixel shader. There is a big for loop over all the pixels, and at each point it computes a pixel value. First it decides which frame number it is on (stored in BP register I think), then calls an "effect" for that pixel.
It then jumps three pixels. This gives that nice "dissolve" transition between effects.
Keyboard controller is wired directly to the bus, so you can read the keyboard with a single instruction.
A MIDI controller is wired directly to address 0x330 (not standard equipment, back in the day this required a Roland card or SoundBlaster 32?), so you can just write MIDI to that.
There is a system timer interrupt configured for the music. The graphics appear to run continously, I can't see a link to the timer or vertical sync in the graphics code, that appears to just run continuously.
It's actually much simpler than that. After you set the right graphics mode (which for most simple dos demos is usually mode 13h, 256 color on 320x200) then there's an area of the memory that you can write to and it will show up as a pixel.
The "flow control" is usually just that you run your effect n times in these simple demos. Which means it will run faster on a faster CPU, but you usually wouldn't bother implement any form of timing in 256 byte.
MS-DOS programming was overall a pain in the... byte but what I miss most about it was the simplicity of graphics.
Wanna draw? Just write to memory. Setting a mode was one instruction
(Wanna play sound? Fumble with 2 levels of IRQ controllers one DMA controller then sob uncontrollably. Or use Allegro. Wanna do multithreading? What's that? )
nchelluri@grugbarn:~/dev/hello $ cat > hello.go
package main
import "fmt"
func main() {
fmt.Println("hello world")
}
nchelluri@grugbarn:~/dev/hello $ go build
nchelluri@grugbarn:~/dev/hello $ strip hello
nchelluri@grugbarn:~/dev/hello $ ./hello
hello world
nchelluri@grugbarn:~/dev/hello $ du -h
1.4M .Here's another awesome 256b demo that I love: http://www.pouet.net/prod.php?which=66372
https://www.pouet.net/prod.php?which=78045
I ported it to a boot sector so you can run it with a single (rather long!) Linux command line in qemu:
https://rwmj.wordpress.com/2019/12/08/pyrit-by-rrrola-incred...
The source code for Pyrit is worth reading too (see first link). It's very clever and quite readable.
Later: Your question made me wonder what the performance of virtual 'target CPU' is - the 'cycles' setting in the config is 20000 and there's a rough estimate of what these numbers translate to here
https://www.dosbox.com/wiki/Performance
So it looks like it's something along the lines of 'a 486 in the prime of its life'.
Usually you don't "handle" it in very small demos of the 256b/64b kind, you just run your effect. And yes this means speed will depend on the CPU speed.
I used this to slow down my computer so that old games were playable. Hooked into the interrupt, wasted cycles, and could enjoy the game. :)
You may argue that JS is sandboxed, but so is DOSBox. At least DOSBox can’t easily connect to remote servers over the internet.
Given that, it's not much of a sandbox.
Or does that require intervention from the host system rather than auto-mounting home and similar?
If we want to run code in browser there is WASM.
So is the proposal that it would it be beneficial to have a DOS-like OS or x86 emulator in WASM for running COM files?
Yes, that would be better and more sandboxed than dosbox running outside the browser.
Let's see if HN accepts the URL:
https://floooh.github.io/tiny8bit/c64.html?prg=AQgLCB4AnjIwN...
another reminder of the 'art' of the demoscene and it's recent recognition as a piece of UN heritage in Finland which I thought was pretty cool http://demoscene-the-art-of-coding.net/2020/04/15/breakthrou... (HN discussion: https://news.ycombinator.com/item?id=22876961)
But, what I did last year was visualize one of the assignments; the assignment was something about overlapping areas on a field, the naive solution (for me) was to create an x by y bitmap and just add the overlaps, which could then easily be converted into a visible image, which helped me with visualizing the problem and my solution.
https://www.pouet.net/toplist.php?type=32b&platform=&limit=5...
For handy reference:
8^8: 16,777,216
8^16: 281,474,976,710,656
8^32: 79,228,162,514,264,337,593,543,950,336
8^64: 6,277,101,735,386,680,763,835,789,423,207,666,416,102,355,444,464,034,512,896
8^128: ...
8^256: ?
8^512: hello from the other side of the quantum dimension
8^8 sounds interesting. 16 million reboots of a real {PC,C64,ST,Amiga,Mac,Z80,...} sounds like a collectively highly entertaining kind of hilarious. The issues only begin when you start wondering if any of the programs wedges the hardware into "interesting" states that are preserved across reboots - or at least the what if of that dimension of entropy... then the problem space becomes 8^8^8...
[I decided to compute 8^8^8. The result is apparently 15 million digits long. (`echo 8^8^8 | bc -ql | wc` -> `222814 222814 15596963`)
The smallest category in pouet for reference is 32b (or 256 bits), so 2^256 combinations to brute force. For comparison usually 128bit encryption is considered "safe" and infeasible to brute force.
You might be able to constrain the search space to only valid IA32 instructions, but realistically I don't see it helping that much
It'd probably not constrain the search space nearly enough though.
But even if it did and you'd somehow manage to even generate every combination, you'd still face the second problem of how to evaluate if they do something "interesting enough" to be worthwhile reviewing.
My original comment ran the numbers against the OP's "wouldn't that be 256^256?", but I got tripped up by the reply refuting that and saying it was 8^256 instead.
For as-yet unknown reasons my brain has always had a hard time mapping between the real world and the mathematical vacuum, so it was honestly less stressful to risk trusting that comment than try and [figure out how to] figure it out on my own. So I just substituted calculations for 8^n.
Here are the original numbers I supplied:
8^8: 167,77,216
16^16: 18,446,744,073,709,551,616
32^32: 1,461,501,637,330,902,918,203,684,832,716,283,019,655,932,542,976
Thanks.
Hopefully I can figure out those mapping problems one day. I think neurological damage may be involved, or something - I had to resort to button-mashing on my calculator while trying to figure out how many vegetables I could buy for $X given that they were $Y/kg one day at the supermarket. I'm 29. </rant>
I think of it just slightly differently -- 256 bytes, 2048 bits -- so 2^2048 (same result as your 256^256).
To give people (who don't spend time with these numbers all the time, I do b/c of cryptography): 2 ^ 256 is on the order of the /number of atoms in the entire universe/ -- every star, moon, comet, black hole, galaxy, etc. across the entire known universe)
Now consider this: 2^512 -- take every single atom in the universe, and imagine that that atom //contains a universe of atoms//. Congratulations, you're only at 2 ^ 512.
Imagine how large 2 ^ 2048 is!
First to set the video mode, and then to set up a timer used to progress time.
There is even more debate in 4k. After all, most rely on graphics drivers that take hundreds of megabytes. But the thing to understand is that in any case, the intro ships all the code that produces the sound and image. The OS is just an abstraction layer. The exception would be fonts and MIDI instruments, that can be stored in the hardware or OS.
But not all intros have text, "Memories" doesn't. And many intros do their own sound synthesis, though in PC 256 bytes you are usually limited to MIDI or to that horrible buzzer.
Anyways, great job. I was there during the compo, it was epic, with everyone double checking the executable size, even the old guys who have seen it all. You got my vote BTW.
Things like self modifying code, using bits of the bios or video ROM in ways they weren't intended by jumping into the middle of them, saving space by using code as data or vice versa, tiny packers which compress or uncompress the code, massive pregenerated buffers to do runtime lookups to generate data in one order but use it in another, etc.
This seems very vanilla in comparison...
Also, the tiny unpackers are generally used from 4096b and upwards. The size of the unpacker takes too much space and doesn't make up for the compression ratio at 256b.
- Jumping into the middle of video ROM and/or the BIOS
- Using substantial amounts of code as data
- Pregenerated buffers and data reordering
oooooo.
But all that doesn't really cut the space down in something as "big" as 256 bytes, it's the approach and the algorithms that do =)