Achieving full-motion video on the Nintendo 64 (2000) [pdf]
ultra64.ca
ultra64.ca
Disclaimer: I wrote PL_MPEG, but not the N64 port.
[1] https://github.com/phoboslab/pl_mpeg
[2] https://www.reddit.com/r/n64/comments/dr15py/i_just_started_...
I've actually since moved to port a H264 implementation to N64. It's been a long journey and I'm now at around 18 FPS, after nights of manual RSP assembly optimizations, vectorizing most of the intra-prediction and inter-prediction algorithms. I want to reach 30FPS so there's still some work to do.
Vector registers are 8 lanes, signed 16-bit, so they map quite well to per-pixel calculations on each plane (YUV), which is what video codecs do, as you can process 8 pixels at a time, and you have 16-bit precision to handle intermediate results.
The most complex hurdle is that RSP only has 4K of RAM so you need to DMA in and out macroblocks a lot (especially since I can't possibly rewrite a FULL h264 decoder in RSP assembly, not in this lifetime: I need to write only specific performance-sensitive algorithms, while the bulk of the decoder stays in C; this means that the same data ends up going in & out the RSP a lot, especially since the H264 decoder I'm using is not aware of this problem).
This said, RSP DMA is even rectangle based, so it's another perfect fit: I can DMA a macroblock by specifying the pointer in RAM, width and height (usually 16x16, but some algos works on sub-partitions of 8x8 or 4x4) and the stride (screen width), so that a single DMA call will transfer the block from the middle a frame, skipping the rest of the data.
Vector multiplications in RSP were designed to write DSP-like filters, so they map quite well to the pixel filters required by H264. There are several different multiplication instructions for different fixed point precisions, and there's even one that automatically adds 0.5 (in the correct fixed point precision) which is also a common pattern in FIR filters, and also used in H264.
Saturation (VCH/VGE/VLT opcodes) is also supported; this is useful as most algorithms eventually need to saturate the calculated value in the 0-255 range, so that's another thing which usually require 1 clock cycle for 8 pixels.
When working with 4x4 partitions, half of the vector lanes are ignored; when writing back to memory, you need to do a read / combine / write sequence (as you may want to write 4 pixels and keep the existing 4 pixels, but vector writes will write 8 pixels); in this case, the VMRG instruction is used, which basically allow to combine two vector registers into one, with a bitmask to specific where to get each lane frame.
For IDCT, it comes very handy that most RSP opcodes allows to do partial broadcasts of the lanes of one of the input registers; this allows to keep a 4x4 matrix into 2 consecutive registers and then play some tricks with broadcast to multiply by rows and by columns (which is required by IDCT where you need to compute A' x B x A, with A&B being 4x4 matrices, so if you expand that you will see that you need to rotate vectors a lot).
So well, it's actually a pretty good fit.
PS: in the Gamasutra article, it shows the RSP code used to do colorspace conversion (YUV->RGB). The article says that it give a big boost (and I can believe it: especially in MPEG1, CSC is like 30% of decoding time), but I brought it to basically 0% by letting the RDP do it (RDP is the GPU in N64). In fact, the RDP supports YUV textures: so in my H264 player, the RSP just does the interleaving (that is, merges the 3 separate Y, U, V planes into one) and then asks the RDP to blit a textured rectangle in the correct format. The RDP even runs in parallel to both RSP and CPU. It might be that, back in 2000, this wasn't fully documented by Nintendo, though I found several references in old Nintendo docs. I can't see otherwise why it wasn't used. Once you reverse engineer how to pass the correct constants, it works really well and brings the CSC cost to basically zero.
So it's quite fast, but you need to remember that the main RDRAM is shared among the main CPU and the whole RCP (eg: it's also used as video memory for textures and frame buffers by the RDP), so contention is really high.
I wouldn't have thought it's possible to get H264 running at reasonable speeds on the N64. Congratz for pulling it off!
https://www.youtube.com/watch?v=IiSFeqedcho
:-)
And also the Windows 95/98 startup screens.
There, we had about an hour of video - on CD, it should be noted, not a cartridge - at 288x320 (using an interlaced display mode, which virtually nobody used beyond splash screens), with a perfectly solid 30fps. Left unregulated, the decoder yielded around 40-70fps. All ARM assembly, and a huge amount of fun to write. ^_^
No hardware acceleration, needless to say, other than page blitting to copy the previous frame to the current buffer.
http://web.archive.org/web/20081221184231/http://www.gamasut... (edit, link fixed, thanks!)
If you like this kind of stuff, check out https://www.reddit.com/r/TheMakingOfGames/ and https://www.reddit.com/r/videogamescience/ I often post games-related stuff I find here to there. But, this might be the first time I’ve seen something posted to HN because it was first posted there :)
Does anyone know if this technique was ever used again on the N64?
edit://Visual difference minimal you say?
https://www.adriancourreges.com/blog/2015/11/02/gta-v-graphi...
Perhaps we cannot get much better perceived quality per bit, but I would strongly argue that we could improve other areas such as encoder/decoder latency.
I've got a little toy project where I am generating 1080p 24bpp frames in a C#/AspNetCore webservice and sending them to client browsers at 60fps over websocket. Nothing serious, I was just trying to test the limits of what I could do with current-year managed language and some crappy fun code. Think of it as a really bad attempt at Google Stadia. I currently see about 5 frames of latency with MPEG-1 using bidirectional pipes against FFMPEG. While this is an extremely encouraging first result, I feel there is some compromise I can make on quality+bandwidth in order to further reduce this frame latency (I.e. if I were to hand-roll some codec for this specific application).
If I could get the frame latency down to the ~30ms range on localhost, this opens the door really wide on some crazy project ideas I've had regarding novel real-time graphics approaches. I will definitely be printing this article off for deeper review.