Nintendo 64 Architecture – A Practical Analysis
copetti.org
copetti.org
https://twitter.com/DebugSteven/status/1054903603985559553
So I guess what I'm saying is that I have pretty hands on knowledge with the system and would be happy to answer any questions I can.
One thing I'll throw out there, is that one of the biggest limitations of the N64 (its 4KB texture memory) gets called a texture cache a lot, but that's a misnomer. It's a manually managed piece of memory, and (IMO) the system would have been much better off if it were actually a cache rather than having to load an entire texture in regardless of what was being sampled. Nowhere I've seen in Nintendo's literature do they call it a cache either. The crazy hacks that Rare did to subdivide their geometry on texture boundaries wouldn't be necessary for instance. I'd maybe even be into a 2KB cache over a 4KB chunk of manually managed memory.
One other aside is that I think the system still has tons of unlocked potential. So much of unlocking it's power seems to be centered around memory bank utilization. Switching which page of DRAM within a bank is expensive in terms of latency, but it seems like if you allocate your memory in 1MB bank chunks you can get around a lot of the limitations of the systems having the slow memory that developers complained about at the time. I don't blame developers at the time, they were coming from SNES where it was single cycle access to RAM, to the N64 that had a very deep, very modern memory hierarchy and what all that means for your code. The industry as a whole didn't really catch on until about halfway through the PS2's development cycle. But applying some of those PS2 techniques back, the system really purrs when you have dedicate a 1MB bank to each streaming source or destination. I can't wait to see what crazy stuff happens when the demoscene folk really start to get their hands dirty with it.
The source behind the text of the most recent slides starts here if you'd like to read it: https://github.com/monocasa/n64-slides-apr/blob/caae25f397c5...
I should probably save off pictures and put it into a pdf or something where it could be more easily accessible.
https://github.com/monocasa/n64-slides-apr/releases/tag/v0.1...
I've only tested it with cen64 and real hardware with a 64drive.
A is forward through the slides, B is back.
I would love to know more about this. Is their texture format tomfoolery written up somewhere?
So TMEM is only 4k. If you want mippmapping, that eats half of it, down to 2k practically. That leaves you with enough room for 1 32x32x16BPP texture at most. So I think what they did in some cases was to take a mesh with a larger texture, and run it through a processor to tesselate the mesh on smaller texture block boundaries in UV space, so they can render each tile of geometry with the same texture block at once, then swap to the next subtexture and render all it's geometry. That'd give you an apparent larger texture than you could fit into TMEM, and is one reason (of many) why their games look so good. They also might not have had tooling for that and just brute forced it by hand, I can't tell just from looking at the wireframe.
https://twitter.com/007goldeneye25/status/109829415491907174...
The rabbit you need to catch beneath the castle in Super Mario 64 is named "MIPS". :)
Aw, they should have brought it back and renamed it ARM…
It might circle back around to being exotic again, if FPGAs take off and more work is done on highly-specialized task-specific hardware, as opposed to building the fastest general-purpose chips you can and beating problems to death with sheer speed.
Sure. The Afterburner Card that Apple is selling to accelerate ProRes decoding on the newest Mac Pro is an FPGA. They’ve opened it up to third parties like Red as well iirc.
I know Gigabyte back in the mid 2000s made a PCIE Card that let you use DDR as a disk drive once upon a time; for the original they actually used a Xilinx Spartan FPGA since it was a smaller run.
The FPGA also can have a massive unique key that allows the designer to create a whitelist algorithm that only lets certain unique IDs run that firmware. Other options involve setting a time limit for how long the firmware will run, disabling certain features, or totally bricking that FPGA forever. Spartans have this feature but it would still allow for someone to build a new design that doesn't check the device ID.
Additionally, the bitstream can be encrypted so that if a field update is necessary or the firmware is stored in a stored in a separate flash chip, someone can't reverse engineer it.
Overall, the more you pay, the more security features there are available. An example secure design would disable JTAG pins permanently and have a microprocessor inside that would handle new updates. The processor would authenticate any new encrypted firmware before programming the internal flash.
The issue with this is that FPGAs are expensive to buy as a hobbyist, expensive to buy at small quantities unless you can negotiate with Avnet or similar, and involve using software from the past if you want to program them.
Beyond FPGAs, ASICs are fairly common in very high margin/high volume electronics.
My main use for C and C++ are actually their shading language derived dialects.
http://www.qwertymodo.com/hardware-projects/n64/nonvolatile-...
Lol this stood out for me (although probably just semantics...):
"Reality Co-Processor running at 62.5 MHz." What big dreams we had back then to try and"simulate reality" with only 62.5Mhz :)
Well done author... well done Nintendo !
I bought one on eBay, and it's a quirky device. It's essentially a glorified floppy drive with proprietary disks, and it makes loud, albeit amusing scanning noises.
Every Nintendo console prior to the Wii had some sort of "expansion" or upgrade capability for future peripherals to be added. Hardware revisions took the place of this model.
Plus, adding more texture memory would have meant respinning the RCP silicon. That would have been a significant expense for marginal returns.
Besides, Nintendo had already sold a memory upgrade in the form of the Expansion Pak. A second memory upgrade which couldn't even be installed into existing consoles would have been a very difficult sell, and could have soured customers on their future consoles. ("Why buy a Game Cube when they'll just release a Game Cube Plus next year?")
the https://en.wikipedia.org/wiki/Nintendo_64_accessories#Expans... raised main memory from 4mb to 8mb.
it was required for one of the Zelda games, but also helped with graphics for some games if it was inserted.
seemed to be a decent success.
https://www.reddit.com/r/n64/comments/ft63zn/dk64_memory_lea...
The lead artist claims they’d always been planning to make use of it. There’s compelling evidence of a memory leak or similar issue since the released version of the game apparently crashes if you leave it running for some amount of time over 10 hours (which wasn’t usually a problem in practice except for very long speedruns until it was released on Wii U Virtual Console and people started using save states which don’t reset the timer the same way as turning the game off and on does).
I _think_ it was on the documentary clips shipped with Rare Replay, but I'm not 100% on that.
That being said, going to a real cache rather than a block of manually managed memory I thin would have been the better design choice. Most textures didn't fill up the memory because of the practicalities of managing that memory. A double buffer scheme to load a block wile rendering from another one, the fact that you have to eat the whole cost of the full texture's load before you can render from it, etc.
It was a multifaceted problem, and was ultimately a design flaw/oversight rather than someone saying "I think 4kB is enough memory to store all the textures". The problem is less that the cache was small, it's more that Nintendo's plans for how awesome RDRAM and a unified memory architecture didn't pan out.
Problem #1: There was no dedicated video memory. All RAM on the N64 was shared RAM. So framerates tanked if you didn't have most of your stuff in cache. Keep in mind the framebuffer also lived in this unified memory area, so the video chip was already very noisy on the memory bus.
Problem #2: The unified shared system RAM was RDRAM, not SDRAM. And the latency on RDRAM is absolutely terrible. So the already expensive cost of using RAM was compounded.
If the N64 did what the playstation and saturn did and just have dedicated video/system RAM, and made this RAM relatively low latency SDRAM instead of the relatively high latency RDRAM, this 4kB limitation wouldn't have mattered.
Larger cartridges (i.e. 32/64MByte) gave them space in ROM to play with tiled textures. Usually this -did- also involve use of the 4MB RDRAM upgrade.
#1 It was still slower than the cache.
#2 You were still using the single shared bus. You would still be using cycles which contribute to data stalls elsewhere in the system.
#3 ROM was expensive. N64 games were typically in the ballpark of $10 more expensive than Playstation or Saturn games because of the manufacturing expense.
#4 I don't fully understand why, but it was all or nothing. You couldn't have uncompressed textures in ROM but also gain the benefit of the cache. Maybe the cache invalidation was poor or something. I wish I knew more.
Later games were more likely to go this route because ROM was cheaper. (Moore's Law and all that)
Additionally, the TMEM could only be loaded from RDRAM, not directly from the cartridge. I think the RDP's DMA master is only connected to the RDRAM slave port and not the main system's bus matrix.
So going back to it, games would a lot of the time store compressed data with a simple algorithm that could run out of the CPU's cache. Then the scheme looks like
* Cart->RDRAM DMA of compressed texture
* CPU decompresses texture into another RDRAM bank, and can be considered a RDRAM->RDRAM transfer. Sometimes the RSP handles this instead. I'm not sure if you could load straight out of RSP DMEM to avoid another bounce to RDRAM. I don't think XBUS works that way, but I could be wrong.
* RDRAM->TMEM DMA of uncompressed texture
Interestingly, games with more advanced texturing schemes like Indiana Jones tended to use uncompressed textures. They did this to avoid the decompression step and it's bandwidth. At that point it's just staging the texture with that cart's DMA, and slurping that into TMEM without any other processors eating bandwidth in between.
The Playstation by comparison used EDO (Based on eyeballing the pictures on wikipedia, baseline was 70ns/60ns for CPU and Video memory.) But, It's main bus was under 133mb/sec, and the fastest it could read from CD was 300kb/sec.
EDO Memory would have kneecapped the N64 from a memory bandwidth standpoint. The cartridge bus alone is over 200mb/sec. SDRAM -might- have done the job but may have wound up being more expensive; PC-66 (we are at the infancy of SDR in 1996) would have meant a PCB with 8 chips laid out for the parallel bus. To be frank I'm not sure Nintendo could have even gotten such a configuration (i.e. 8 512KB PC-66 chips.)
RDRAM was definitely a design compromise, but in retrospect I understand it's use in keeping overall costs down.
Dedicated video ram would have been a better option however, but I think it was another cost issue.
The N64 launched at $200, the Voodoo at $300. Of course you would additionally need a computer to run the Voodoo, but I remember thinking the N64 was already way too expensive back in the day. It would've been even more expensive to support a 64-bit memory bus.
[1] There are other factors of course: SD TV resolutions are low so when you hook a 32-bit 90s console up to a large modern flat panel HD or 4K TV with games running at a resolution of around 320 x 240 the pixels are MASSIVE. In addition polygon counts are low, draw distances are often low, and so it goes on. Depending on your setup games can look considerably worse on a modern TV than they would have done on more modestly sized 90s CRT screens. To be clear I'm talking about SD TVs here, not CRT monitors, which could support much higher resolutions and would therefore suffer from some of the same problems as modern flat panels in terms of making the graphics look too sharp.
(I think you're getting distracted by the terms "digital" and "analog". It may help to think about this in terms of discrete and continuous inputs instead.)
to https://bitbuilt.net/forums/index.php?threads/trimming-your-... revercse engineered pcb layout here https://gmanmodz.com/2020/01/30/2020-the-year-of-n64-again/
One was thrown together by committee with some function goal in mind, the other designed top to bottom with huge influence from process engineers. Simplified layout and reducent component count/variety means less time in pick&place, faster optical alignment, faster optical inspection, less opportunity for process flaws.