Project-Nested: An NES Emulator Running on the SNES
github.com
github.com
IMHO that's not quite right — there's a lot of simple things Nintendo could have done to make the SNES a lot more backward-compatible with the NES, that would have had to have been done right at the beginning (and so stuck around as additional evidence of this), just as the few places that do line up (e.g. controller read-out + MMIO address) were certainly done right at the beginning.
Instead, I think what Nintendo was imagining, was that — as long as the NES and SNES were still both "alive" in the market concurrently — then companies would want to concurrently release both NES and SNES versions of their games. And, with some careful planning, development studios would be able to have a single 6502 macro-assembler codebase that compiled to either a NES or SNES target. Developers could either add MMIO address pokes / JSRs inside #ifdef-like macro structures; or just write two subroutines, one for NES and one for SNES, and then compile-time branch to determine which would get compiled in.
This makes sense of where the NES and SNES line up in compatibility, and where they diverge: they line up where there's an obvious way to support both consoles with one block of instructions; they diverge where developers would have to write separate subroutines to support the different hardware anyway, so instruction-level / memory-map-level compatibility isn't so crucial.
This seems so obviously "possible" (though definitely challenging!) that I've always wondered whether there are any games that are actual examples of this — where the game came out at the same time in both NES/FC and SNES/SFC releases, and binary analysis reveals that the majority of the game's engine code is shared between the two, with differences mostly in the subroutines/coroutines composing the "graphics engine", and in the static data.
—————
On a separate note, I feel like it would be even simpler — and a lot more performant! — to ahead-of-time recompile a NES game to run on the SNES. The NES doesn't tend to do tricks with dynamic runtime code generation, so NES games are good candidates for being statically analyzed and transpiled. (And the SNES has such a similar architecture, that even those dynamic runtime tricks might map cleanly onto the SNES, too.)
You might be interested in these articles covering each console's architecture:
* NES https://www.copetti.org/writings/consoles/nes/
* SNES https://www.copetti.org/writings/consoles/super-nintendo/
* Game Boy https://www.copetti.org/writings/consoles/game-boy/
I'm far from an expert player, but I really think the physics are identical once the patch is in place. I sometimes wonder if it was really a mistake at the factory—something they didn't catch until 20,000 PCB's in, and didn't want to change in revisions for consistency.
> The actual mechanics behind the issue were discovered a while ago (the Y-reverse thing isn't right). Something to do with the block getting replaced immediately instead of a one-frame queue/delay.
Unfortunately I can't find a fuller explanation for this claim.
Having used the patch, it does feel right...
Maybe brick-blocks taking an extra frame to swap from tiles to sprites for their animations, was something perceived as an unfixed bug (or a necessary compromise, due to per-frame CPU-time limits) at the time of the original's development — one with, perhaps, a "TODO" comment sitting there in the original assembly source. And the development team of the port noticed that, with either more dev time, or with the increased cycles-per-frame available on the SNES, the "bug" was able to be "fixed".
But perhaps — since that team wasn't aware that the change in physics this created was unexpected by the original devs — they thought that the physics resulting from "fixing" this "bug" were the originally-intended physics of the game (i.e. that the animation "bug" was a regression, and fixing it put the game back the way it was in some earlier build), rather than that they were introducing a regression themselves.
These sorts of things are the reason I've always wished version control got used more heavily in the gamedev industry back in the 80s. Then any leak of a game's code-base would actually have a productive use (rather than just serving as poison fruit to aspiring developers): it could be a treasure-trove of data for Software Engineering academics to study the in-industry SwEng practices that lead to "good" games vs "bad" ones.
(I've had a longstanding hypothesis that many "bad" games from the 80s/90s were "good" games for much of their development, but got broken very close to release as they tried to merge in tons of longstanding feature-branches — i.e. the exact problem Continuous Integration tries to address.)
Personally, I disagree with the author of that article's conclusion here, that the [only] "solution is to embed an interpreter runtime in the generated binary." You can statically recompile this code!
You "just" need to run a concolic interpreter over the code, generating out new static instruction streams for each possible phi-node (branch) state, and then coalescing them back together where you find you've generated the same instruction stream. You can early-terminate these spec-exec paths when you find yourself back at an instruction postdominated by entirely static code [within the body of the unrolled runtime coroutine sequence].
Might sound like an unbounded exponential check, but in practice 1. merging back together will usually happen within, at most, five instructions, because these are almost always last-minute bit-packing-hacks to get code to fit within a given constrained ROM image size; and 2. there are always only two concolically-discoverable valid interpretations of a given instruction stream (if doing mid-instruction jumps), and only one concolically-discoverable valid interpretation of data-as-code / RAM-as-code. (NES games don't assemble and JIT arbitrary code streams. They don't have the work RAM for it. They're at most unpacking existing, non-arbitrary code streams; or taking something that is primarily a code stream, and then reusing it as data elsewhere, maybe a PRNG seed, or a 1-bit noise texture.)
A large part of my job lately is decompiling and "reconstructing meaning" for object code compiled for lousy underspecified abstract-machine architectures. There are a lot of modern compiler-theoretic tools and tricks that can make short work of "statically" recompiling any but the hardest of hard cases. They're sadly virtually unknown outside of academia, though (probably because they're based on various graphical transformations of the code proposed after the Dragon Book was published, which is the point in history where most people in industry's knowledge of compilers seems to stop.)
Pretty sure you’re still going to be stymied by things like cycle-counted code to line up I/O writes with the state of the hardware, though. I don’t suppose there’s tricks for that.
Most of the point of recompiling (JIT or AOT) is to amortize instruction decode and dispatch costs versus an interpreter. If you have to do the cycle accurate bookkeeping anyway, you're most of the way to the overhead of an interpreter again. There'll be some savings around partial instruction streams that don't have effects outside the CPU core that can just be run as a burst, but I'm not sure it'll be a game changing amount.
That’s correct — but one of those events is external interrupts (from the graphics, audio, or cartridge hardware), which complicates things a bit since your “checkpoints” can be anywhere in the code rather than just at a specific set of register accesses.
You could work around this by tracking cycles per-basic-block while running transpiled code, and falling back to cycle-accurate software emulation when you get close to a hardware interrupt.
Instructions are only written once but could potentially be written by any store with zero page indirect[1] or indexed addressing.
[1] that is ($AA,X) or ($AA),Y
That said, the stack was used as a fast indirect jump address location pretty frequently, as described here, essentially abusing the fact that PHS and RTS are 1 byte instructions: http://www.6502.org/tutorials/6502opcodes.html#RTS
And really since basically all PC-modifying instructions operate on either relative (b*), absolute-indirect (jmp), absolute (jmp, jsr), or stack-indirect (rts, rti) there's no special benefit you get from having code in the zero page anyways afaik.
You save one cycle off each instruction that modifies another instruction, which can important for inner loops. See eg http://www.linusakesson.net/programming/gcr-decoding/index.p..., although that's a disk-induced realtime requirement, rather than hblank/beam-induced.
Thanks for clarifying :)
Runtime code generation is rare, but dynamic dispatch (i.e. "JMP indirect" through RAM), jump tables, and certain instruction-level hacks are fairly common. This can make it tricky to reliably identify all subroutines or traces [1].
Edit: Thinking about it more, that all would also make sense given their strategy past the SNES for back compat. On the GBA they literally just have a complete GameBoy SoC on the die. It's not accessible to GBA software, is probably clock gated in GBA mode, but is there with as few changes as possible from GB hardware to run GB software. Then for the Wii, the companion ARM processor (Starlet as it's known by the Wii hackers) will go so far as to patch problematic GameCube games on load even though it's very nearly identical hardware with a few extensions. It seems like they either drop down gate for gate compat, leave a back door for themselves to get really dirty with patching, or these days just keep an emulator around that they've re-QAed each game on that they'll allow. They're very intentional with back compat in a way that feels like an ancient learned lesson internalized by the company.
I mean, we'll never know for sure, but what was the cpu landscape like? Were the other options clearly better?
Other 16bit consoles used an off-the-shelf 68000, combined with custom support chips. Nintendo made their own custom cpu silicon, an off-the-shelf soft-cpu and custom support logic for bus control and DMAs.
And there simply weren't that many 16bit soft-cpus available for Nintendo to license. You couldn't license the 68000 HDL, nor the 8086 HDL.
It was a very different era back then. Tooling sucked and everything was a lot more manual. But we didn't mind as were still pushing boundaries of our imaginations and didn't know any better.
And then I think you'd have seen a mapper for the NES that more closely resembled the 65816's native addressing, since that would definitely have eased such porting efforts. But even the MMC5, which was first used the same year the SNES came out didn't move much in that direction.
Tbh I think you're also overestimating the assemblers they used at the time as well, both in terms of the expressiveness of the macros they had available and and kind of dead code removal you might be imagining happening.
I think a much simpler explanation is just that they had it in mind as a possibility from the start and they made some design choices that were both easy and would have allowed it. They may have been leaving open the possibility of an adapter cart like the genesis power base converter (which had much less work to do) too. But they also decided fairly early on that it wasn't really worth it economically and nixed it early enough that it didn't make a huge impact on the architecture.
Nintendo stuck with pretty similar PowerPC chips for three generations while Sony went MIPS->PowerPC->x86. And when they finally switched off of PowerPC, they went to ARM, which they'd been using for handheld systems for two decades.
[1]: https://en.wikipedia.org/wiki/Gunpei_Yokoi#Lateral_Thinking_...
Ive heard something maybe possible with retroarch but Im surprised there is no pi zero with an easy multiplayer setup yet ( for instance to play 2 player Mario Tennis on the Gameboy )
You can add a small delay to the host to even things out but to properly account for latency is a major reverse engineering project which afaik off the top of my head has only been done for super smash bros melee
In addition, it's very easy to have your rom inside the retroarch folder, then send the whole folder to another person / computer. That way there is no confusion about rom versions, in addition you can pre-configure the controls if the person you're sending the folder to isn't as tech savy.
No need for this, it's natively supported and runs game cube games (and with the right setup, GC homebrew)
Wii -> GameCube (similar hardware) -> [not N64] SNES -> NES (demonstrated by this project)
example. Heck, if console manufacturers kept two separate branches "every other generation" emulation (odd/even generations) and then made sure each console generation had extra hardware to play the immediately prior generation's games (as some sometimes do), then every generation could play every other generation. That would be nice for players, but maybe not nice for the companies though.
https://external-content.duckduckgo.com/iu/?u=https%3A%2F%2F...
Epeople put things on GitHub because it's free and useful
Step 1, click "Open Nes" and select a game.
Step 2, (optional) select or create a profile.
Step 3, click "Save Snes", a file will be created in the same folder as your Nes game.
Step 4, play on Snes hardware or Snes emulator.
Step 5, (Optional) click "Load SRM" and select the SRM file(s) generated by the Nes emulator,
feedback will be saved to the profile so the game can run faster after repeating steps 1-4.
[1]:https://github.com/Myself086/Project-Nested/blob/fc108119/Pr...