LaiNES – Cycle-accurate NES emulator in around 1000 lines of code
github.com
github.com
According to https://en.wikipedia.org/wiki/Instructions_per_second , a modern i7 processor handles north of 100,000 MIPS.
According to https://en.wikipedia.org/wiki/Transistor_count, the 6502 had 3510 transistors.
At 100,000 MIPS, a modern CPU would have a budget of ~56k instructions to process one cycle of that CPU, or about 16 instructions per transistor.
So it would seem it might now be possible to simulate those old processors, at the transistor level, in real time. Is anyone aware of experiments in this domain? If not really useful, that sounds like an interesting fun (side) project.
http://arstechnica.com/gaming/2011/08/accuracy-takes-power-o...
I'm a daily/weekly higan user, nut I make no excuses for how it's written. take a peek into the code.
And if you liked that you'll probably also enjoy http://www.megaprocessor.com/progress.html
It takes a lot of grunt to simulate a modern GPU or CPU, and of course getting near realtime is impossible, but it's handy if you want to run a load of code through it to see if it actually works.
So... yes, you can unit test the design of your processor. If you have a big enough server farm.
This comes to the host CPU power. The more accurate the emulator, the more power from the host will be needed.
UPD: Found a wiki article: http://emulation-general.wikia.com/wiki/Emulation_Accuracy
In this case, a 'cycle' is a processor cycle. Most emulators are imperfect, usually due to speed considerations, or to lack of access/understanding of the specifications of the original hardware. A short codebase that is entirely accurate is an achievement.
Cycle-accurate means the emulator is emulating all of the little steps that make up an individual instruction. Instruction-accurate emulators ignore the smaller steps and treat instructions as indivisible.
Cycle-accurate NES emulators only really matter for emulating certain graphical and sound effects.
In reality, the clocks for the various components don't necessarily run at the same rate, or even at integer multiples of the CPU clock. You can have a situation where, for example, there are 3.5 clock cycles on the graphics hardware for every one CPU clock cycle. For a lot of the classic systems, this happened because a single higher master clock is divided down for each component.
A "cycle-accurate" emulator is one that operates as if the emulated state of all hardware were updated on every tick of the master clock. This wasn't generally done in the past because it was far too slow ~20 years ago when emulation of classic consoles and computers really took off.
More sophisticated hardware doesn't necessarily have any single master clock in this sense, so it doesn't make much sense to talk about a "cycle-accurate" emulator of a modern PC, for example.
In fact, doubling (or more) the CPU clock enhanced some of the games I tried. Animations in Kirby's Adventure became more smooth (e.g. the Spark ability). Screens full of enemies in Metroid ran with no slowdown. The glitchy first scanline of status overlays in various games cleared up. I haven't yet observed negative effects in major-party titles.
I had designed my emulator as an experiment: treat the NES as an abstract specification, rather than a concrete implementation. Turns out that a lot of games seem to have actually been designed following the same principle. It was a wonderful feeling to see these games as I imagine they were intended to be experienced, as if I opened a letter from the game designers left unopened for thirty years.
Did you mean the way they were not intended to be experienced? Otherwise it doesn't make sense. The way they were intended to be experienced is on a real console connected to a CRT TV.
We can objectively talk about various metrics of accuracy compared to the original hardware. Authorial intent is, pretty much by definition, a matter of opinion (and authors themselves, over the years, often change their ideas of what their intent was).
* no framerate drops
* audio synthesized with band-limited step functions (including the "triangle" channel)
* video composed of band-limited scanlines
* video interrupt running at NTSC rate
* audio interrupt running at 240 Hz
* free-running audio synthesis
Like you said, it is entirely a matter of opinion what authorial intent is. For my purposes this definition provided an interesting experiment, and the result was aesthetically pleasing.
All told, it's all a part of game's version of the "the artwork is never quite finished/realized, it's just eventually published".
I disagree with this statement. It takes a considerable amount of effort and research to reach cycle-accuracy, especially on the PPU side. Understanding how the PPU pipeline works will take you more time than everything else.
If you settle for frame-based emulation you can write the functions to emulate the opcodes without worrying about the order and duration of the internal operations. Once the opcode is emulated, you just add cycles to a global counter based on a table that contains the number of cycles per instruction.
After 29781 cycles (you may end up being a little off each time, if you are not cycle-accurate), you call a function to update the state of the PPU. There you can use familiar iterative constructs to perform the rendering.
Compare this with the state machine approach in my code (ppu.cpp) and how much more careful I have to be composing microinstructions to form opcodes (cpu.cpp). I came up with a (I think) clean design but it wasn't trivial.
Even in a modern PC, many of the clocks are divided down using PLLs which maintain a fixed frequency and phase ratio, so it is theoretically possible to make such an accurate emulator, but it would be difficult to write and many orders of magnitude slower than the real hardware.
For an example of PC software which doesn't run correctly on anything but real hardware or really accurate cycle-emulation, see this amazing demo:
https://trixter.oldskool.org/2015/04/07/8088-mph-we-break-al...
You generally only care about cycles in real-time stuff. In simpler CPUs like those used for microwaves and such, each CPU instruction takes constant time. You can count the number of instructions in your assembler loop and know how long the loop will take. Mind you sometimes each instruction takes constant time but some can take more time than others. A cpu cycle is defined as 1/clockspeed . A fast instruction can take 1 cycle, others 3 or more.
Usually there's multiple cpu instructions for GOTO and loading from memory, a fast set for things "close by" and a slow set for far away code or data. Loops and conditional instructions even on simple cpus can take different amounts of time depending on the outcome, so it gets pretty complicated to time things sometimes.
So when doing really low level timing critical code, you can use a super accurate signal generator to drive your CPU clock, then count your instructions to time things. Common uses include generating audio from bit banging and similar.
In more complex CPUs a number of things made it difficult or impossible to figure out exactly how long an instruction will take. None of the software for these CPUs is written to depend on exact cycle timings. Because nothing depends on cycle accuracy you don't have to worry about the timings when emulating these kinds of CPU's.
Or in the graphics processor, maybe you'll blit out whole sprites/tiles at once, instead of rendering them pixel-by-pixel, the way that the hardware does.
In a cycle-accurate emulator, you're going to run every opcode, process the graphics pixel-by-pixel, etc.
The rule of thumb is that if it's short but in no way obfuscated, it's terse. If it's short, and impossible to read because of it, it's gratuitously compact.
/* CPU state */
u8 ram[0x800];
u8 A, X, Y, S;
u16 PC;
Flags P;
bool nmi, irq;
I like it.The reason for the single letters is that those are the actual names of the 6502 registers.
Glad you like it overall. :)
I also love the way you use C++ templates to simplify things!
program counter
flagPort
non maskable interrupt, interrupt request
I'm sure you knew that though, it's common micro processor nomenclature.
Additionally, the zeroth "page" of memory (the lowest 256 bytes) are accessible with a dedicated addressing mode which saves space and time in the program, and can be used as a form of cache.
I'm always kind of in awe of this sort of thing. Maybe I should try to do it. Should take some of the awe away.
But I should probably focus on sucking less, first.
Everything you need is on the NESdev wiki or it's linked there. A couple of particularly helpful resources are linked in my README. The PPU diagram and the 6502 reference were especially useful.
I definitely encourage you to do it, it's a great learning experience and very rewarding. When games start running is pure programming ecstasy :)
In my case the Sinclair Spectrum, which I did cycle-accurately on about a 25MHz 386 and fast enough to play Jetpac on the slowest 386 ever sold. Cycle accuracy really only helped with pitch-perfect sound.
It was fun, and pushing something from a blank sheet to a working program is... shall we call it a practical exercise in not sucking?
However, putting this repo through `cloc` reveals the HN title to be rather misleading.
EDIT: I had initially read through {cpu,apu,ppu,gui,joypad,mapper}.{cpp,h} and noticed that the mental tally I was taking had run well over 1000LOC. In my haste I quickly cloned the repository and ran cloc against the current dir which massively inflated the result (Doh!). See author's comment below for a more sensible figure.
Maybe calm down with the "karma farming" accusations, my problem was with the HN title, not the repository.
CPU and PPU implementations tend to be in the order of the thousands of lines -- they are around 200 and 300 lines respectively in LaiNES. In fact, most of the code is in the GUI that I could easily strip away if this was a competition. And this wasn't written to be small - it was written to be simple. It also came out small, but that's incidental.
Here's how I counted the lines and how I decided on the description for the repository, which by the way has been catapulted from totally unknown to worldwide attention overnight, and it's now object of unexpected, ruthless scrutiny that I couldn't foresee.
[andrea@manhattan src]$ rm -rf boost nes_apu Sound_Queue.*
[andrea@manhattan src]$ cloc .
24 text files.
24 unique files.
1 file ignored.
github.com/AlDanial/cloc v 1.70 T=0.03 s (780.3 files/s, 63170.2 lines/s)
-------------------------------------------------------------------------------
Language files blank comment code
-------------------------------------------------------------------------------
C++ 11 210 110 1163
C/C++ Header 12 87 7 285
-------------------------------------------------------------------------------
SUM: 23 297 117 1448
-------------------------------------------------------------------------------Still though, my gripe with the HN title still stands: to say this is ~1000 LOC is a bit rich. It's still bloody small so why try to shoehorn it into this category?
> And this wasn't written to be small - it was written to be simple. It also came out small, but that's incidental.
I think that why this is pure gold! Because it's simple, it's easy to understand, I would have been hopping with glee if this had been available to me as a teenager, instead of having to read tons of articles/textfiles of varying quality with lots of trial, error and head-scratching. It's size is besides the point, and why the title is still - in my opinion - misleading. I can't help but feel the very people that would benefit the most from this might possibly be put off that they're going to be presented with some indecipherable demo comp entry.
Either way keep up the good work.
I believe there is still a lot of room for improvement in terms of accuracy, clarity and code size. Note that this repository was more than 3 years old. Maybe this will motivate me to improve it even further.
While reading through the repo one thing I kept thinking was it would be nice IMO would be decoupling the everything from the GUI so that it was a little 'flatter'. So `main` would call `NES::run()` (or something), and `NES` would leverage GUI. GUI would be just drawing stuff. (In my head at least,) it feels like that way it might be easier to mentally partition things, for those using it as a learning project. As `NES` would be responsible for ownership and interop that GUI is doing now. I'll add it to my todo list and perhaps in god-knows-when I'll fork it and do this if you haven't gotten round to it :) Having said that it's inspired me to finish my first Go project which was a GameBoy emulator. Was started mainly to deep-dive the language, but I think it could be useful in a similar way if cleaned up and documented for folks.