I used profilers: Chrome developer tools and Firebug. The Chrome profiler is very nice and easy to use, but it looks like it is a statistical profiler. When the usual invocation takes less than 1ms, it might get the times wrong (there might be a bottleneck that Chrome profiler is not showing).
The Firebug profiler is more precise, I believe it traces every executing instruction, which also makes it much slower to run you program in, but does sometime gives you interesting insights.
Also, there's intuition. I wanted to keep the code readable. The interpreter loop is basically something to the effect of
while(true) {
op = decode(mmu.read_32(cs.pc));
cs.pc += 4;
(optable[opcode])(op);
}
which I believe is very readable, but it counters everything one expects about writing fast javascript code that runs in a tight loop. Bellard told me he thinks my "dynamic interpreted" version might be slower than a throughly optimized interpreter version, and he might very well be true. But I found the readability of the interpreter to be of the upmost importance, at least until I had a working "dynamic recompilation" version.To optimize the interpreter version, the profiler showed me that basically three things take most (about 65, 70% of the time: * Instruction decoding * Memory access
There's not much I found to make instruction decoding faster , but the memory access time could be improved. In the beginning, I used the form mmu.read_32(addr) to read from any region of the memory, including RAM. Most of the accesses are of course from RAM, so it made sense to optimize for this fact (despite it being bad for readability).
So nowadays I do: op = (v8[rpc] << 24) | (v8[rpc + 1] << 16) | (v8[rpc + 2] << 8) | (v8[rpc + 3]);
For memory accesses.
Some other things I found randomly. The processor in the board I'm emulating has a 75Mhz clock. With that emulated frequency, Linux was spending a lot of time on the "idle task". Whan I told the emulator that the frequency was about 8Mhz, the problem went away.
Then, I thought to myself: it would be neat if I could implement parts of the libc in Javascript. That's how https://github.com/ubercomp/jslm32/blob/master/src/lm32_libc... came to be. So I implemented some string.h functions in javascript (memcpy, memmove, and memset), and voilá, the boot time went down about 3 times.
But in the end of the day, lm32_libc_dev is a hack with very negative consequences. I would have to modify every program I wanted to run on the emulated board to use my libc functions instead. I did this with the linux kernel and it was a bad experience. So I focused my energy on the dynamic code generation side of things.
I can tell you I have succeeded as the jslibc thing is now deleted on my local source three and I didn't experience any slowdowns whatsoever.