Parallel Roguelike Lev-Gen Benchmarks: Rust, Go, D, Scala and Nimrod
togototo.wordpress.com
togototo.wordpress.com
> I think there may be a more concise way to parallelise parts of the problem in Rust, using something like (from the Rust docs): > > let result = ports.iter().fold(0, |accum, port| accum + port.recv() );
We plan to have convenient fork/join style parallelism constructs, so that you don't have to build it yourself using message passing or unsafe code. There is a prototype in the `par.rs` module in the standard library, although it's pretty rough at this point.
I'd be interested in seeing what the benchmark looks like with Rust 0.8, which has a totally rewritten (and at this point much faster) scheduling and channel implementation.
4.8 has been out for quite some time, I did a quick comparison using your benchmark between gcc 4.8.1 and clang 3.3 on my system and clang still won but the difference was ~3.5% on my machine.
But that's preferences for you, everyone has them :)
Any chance you would consider putting the compiler version used (for all compilers, not just c/c++) next to the compiler in the column for better disclosure?
Good luck on your game!
https://github.com/logicchains/Levgen-Parallel-Benchmarks/pu...
Go performance improves from ~470 ms to ~360 ms on my Quad core MBP 15 (Late 2011)
Edit: Also, how is this relevant for anything other than Scala specifically? It seems odd for C and Go to end up on one side and C++, Rust, D and Nimrod to be on the other side, if memory consumption is a big concern.
For many applications this doesn't matter, especially with how cheap RAM is. But for some applications, such as mobile games and rented server space where your profitability is a function involving RAM cost per user, this can be very significant.
That said, you can lower your memory requirements dramatically using pretty standard high performance JVM techniques in Scala much as you can in Java. I've designed server systems in scala that produce no garbage for weeks at a time.
This can be a good option if you have some small subset of your system that needs to be memory sensitive/high performance without having to lose the benefits of the JVM in the rest of the system.
- Modula-2
- Modula-3
- Ada
- Turbo/Apple/Object Pascal
- Delphi
- Oberon, Active Oberon, Oberon-2, Component Pascal
$ command time -f 'max resident:\t%M KiB' ls /
bin etc initrd.img.old lib32 media proc sbin tmp vmlinuz
boot home iso lib64 mnt root srv usr vmlinuz.old
dev initrd.img lib lost+found opt run sys var
max resident: 968 KiBhttp://tech.brightbox.com/posts/2012-11-28-measuring-shared-...
If you're building a piece of software that needs to co-exist with other memory intensive processes, the JVM's policy of taking memory and rarely giving it back can cause unnecessary swapping. Heaven forbid you need to run multiple JVMs on the same machine (like with hadoop).
Then you start having to do silly things like specifying how much memory your app is allowed to use which makes me feel like I'm on classic MacOS.
I used this approach in the past to keep memory usage lean. I had some spikes in application that required a lot of memory, so had to keep Xmx high, but configured the app to trigger deallocation to the OS when there's more that 20% of free memory. Worked like a charm even on and old ere 1.5 Sun JVM.
[1] http://www-cs.canisius.edu/~hertzm/gcmalloc-oopsla-2005.pdf [2] http://sealedabstract.com/rants/why-mobile-web-apps-are-slow...
Also there is a small bug in the C version. When it goes to print out the level, it only compares the first 100, so it'll print a different level from the C++ version with a seed of say, 20.
Also appears to be a small benefit from modifying the Lev struct to have Room first.
uint32_t genrand(uint32_t *seed)
{
uint32_t x = *seed;
x ^= x << 13;
x ^= x >> 17;
x ^= x << 5;
*seed = x;
return x;
} where
noFit = genRooms (n-1) (restInts) rsDone
tr = Room {rPos=(x,y), rw= w, rh= h}
x = rem (U.unsafeHead randInts) levDim
y = rem (U.unsafeIndex randInts 1) levDim
restInts = U.unsafeDrop 4 randInts
w = rem (U.unsafeIndex randInts 2) maxWid + minWid
h = rem (U.unsafeIndex randInts 3) maxWid + minWid
And change:let rands = U.unfoldrN 10000000 (Just . next) gen
to:
let rands = U.unfoldrN 20000000 (Just . next) gen
The running time should double. Does it? Or does it increase by orders of magnitude? The latter is what happens to me.
https://github.com/logicchains/Levgen-Parallel-Benchmarks/pu...
> 120 ms improvement over the origin implementation
https://groups.google.com/d/msg/golang-nuts/CZVymHx3LNM/esYk...
(Dmitri is one of the main people currently responsible for the further development of the Go scheduler). If you only have 8 goroutines running at a time with 8 cores, and one stalls, that wastes CPU usage. If you have, say, 80, the scheduler will put another one to work so the CPU is idle less often.