300% faster float point ops on ARM Linux
engadget.com
engadget.com
This both reduces register pressure and means you don't have to marshal stuff to and from the integer registers at function boundaries.
The reason this hasn't happened before is because VFP is an optional part of the instruction set, and you can still configure compliant ARM cores without a FPU.
I'm not sure if their port uses NEON (the ARM SSE equivalent) at all.
OpenEmbedded/Angstrom would be ideal targets.
Props to this guy (and his team-mates) for taking on such a project and seeing it through. I'm sure there was a lot of detractors saying things like 'It won't matter enough" and "Recompiling everything is going fragment the ARM repositories and it's not worth it."
Btw, I'm very satisfied with the genesi products - OpenGL ES 2.0 works just fine, with some little bugs, but overall very cool and I think cheap system (Ubuntu 10.10, and they are working toward Ubuntu 11.04).
LuaJIT works also there (and jits), giving 4-5 speedup over standard lua (It would be better, once LuaJIT too starts using hard floats, it's still relying on software math libs)