CorePy: Assembly Programming in Python
corepy.org
corepy.org
Ruby is an excellent language for this, because it gets out of the way. Here's a Rasm snippet:
@epilog ||= @prog.add {
push ebx
mov ebx, retv
mov [ebx], eax
pop ebx
xor eax, eax
ret
}
Obviously, that's all pure Ruby.Once you have a class mapping for each of the instructions in your target ISA --- way easier than it sounds --- getting your code to jump into a runtime generated buffer is pretty easy.
We're starting to throw code onto Github; I've been thinking about publishing this, but didn't think it would be very interesting, except as a hack.
Significantly, this means that you can write assembly macros with the full power and elegance of lisp. A trivial example is the move macro, which is used extensively in VOPs:
(defmacro move (dst src)
"Move SRC into DST unless they are location=."
(once-only ((n-dst dst)
(n-src src))
`(unless (location= ,n-dst ,n-src)
(inst mov ,n-dst ,n-src))))
Check out the SBCL source if you want to see some wild uses of VOPs and macros that expand into assembly code. In compiler/x86/arith.lisp there is a lisp/assembly implementation of the Mersenne Twister.VOPs can optionally handle argument type checking and can be tuned to emit specialized code in various circumstances, such as when one of the arguments to the call is a compile-time constant.
So rather than doing...
def doit(option1, data): for x in data: if option1: x += 7
You can create a function that doesn't include the "if option1" part. This is great if you have many different options... and instead of bloating your code for many different optimized function, or one slower generic function, you can create the optimized functions at run time.
Looks very nice. I just wish it had support for 3dnow, and windows :)
In most cases the compiler will do a better job at optimizing machine code than humans will.
For example, numerical calculations that are highly SIMD can be improved substantially by using SSE instructions. Autovectorization is ok, but not great. Furthermore, a programmer who is familiar with SSE instructions can alter/swizzle data to make things easier to use with SIMD instructions. Yet further, a programmer can take advantage of things like non-temporal storing which compilers will not do on their own.
Now, granted, you can massage gcc into giving you "good" code with alot of hints but to do it "right" you are still peering at the assembly and making sure gcc isn't doing anything "stupid". Highly optimized C code is so dense with compiler directives as to be unrecognizable to someone unfamiliar with the underlying architecture.
The belief that your naive for-loop computation is somehow transformed automagically into perfection by the compiler is a pure and unadulterated myth.
Let's go back to the original statement:
I'd assume any intense math operations that happen inside a loop would be much faster.
That's what I was responding to. Just trasnlating the logic to asm won't make it fast. If you look at my earlier followup, I mentioned that if you could make assumptions about your code that the compiler can't know, then you're back in the land where asm optimizations can make sense. That seems to be what you're getting at with the rest of your points.
SIMD is kind of a special case. The vast majority of code I've written in my life has not been SIMD'able. Heck, the vast majority of the code I've written hasn't even been for-loops.
I'm interested to see what kind of (non-SIMD) code can be done better in hand-coded assembly than from an optimizing compiler, for real programs.
Now, if a project like this made it easier to embed C or C++, you might get the best of both worlds...
http://www.cosc.canterbury.ac.nz/greg.ewing/python/Pyrex/ver...
http://sourceforge.net/projects/cxx/
http://www.boost.org/doc/libs/1_37_0/libs/python/doc/index.h...
swig for completeness but I don't think it qualifies for easier
or just let the jit do it http://psyco.sourceforge.net/introduction.html
Ultimately though, I agree that a JIT is the way to go.
BTW, while that linked code is a good example of how weave works, it's a bad example of when to use it. I've no doubt that the code runs faster with that portion written in C, but a different algorithm could blow it away. A prime generator like that would be MUCH faster with a sieve, even in pure Python.
I'm glad you found it instructive.
Also, do you compile your python programs?
Note that I'm not so much comparing against Python as comparing to calling a C function from Python (which is easy).
The cases where you tend to want to do things in assembly are usually where there are assumptions that you can make about the execution that the compiler can't know.
Thanks for this news.