SIMD in Pure Python
da.vidbuchanan.co.uk
da.vidbuchanan.co.uk
There is also a very different “write SIMD assembly in Python” approach available through the PeachPy library, one of the least known gems between Python and HPC worlds: https://github.com/Maratyszcza/PeachPy
This is what a dot-product would look like in PeachPy: https://unum-cloud.github.io/usearch/python/index.html#id4
PS: Cppyy and Numba are also fun to use in such projects :)
Is it just me (I'm far from an expert here) or is this code really weird? Why does ymm_one appear to contain the number zero? Why do we subtract what looks like it should be the inner product we want from ymm_one at the end?
# Negate the values, to go from "similarity" to "distance"
VSUBPS(ymm_c, ymm_one, ymm_c)
Is actually just subtracting zero from ymm_c for no reason I understand.It is natural to assume vectors have distance 0 from themselves at least, which this thing doesn't (in general). E.g. if you compute this "distance" for a = b = (2,0) then you get that the distance from a to itself is -3 which seems pretty weird.
>* the state of the cells is being stored in a big array, accessed via the get_cell and set_cell helper functions. What if instead of using an array, we stored the whole state in one very long integer, and used SWAB arithmetic to process the whole thing at once?*
I’m really curious as to how the unpacking of this long integer into pixels on the screen doesn’t add more overhead than it saves. I guess I’ll have to wait for your next one on the compressed gzip stream hack.
Of course, this won’t be anywhere near as efficient as just implementing the decoder in C, as Huffman decoding needs to be much more general than just unpacking fixed-width chunks. But it will definitely be an improvement over naive loops.
I wonder whether .hex() could be pressed into service as a very scuffed 4-bit unpacker? Maybe something like .hex().encode().translate() to get an arbitrary palette mapping?
.hex() is a good idea, and probably has a decent chance of beating gzip, assuming .translate() is fast enough.
The fastest approach for the bitsliced AES impl ended up being pretty cursed: https://github.com/DavidBuchanan314/python-bitsliced-aes/blo...
> It is quite easy to add new built-in modules to Python, if you know how to program in C. Such extension modules can do two things that can’t be done directly in Python: they can implement new built-in object types, and they can call C library functions and system calls.
Then why are you talking about this instead of extension modules?
I just recently started going through a performance programming course (https://computerenhance.com/) and have learned about SIMD and other techniques and it is awesome to see something out in the wild.
> The general term for this concept is SWAR, which stands for SIMD Within A Register. But here, rather than using a machine register, we're using an arbitrarily long Python integer. I'm calling this variant SWAB: SIMD Within A Bigint.
Thanks to Peano and Godel, it's safe to say we may encode any compute with operations on natural numbers. So if anything is slow in Python for you, you may always encode it in Bigint and hope for the best.
then you can apply this bit parallelism to a tensor.