S6: A standalone JIT compiler library for CPython
github.com
github.com
- Pythran (https://pythran.readthedocs.io) (ahead of time compiler) - mypyc (https://mypyc.readthedocs.io/en/latest/) (ahead of time compiler) - cython (ahead of time compiler) - Numba (JIT) - Pypy (An interpreter doing JIT compiling) - Nuitka (Ahead of time, IIRC)
At this point we should build a "Awesome Python Compilers" repo ... oh wait, it obviously already exists: https://github.com/pfalcon/awesome-python-compilers
e.g. PyPy doesn't work with PyTorch https://github.com/pytorch/pytorch/issues/17835
> We have stopped working on S6 internally. As such, this repository has been archived and we are not accepting pull requests or issues. We open-sourced the code and provided a design overview below to spur conversations within the Python community and inspire future work on improving Python.
> Python is slow, and a lot of researchers use Python as their primary interface to build models.
Why not just build models in JS?
I'm hopeful that honest efforts to port over useful bits of the scientific/mathematical ecosystem to JS will start to take off someday soon.
And have a look at, say, Julia.
It's more likely that something like Julia ends up replacing Python, even though I still believe it's a low probability thing.
Pytorch readme quote
pytorch is not a Python binding into a monolithic C++ framework. It is built to be deeply integrated into Python.
could be applied to many key libraries of this ecosystem.For sure, a lot of those libraries rely on tons of compiled code in C/C++/Fortran, but the python code is not mere wrapping on top of those libs. And most of the compiled code inside those repos has a lot of internal logic as well.
Some examples (using cloc on git main branch, as tokei does not understand cython):
* sklearn: 150 kloc python, ~10kloc cython, not much of other languages
* NumPy: 200kloc python, 350kloc for C/C++
* scipy: 800 kloc C, 200 kloc python (the 800kloc surprised me and it looks like a lot of it is C generated code from cython)Because ML/Data source has particular numerical requirements. JavaScript has poor mathematical defaults and still misses useful datatypes like float16*.
* Yes, this can emulated. Emulation is slow. You've gained some speed and then paid it back for emulation.
IMHO, you can't beat Python for ML by being just a slightly faster Python (there aren't enough newcomers for that, the Python ecosystem is enough to assimilate them before other languages can). You need to offer something more significant to make migration worthwhile. Julia tries to do that by allowing to implement the main stack in Julia, if it had JS speed and DX it would have eaten Python's lunch.
JavaScript is 43 times faster for this random workload:
(env) bwasti@bwasti-mbp ~ % time python bleh.py
python bleh.py 9.27s user 0.06s system 99% cpu 9.352 total
(env) bwasti@bwasti-mbp ~ % time bun bleh.ts
bun bleh.ts 0.16s user 0.02s system 83% cpu 0.216 total
(env) bwasti@bwasti-mbp ~ % cat bleh.py
import random
v = [0 for i in range(1000000)]
for i in range(int(1e8)):
v[i % len(v)] += 1
(env) bwasti@bwasti-mbp ~ % cat bleh.ts
const v = new Float32Array(1000000)
for (let i = 0; i < 1e8; ++i) {
v[i % v.length] += 1
}
Try as many as you want and you'll see that it just blows python out of the water