Taichi lang: High-performance parallel programming in Python
taichi-lang.org
taichi-lang.org
To make it work with taichi, I had to change the declaration of the sieve from sieve = [0] * N to sieve = ti.field(ti.i8, shape=N) but the rest of the code remained the same.
Ordinary Python:
time elapsed: 0.444s
Taichi ignoring compile time, I believe:
time elapsed: 0.119s
A slightly more realistic example than the 10x+ improvement they show on the really toy example with results that aren't too bad. I'd take a 3x improvement for tiny changes. Pretty neat!
(I tried some other trivial things like using np.int8 and it was slower. One can obviously make this a ton faster but I was interested in seeing how the toy was if we just made it slightly more memory-bound).
A negative was that throwing list comprehensions in made the python version faster - about 0.3 seconds - (and shorter and arguably more "pythonic") and simultaneously broke the port to taichi.
Could you share your code as well.
N = 1000000
isnotprime = [0] * N
def count_primes(n: int) -> int:
count = 0
for k in range(2, n):
if isnotprime[k] == 0:
count += 1
for l in range(2, n // k):
isnotprime[l * k] = 1
return count
import taichi as ti
ti.init(arch=ti.cpu)
isnotprime = ti.field(ti.i8, shape=(N, ))
@ti.kernel
def count_primes(n: ti.i32) -> int:
count = 0
for k in range(2, n):
if isnotprime[k] == 0:
count += 1
for l in range(2, n // k):
isnotprime[l * k] = 1
return count import time
import math
import numpy as np
N = 1000000
SN = math.floor(math.sqrt(N))
sieve = [False] * N
def init_sieve():
for i in range(2, SN):
if not sieve[i]:
k = i*2
while k < N:
sieve[k] = True
k += i
def count_primes(n: int) -> int:
return (N-2) - sum(sieve)
start = time.perf_counter()
init_sieve()
print(f"Number of primes: {count_primes(N)}")
print(f"time elapsed: {time.perf_counter() - start}/s")
Taichi: import taichi as ti
import time
import math
ti.init(arch=ti.cpu)
N = 1000000
SN = math.floor(math.sqrt(N))
sieve = ti.field(ti.i8, shape=N)
@ti.kernel
def init_sieve():
for i in range(2, SN):
if sieve[i] == 0:
k = i*2
while k < N:
sieve[k] = 1
k += i
@ti.kernel
def count_primes(n: int) -> int:
count = 0
for i in range(2, N):
if (sieve[i] == 0):
count += 1
return count
start = time.perf_counter()
init_sieve()
print(f"Number of primes: {count_primes(N)}")
print(f"time elapsed: {time.perf_counter() - start}/s")
(The difference of using 0 vs False is tiny; I had just been poking at the python code to think about how I'd make it more pythonic and see if that made it worse to do taichi)It looks super interesting except “almost the same syntax as Python” part here seems like such a foot gun for everything from IDE integration to subtle bugs and more.
I was super into the idea of a strict Python subset that gets JIT compiled inline based on just a decorator.
For this we had to calculate forces to animate some kind of polygon with a lot of joints and we could not just call sycipy from the taichi code. I had to implement a very dirty polynomial equation solver in taichi for the demo
Meaning?
- their approach is still bizarre and exploratory and they still don't know how to structure their APIs and are making it up as they go?
or:
- there are still some rough edges, bugs, and no full documentation yet?
as those are quite different cases...
CUDA-only, no mention of Metal.
[0] https://developer.nvidia.com/nvidia-triton-inference-server
they didn't get range composition right
for(int k = 0; k != N; ++k) { dowork(k); }, similar in Python
ranges are [0...N). If I want just a sum, and would like to split work among multiple CPUs, I could write
s = sum(0...N) or s = sum(0...N/2) + sum(N/2...N) or s = sum(0...N/3) + sum(N/3...5N/7) + sum(5N/7...N)
and it will work beautifully. Ranges are decomposable and composable with any number(s) splitting original range.
[0...N/2) + [N/2...N) = [0...N)
When I first looked at Julia, I was ... unhappy. They didn't get simple ranges composability right. You could do it, sure, but it looks so ... unnnatural.
This is somewhat achievable with separate languages. But it ultimately takes a very similar shape with additional tooling/ceremony and diminished benefits particularly from colocation.
but there's a better way, there's a better way.
like python used to be less popular than perl but look at where perl is now vs python. things do take time to change though
Want to use NumPy? Better hope you have code that can be expressed as array computations (a tiny minority of performance sensitive tasks in my field of work).
Cython? Now you're just writing another language, and get the joy of having to distribute compiled code in your Python package - for example, like the very fun segfaults I experienced because Conda will automatically override the system linker, breaking all compilation including my Cython modules.
Numba? Hope you don't make custom classes - you know, one of the basic features of the language.
If you look at ML, Python is completely fine because all the processing that happens with matrix multiplication, even on CPUs, far, far, FAR outweighs all the setup stuff in volume of operations.
On the other hand, if majority of your application relies heavily on processing speed (i.e you need compare/jump operations rather just add/multiply/load/store of the GPUs), Python is going to be slow. In this case, if you want custom performant code, you write C extensions for the performant critical code, and launch them from higher level python code.
That being said, there is generally a library (like Taichi) that already does this for you.
>> Reference Implementation >> None as yet.
> but they're still slightly janky
maybe you're from the future?
i bet the python macros will be ugly af too.
also walrus operator = bdfl no more.
Julia is a great niche language for what it is good at. It will survive but won't gain much popularity.
[1] https://discourse.julialang.org/t/very-slow-time-to-first-pl...
Packages need to precompile, and they don't. They need to fix invalidations, and they don't. They need to fix inference issues in their packages and they don't. "Using" time remains quite high.
These are all real limitations, but users of these languages learn to live with them. Rustacenans learn to start a build, then so something else while it builds. That is, on its face, totally unacceptable to a Pythonista. Pythonistas learn to always ship performance sensitive applications with _another language_ doing all the hard work: Totally unacceptable for a Rustacean.
If someone tries out Python, spends five minutes getting package management to work and fails, they have not seriously thought about Python as a language. I feel people do just that with Julia: Try it out, then reject it on the first rough edge.
Fair enough: There are many people for whom lack of static type checking or Julia's latency is a showstopper, making the language unsuitable. But I'm still firmly convinced that for scientists/engineers at least, Julia, on balance, offers a better language than Python. I hope you're wrong that it won't gain much popularity.
when we looked at it, it was because c++ is hard af to learn and the developer dont use it to its full potential. they just use some quantlib which where you look under the hood has many unoptimized parts.
with julia, the code is so simple and clean we even put in some GPU code in one place using CUDA and complete blows the C++ out of the water.
I did achieve some good performace with numba once thought with the avx so pythoni snot all bad but numba is only a small subset of python but with julia i can do crazily fast stuff that looks like python and is readable.
I don't doubt that what you say it's true, but to me, it comes down more to lack of familiarity with other languages than any actual merits of Python syntax and semantics. Frankly, I am glad the didn't take more from Python.
Still, most things worked in Julia, and there have been many improvements since then so I suspect the few remaining rough spots are being smoothed out. In the future I will be happy if I get to work with Julia more.
But then there's not enough users... so the cycle continues until one day Julia hits critical mass and a tipping point is reached.
all serious numerical languages have that. it's more natural
0 base indexing is only good for calculating memory offsets. Nothing else. like in go `vec[a,b]` is indexing `a` to `b-1` which is purely because it's more convenient due to 0-indexing. this `b-1` is hugely confusing and big gotcha for the layman.
And almost all serious general-purpose languages use 0-based. It is more natural at least to me. You see, this is exactly why 1-based index of Julia is disliked by many.
you may not like rachel macadams, but she was never gonna be urs anyway. so meh
Style aside, you can't simultaneously complain that Julia doesn't get the resources of a general-audience language and then talk down to "general-purposes folks". Julia could be the bees knees for math-y stuff. And if so, good for it. But Python's success is because its flexibility makes it pretty good for a wide range of things, even if it'll never be great for any one of them.
i love python. in fact, in my daily run on cloud run it's one of the key component since it has a decent gcp client library for writing query code to bigquery.
but then i glue it up with bash and run some other analysis in julia.
i think you can have python, julia, bash, r, whatever. heck, i even dabled in go and react.
just saying julia is aimed at numerical first.
and i think 0-based indexing was merely a historical accident anyway due to lack of memory in early computers so use every last bit including the 0 even though it's not suitable for indexing but is suited for calculating mem offsets.
whatever, use whatever u like, not trying to promote julia to you or anyone else. just having a good argument on friday afternoon
If you look at compute in general, it can be pretty much be summed up as add/multiply/load/store/compare/jump (straight from Jim Keller). All the other instructions are more specialized versions of that, with some having dedicated hardware in CPUS.
If you need to do those 6 things as fast as possible, on single piece of hardware, you are most likely writing a video game. Thus video game development is pretty much C/C++ with a few things of Swift/C# sprinkled about.
If a single piece of hardware requirement goes away (i.e you are writing a distributed system to serve a web app), people quickly figured out that hardware is cheaper than developer salary, and also network latency is going to be the dominant thing for speed. This is the reason Python took off - its super quick to write and deploy applications, and instead of paying a developer $10k+ a month, you can just spend half that on more EC2s that handle the load just fine, even if the end user has to wait 1.5 seconds instead of 1.1 for a result.
If you don't need compare/jump, your program is essentially better off suited to running on the GPUs. OpenCL/CUDA came about because people realized that a lot of applications simply need to do math without any decisions along the way, and GPUS are much better at this. The paradigm is that you write kernels that you then load onto the GPU - this can be done in any language since you really just need to run the code once. I.e Python, despite being slow is used primarily for ML because of this.
Then there is multiply/add only, which you probably best know in implementation as ASICs for bitcoin mining that blew GPUs out of the water. When you don't have memory controllers and just load/store from predefined locations, your speed goes through the roof. This is the future of ML chips as well, where your compiler looks a lot like the verilog/hdl compilers for FPGAs.
Furthermore, with ML, the compare/jump and even load/store is being rolled into multiply/add. You have seemingly complex algorithms like GPT that make decisions, but without any branching. Technically speaking, a NAND gate is all you need to make a general purpose CPU and you need 2 neurons to simulate a NAND gate. So you can build an entire general purpose CPU from multiply/add.
So in the end, its absolutely worth investing in Python and making it better. Languages like Julia are currently better suited to performant tasks, but the necessity of writing performant code to run on CPUs is going away slowly with every day. Its better to have a high purpose language that allows you to put ideas into code as quickly as possible, and then have different specializations for more generic tasks.
I don't care much about how the interface over LLVM looks like. As soon as I have the same result in the end, I'd rather stick to whatever has more users.
Julia language features are really weak.
(taichi) [X@X taichi]$ python primes.py
[Taichi] version 1.4.1, llvm 15.0.4, commit e67c674e, linux, python 3.9.14
[Taichi] Starting on arch=x64
Number of primes: 664579
time elapsed: 93.54279175889678/s
Number of primes: 664579
time elapsed: 0.5988388371188194/s
fwiw, i tried the gpu target (cuda) and it was faster than vanilla, but slower than accelerated cpu target by about 4x.
(I don't know enough about the python ecosystem, but I have to tweak code from one of my coworkers and he uses numba.)