HNHacker News
TopNewBestAskShowJobs

certik

519 karma · joined March 22, 2019

https://ondrejcertik.com/
submissionscomments
certik··on LPython: Novel, Fast, Retargetable Python Compiler
Definitely let us know any feedback once you try it. You can open up an issue at lpython github.
certik··on LPython: Novel, Fast, Retargetable Python Compiler
Then it's an optimizing compiler for sure.
certik··on LPython: Novel, Fast, Retargetable Python Compiler
I would say LPython, Codon, Mojo and Taichi are structured similarly as compilers written in C++, see the links at the bottom of https://lpython.org/.

Internally they each parse the syntax to AST, then have some kind of an intermediate representation (IR), do some optimizations and generate code. The differences are in the details of the IR and how the compiler is internally structured.

Regarding the type inference, this is for a blog post on its own. See this issue for now: https://github.com/lcompilers/lpython/issues/2168, roughly speaking, there is implicit typing (inference), implicit declarations and implicit casting. Rust disallows implicit declarations and casting, but allows implicit typing. As shown in that issue they only meant to do single line implicit typing, but (by a mistake?) allowed multi-statements implicit typing (action at a distance). LPython currently does not allow any implicit typing (type inference). As documented at the issue, the main problem with implicit typing is that there is no good syntax in CPython that would allow explicit type declaration but implicit typing. Typically you get both implicit declaration and implicit typing, say in `x = 5`, this both declares `x` as a new variable as well as types it as integer. C++ and Rust does not allow implicit declarations (you have to use `auto` or `let` keywords) and I think we should not do either. We could do something like `x: var = 5`, but at that point you might as well just do `x: i32 = 5`, use the actual type instead of `var`.

certik··on LPython: Novel, Fast, Retargetable Python Compiler
> Looks interesting! Thank you for focusing on AoT compilation and not just JIT compilation. To be honest, I'm sick of JIT compilation. In theory it seems like the best of both worlds, but in practice it turns out to be the worst of both worlds, especially for larger projects.

Indeed, I think for large projects you want:

* generate a binary

* fast compilation in Debug mode

* fast runtime in Release mode

Sometimes you want JIT, so we support it too, but I personally don't use it, I write the main program in LPython and just compile it to a binary.

> Will LPython have the ability to generate AoT compiled libraries in addition to executables?

Yes, you can do it today. Just compile to `.o` file and create a library. We use this library feature in production at my company today. If you run into issues, just ask us, we'll help.

certik··on LPython: Novel, Fast, Retargetable Python Compiler
You are right, we should have made that clearer. If you look at all the 25 compilers that we list at the bottom of: https://lpython.org/, some of them are supersets, some of them are subsets. Sometimes the distinction is blurry, since even with LPython you can call arbitrary CPython today, it just doesn't get compiled right now, but perhaps later we can actually compile it like Cython does, just not speed it up much. In that case we become a subset that is fast, the rest is slower, but I think every superset of Python behaves like that too. In a way, I think all of the 25 compilers supports all of Python one way or the other (LPython as well), and I think each of them only gets a subset to be fast.

If you have some ideas how to best communicate this, let me know.

The best approach that I know right now is to say that we support a strict subset of Python, and if it compiles, it will be fast. The rest of Python you have to call explicitly and it will be slow and you get a CPython dependency in the binary that we generate, but you can do it. That way there is a clear distinction what gets compiled via our compiler into high performance machine code and what gets dispatched via CPython (it will eventually get "compiled", but it will just call CPython).

certik··on LPython: Novel, Fast, Retargetable Python Compiler
Yes, try `conda install -c conda-forge lpython`. It should work on Windows, but it's not as extensively tested as macOS and Linux. If it doesn't work, please report it, and we'll fix it. Once we get to beta, we will support Windows very well.
certik··on LPython: Novel, Fast, Retargetable Python Compiler
It was the easiest for us to deliver a binary that works on Linux, macOS and Windows. Others can then use this binary as a reference to package LPython into other distributions. You can also install LPython from source, but it's harder than just using the binary that we built correctly (with all optimizations on, etc.).
certik··on LPython: Novel, Fast, Retargetable Python Compiler
Mojo is a strict superset of Python, LPython is a strict subset of Python.

We could target MLIR later, right now we are just targeting LLVM.

certik··on LPython: Novel, Fast, Retargetable Python Compiler
Yes, the Numba use case is a subset of LPython. We want to support what Numba does, that is, you decorate your function and JIT it. But in addition, we also want to compile to binaries (ahead of time) that have no CPython dependency, and support high performance optimizations. Numba speeds up Python a lot, but doesn't seem to quite get the Fortran/C++ level of performance sometimes. One of our main goals is to be able to get maximum performance, so that eventually as a user you can depend on LPython that if it compiles it, it will run at least as fast as C++ or Fortran would.
certik··on LPython: Novel, Fast, Retargetable Python Compiler
Yes, LPython is a strict subset of Python, while Mojo is a strict superset of Python.

Both are valid and consistent approaches with their pros and cons, I listed some of them here: https://fortran-lang.discourse.group/t/fast-ai-mojo-may-be-t....

certik··on LPython: Novel, Fast, Retargetable Python Compiler
Yes, we can compare with pythran too. We should really compare against all the other compilers, but it's a lot of work to compare meaningfully: we don't want to put up benchmarks unless we are really sure they are solid, and as everyone knows, it's really hard to do benchmarks that are fair, as one has to have solid experience with both projects being benchmarked. But we did Numba and C++, so you for now you can compare pythran against them to get a decent idea.
certik··on LPython: Novel, Fast, Retargetable Python Compiler
No, we had LFortran, so naturally we have LPython now as the second frontend. We chose "L" in LFortran to be unspecified, although let's just say I live in Los Alamos, and we use LLLVM.
certik··on LPython: Novel, Fast, Retargetable Python Compiler
> Nitpick: The "Documentation" button on the header links to LFortran, not LPython.

Yes, I noticed too, thanks. We currently don't have a dedicated documentation for LPython and a lot of the LFortran documentation applies. We will eventually have a dedicated LPython documentation.

> With that said, I love the fact that this can generate executables, and I'm looking forward to trying it out in the future! Python compilers are really cool. I recently used Nuitka to build a standalone version of an app I wrote for my own use, so that I can run it on hosts that don't have Python installed (or don't have the right Python packages installed yet). This seems to be focused much more on speed.

We focus on speed, but we definitely create executables (no Python dependencies), just like any other C or Fortran compiler would. You only get a CPython dependency if you call into CPython (explicitly). Indeed creating such standalone executables simplifies the deployment and packaging issues: all the packages must be resolved at compile time, once you compile your application, there are no more dependencies on your Python packages (unless any of them calls into CPython of course).

> One thing which I didn't understand from the homepage: can I take vanilla Python code and AOT compile it with this?

Only if that vanilla Python compiles with LPython (since any LPython code is just Python code, a subset). So in general no. If you call CPython from your LPython main program (let's say), then you will get a small binary that depends on CPython to call into your Python code. I thought about somehow packaging the Python sources into the executable, similar to how PyOxidizer does it, but that's a project on its own almost. We will see what the community wants. If we can make LPython support a large enough subset of CPython, I think I would like to use LPython as is, since it's nice to not have any Python dependency and everything being high performance, essentially equivalent to writing C++ or Fortran.

certik··on LPython: Novel, Fast, Retargetable Python Compiler
Awesome, thank you. I knew about mypyc, but forgot. I just put it in:

https://github.com/lcompilers/lpython.org-deploy/pull/37

So now we have 25 compilers there.

Yes, the current syntax to call CPython is low level, you have to create an explicit interface. We can later make it more straightforward, such as using a `@python` decorator to a function, where inside you just do CPython. We always want to make it explicit, since it will be slow, and by default we want the LPython code to always be fast.

certik··on LPython: Novel, Fast, Retargetable Python Compiler
A compiler doesn't need to optimize. I think if it takes Python code, and translates it to something else, it's a compiler. An optimizing compiler is the one that will give you speedups.
certik··on LPython: Novel, Fast, Retargetable Python Compiler
Nuitka is a compiler. With list it at the bottom of https://lpython.org/, together with the other 23 Python compilers, now 24. :)
certik··on LPython: Novel, Fast, Retargetable Python Compiler
Right we support (currently a subset) of NumPy (just `from numpy import ...`) and SymPy (`from sympy import ...`) and some parts of the Python standard library. We want to support PyTorch, CuPy and other such libraries in a similar way, at least the subset that can be ahead of time compiled, which is quite large.

Yes, offloading to GPU we want to support naturally via NumPy syntax. We will look at this very soon, most likely via annotating that a given array lives on a GPU or host, and then array copy will copy it from host to device, etc.

certik··on LPython: Novel, Fast, Retargetable Python Compiler
> - presumably, since it is compiled, it does static checks on the code? How many statically-detectable bugs that are now purely triggered at runtime can be eliminated with LPython?

Yes, it does static checks at compile time. The only thing that we do at runtime (imperfectly right now, eventually perfectly) are array/list bounds checks, integer overflow during arithmetic or casting and such. Those checks would only run in Debug and ReleaseSafe modes, but not in Release mode. So for full performance, you choose the Release mode. For 100% safe mode, you would need to use ReleaseFast.

- does it deal with the unholy amounts of dynamism of Python? Can you call getattr, setattr on random objects? Does eval work? Etc. Quite a few Python packages use these at least once somewhere...

It deals with it by not allowing it. We will support as much as we can, as long as it can be robustly ahead of time compiled to high performance code. The rest you can always use just via our CPython "escape hatch". The idea is that either you want performance (then restrict to the subset that can be compiled to high performance code) or you don't (then just use CPython).

certik··on LPython: Novel, Fast, Retargetable Python Compiler
We put this sentence there to drive the point home that LPython competes with C++, C and Fortran in terms of speed. The internals are shared with LFortran, and LFortran competes with all other Fortran compilers, that traditionally are often faster than C++ for numerical code. I've been using Python for over 20 years and it's hard for me to imagine that writing Python could actually be faster than Clang/C++, somehow I always think that Python is slow. Right now we are still alpha and sometimes we are slower than C++. Once we reach beta, if an equivalent C++ or Fortran code is faster than LPython, then it should be a bug to report.
certik··on LPython: Novel, Fast, Retargetable Python Compiler
Hi William, nice to hear from you. We mention Cython at our front page (at the bottom): https://lpython.org/, together with the other 23 Python compilers that I know about (all of them are competitors, in a way). I am very familiar with Cython from about 10 years ago, but I have not followed the very latest developments. We can do a compilation time and runtime speed benchmarks against Cython in the next blog post. :)
certik··on LPython: Novel, Fast, Retargetable Python Compiler
We are interested for outside use, but right now we just copy (and sync) the libasr directory with LPython/LFortran manually. All the commits thus go first into either one of the frontends. We decided on this approach as it is currently the easiest for us to manage. Once the frontends are finished, we'll be able to just depend on the libasr as a 3rd party library.
certik··on Fast GPT-2 inference written in Fortran
The author here. If you have any questions, let me know.

If there is anybody here who wants to help parallelize this, let me know!

certik··on Advertise on DuckDuckGo Search
Thank you, we really appreciate it!
certik··on Advertise on DuckDuckGo Search
If you are not a Bing frontend, do you think it would be please possible to unblock our Fortran webpage at fortran-lang.org?

See here for details:

https://fortran-lang.discourse.group/t/fortran-lang-no-longe...

Bing decided to remove it for some reason and DDG now also removes it from search results, presumably because it uses Bing underneath. For Google it's #1 page when you search "fortran".

certik··on Optimization Without Derivatives: Prima Fortran Version and Inclusion in SciPy
This is a very important effort: if SciPy accepts this modern Fortran codebase, then I think the tide for Fortran will change.

My own focus is to compile all of SciPy with LFortran, we can already fully compile the Minpack package, here is our recent status update: https://lfortran.org/blog/2023/05/lfortran-breakthrough-now-...

And we have funding to compile all of SciPy. If anyone here is interested to help out, please get in touch with me. Once LFortran can compile SciPy, it will greatly help with maintenance, such as on Windows, WASM, Python wrappers, even modernization later on.

certik··on FastGPT: Faster than PyTorch in 300 lines of Fortran
I don't have much GPU experience myself. As the sibling comment said, there are Fortran compilers that can offload to GPU, there is also Cuda Fortran. There is OpenMP offloading. I think LLVM can also target it somehow, and I would like to support it in LFortran, a compiler that we are developing. In general I am hoping people more experienced with GPU would be interested in helping out.
certik··on FastGPT: Faster than PyTorch in 300 lines of Fortran
It's here: https://github.com/certik/theoretical-physics/, I was hoping more people would contribute to the effort, but so far I didn't manage to spark enough interest. It's open source, it's out there and if I find at least a single person willing to contribute to get it polished, the development will pick up.

In the absence of it, I would have to drive it hard as my main effort, but right now my main effort is LFortran/LPython, we are making excellent progress there, so I am not spreading too thin until the compilers are delivered.

If anyone is interested in the theoretical physics book, please let me know!

certik··on FastGPT: Faster than PyTorch in 300 lines of Fortran
Excellent questions. One is import time and model loading time where PyTorch is very slow, and it gets much worse for the larger models, for the 1558M model PyTorch is 24s to start, while fastGPT is 1s, about 24x speedup.

I am still studying the performance of the inference itself, it's really hard to do meaningful benchmarks that I can trust. The ones in my blog posts should be solid, I've eventually managed to control all variables. For example, my faster tanh() implementation initially showed around 20% speedup, but after I controled everything, I only see 4% speedup without caching, and less than that with caching.

I think the main advantage of Fortran is that all I did was a rewrite (two afternoons) and I right away saw performance better than PyTorch, which is a highly optimized production code, developed by thousands of professionals. After controling for everything and doing a fair comparison, it's only slightly faster (at the moment!), but that's still quite an impressive result I think. And using Accelerate, it's a lot faster. I am guessing this problem is limited by matrix-matrix multiplication, in which case even Python is fast on single core (even pure Python/NumPy picoGPT is competitive after my PR), which the results seem to show.

Thanks for the links, I'll try PyTorch with Accelerate and report back.

I don't know regarding GPU, we'll have to see.

But in general, right now the code is not parallel, it runs on single core, the only parallelism comes from OpenBLAS. It's a great foundation to now parallelize it and see how it performs. In other words, with Fortran you start "fast" right away, and then you can try speeding it up from there. While in Python it is quite a lot of work to even get it to this performance.

certik··on FastGPT: Faster than PyTorch in 300 lines of Fortran
Sorry about that. I was using some Hugo theme and I haven't checked on mobile. I should redo my webpage to be mobile friendly.
certik··on FastGPT: Faster than PyTorch in 300 lines of Fortran
The author here. I am happy to answer any questions.
← PreviousPage 3 of 5Next →