519 karma · joined March 22, 2019
Internally they each parse the syntax to AST, then have some kind of an intermediate representation (IR), do some optimizations and generate code. The differences are in the details of the IR and how the compiler is internally structured.
Regarding the type inference, this is for a blog post on its own. See this issue for now: https://github.com/lcompilers/lpython/issues/2168, roughly speaking, there is implicit typing (inference), implicit declarations and implicit casting. Rust disallows implicit declarations and casting, but allows implicit typing. As shown in that issue they only meant to do single line implicit typing, but (by a mistake?) allowed multi-statements implicit typing (action at a distance). LPython currently does not allow any implicit typing (type inference). As documented at the issue, the main problem with implicit typing is that there is no good syntax in CPython that would allow explicit type declaration but implicit typing. Typically you get both implicit declaration and implicit typing, say in `x = 5`, this both declares `x` as a new variable as well as types it as integer. C++ and Rust does not allow implicit declarations (you have to use `auto` or `let` keywords) and I think we should not do either. We could do something like `x: var = 5`, but at that point you might as well just do `x: i32 = 5`, use the actual type instead of `var`.
Indeed, I think for large projects you want:
* generate a binary
* fast compilation in Debug mode
* fast runtime in Release mode
Sometimes you want JIT, so we support it too, but I personally don't use it, I write the main program in LPython and just compile it to a binary.
> Will LPython have the ability to generate AoT compiled libraries in addition to executables?
Yes, you can do it today. Just compile to `.o` file and create a library. We use this library feature in production at my company today. If you run into issues, just ask us, we'll help.
If you have some ideas how to best communicate this, let me know.
The best approach that I know right now is to say that we support a strict subset of Python, and if it compiles, it will be fast. The rest of Python you have to call explicitly and it will be slow and you get a CPython dependency in the binary that we generate, but you can do it. That way there is a clear distinction what gets compiled via our compiler into high performance machine code and what gets dispatched via CPython (it will eventually get "compiled", but it will just call CPython).
We could target MLIR later, right now we are just targeting LLVM.
Both are valid and consistent approaches with their pros and cons, I listed some of them here: https://fortran-lang.discourse.group/t/fast-ai-mojo-may-be-t....
Yes, I noticed too, thanks. We currently don't have a dedicated documentation for LPython and a lot of the LFortran documentation applies. We will eventually have a dedicated LPython documentation.
> With that said, I love the fact that this can generate executables, and I'm looking forward to trying it out in the future! Python compilers are really cool. I recently used Nuitka to build a standalone version of an app I wrote for my own use, so that I can run it on hosts that don't have Python installed (or don't have the right Python packages installed yet). This seems to be focused much more on speed.
We focus on speed, but we definitely create executables (no Python dependencies), just like any other C or Fortran compiler would. You only get a CPython dependency if you call into CPython (explicitly). Indeed creating such standalone executables simplifies the deployment and packaging issues: all the packages must be resolved at compile time, once you compile your application, there are no more dependencies on your Python packages (unless any of them calls into CPython of course).
> One thing which I didn't understand from the homepage: can I take vanilla Python code and AOT compile it with this?
Only if that vanilla Python compiles with LPython (since any LPython code is just Python code, a subset). So in general no. If you call CPython from your LPython main program (let's say), then you will get a small binary that depends on CPython to call into your Python code. I thought about somehow packaging the Python sources into the executable, similar to how PyOxidizer does it, but that's a project on its own almost. We will see what the community wants. If we can make LPython support a large enough subset of CPython, I think I would like to use LPython as is, since it's nice to not have any Python dependency and everything being high performance, essentially equivalent to writing C++ or Fortran.
https://github.com/lcompilers/lpython.org-deploy/pull/37
So now we have 25 compilers there.
Yes, the current syntax to call CPython is low level, you have to create an explicit interface. We can later make it more straightforward, such as using a `@python` decorator to a function, where inside you just do CPython. We always want to make it explicit, since it will be slow, and by default we want the LPython code to always be fast.
Yes, offloading to GPU we want to support naturally via NumPy syntax. We will look at this very soon, most likely via annotating that a given array lives on a GPU or host, and then array copy will copy it from host to device, etc.
Yes, it does static checks at compile time. The only thing that we do at runtime (imperfectly right now, eventually perfectly) are array/list bounds checks, integer overflow during arithmetic or casting and such. Those checks would only run in Debug and ReleaseSafe modes, but not in Release mode. So for full performance, you choose the Release mode. For 100% safe mode, you would need to use ReleaseFast.
- does it deal with the unholy amounts of dynamism of Python? Can you call getattr, setattr on random objects? Does eval work? Etc. Quite a few Python packages use these at least once somewhere...
It deals with it by not allowing it. We will support as much as we can, as long as it can be robustly ahead of time compiled to high performance code. The rest you can always use just via our CPython "escape hatch". The idea is that either you want performance (then restrict to the subset that can be compiled to high performance code) or you don't (then just use CPython).
If there is anybody here who wants to help parallelize this, let me know!
See here for details:
https://fortran-lang.discourse.group/t/fortran-lang-no-longe...
Bing decided to remove it for some reason and DDG now also removes it from search results, presumably because it uses Bing underneath. For Google it's #1 page when you search "fortran".
My own focus is to compile all of SciPy with LFortran, we can already fully compile the Minpack package, here is our recent status update: https://lfortran.org/blog/2023/05/lfortran-breakthrough-now-...
And we have funding to compile all of SciPy. If anyone here is interested to help out, please get in touch with me. Once LFortran can compile SciPy, it will greatly help with maintenance, such as on Windows, WASM, Python wrappers, even modernization later on.
In the absence of it, I would have to drive it hard as my main effort, but right now my main effort is LFortran/LPython, we are making excellent progress there, so I am not spreading too thin until the compilers are delivered.
If anyone is interested in the theoretical physics book, please let me know!
I am still studying the performance of the inference itself, it's really hard to do meaningful benchmarks that I can trust. The ones in my blog posts should be solid, I've eventually managed to control all variables. For example, my faster tanh() implementation initially showed around 20% speedup, but after I controled everything, I only see 4% speedup without caching, and less than that with caching.
I think the main advantage of Fortran is that all I did was a rewrite (two afternoons) and I right away saw performance better than PyTorch, which is a highly optimized production code, developed by thousands of professionals. After controling for everything and doing a fair comparison, it's only slightly faster (at the moment!), but that's still quite an impressive result I think. And using Accelerate, it's a lot faster. I am guessing this problem is limited by matrix-matrix multiplication, in which case even Python is fast on single core (even pure Python/NumPy picoGPT is competitive after my PR), which the results seem to show.
Thanks for the links, I'll try PyTorch with Accelerate and report back.
I don't know regarding GPU, we'll have to see.
But in general, right now the code is not parallel, it runs on single core, the only parallelism comes from OpenBLAS. It's a great foundation to now parallelize it and see how it performs. In other words, with Fortran you start "fast" right away, and then you can try speeding it up from there. While in Python it is quite a lot of work to even get it to this performance.