HNHacker News
TopNewBestAskShowJobs

certik

519 karma · joined March 22, 2019

https://ondrejcertik.com/
submissionscomments
certik··on Lfortran: Modern interactive LLVM-based Fortran compiler
There is old Flang, which motivated me to start LFortran. The new Flang, which presumably you are referring to, started possibly in the same month as LFortran, but we didn't know about each other.

It's best if you ask Flang developers what they see as the advantages of Flang over LFortran. From my biased perspective, LFortran can run interactively, it is fast to compile the compiler (30s on my laptop) and LFortran compiles your code very quickly (especially with our direct x64 or C backends). It runs in a browser: https://dev.lfortran.org/. We have many backends (LLVM, C, C++, Julia, WASM, x64). We plan to add Python and Fortran (the latter could be used to modernize your old Fortran code). It is easy to add new backends, and it is also easy to add new frontends, so we have LPython and LFortran as two thin frontends, to our intermediate representation that we call ASR (Abstract Semantic Representation). The internal design is simple, so a small team can develop LCompilers at a fast pace. New contributors without any prior compiler experience get up to speed very quickly (typically a few weeks or even less).

We are still in alpha, which means it is expected to break for your code (and when it does, please report all bugs!). To choose between them, I recommend to test them out and pick the one that you like the most, based on your criteria. Note that the most mature and widespread open source Fortran compiler is GFortran.

certik··on Lfortran: Modern interactive LLVM-based Fortran compiler
Turns out Fortran is a great fit for LLM, I am not joking: https://github.com/certik/fastGPT/.
certik··on Lfortran: Modern interactive LLVM-based Fortran compiler
Yes, we could make the WASM->x64 standalone. The main motivation is speed of compilation. We do not do any optimizations, but we want to generate the x64 binary as quickly as possible, with the idea that it would be used in Debug mode, for development. Then for Release mode you would use LLVM, which is slow to compile, but good runtime performance. And since we already have ASR->WASM backend (used for example at https://dev.lfortran.org/), maintaining WASM->x64 is much simpler than ASR->x64 directly.
certik··on Lfortran: Modern interactive LLVM-based Fortran compiler
My experience with LLVM so far has been that is possible to get maximum speed as long as we generate the correct and clean LLVM IR, and do many of the high level optimizations ourselves.

If LLVM has any downsides, it is that it is hard to run in the browser, so we don't use it for https://dev.lfortran.org/, and that it is slow to compile (both LLVM itself, as well as it makes LFortran slow to compile, compared to our direct WASM/x64 backends). But when it comes to runtime performance of the generated code, LLVM seems very good.

certik··on Fortran
> I didn’t intend to say that «not official» means «bad».

Yes, I know, I understood that. Thank you. I was reacting to the term official, since it is an interesting question what official means for Fortran, since in a way it was "abandoned", there was the ISO committee as the only official body, but nobody wanted to provide an official place for it on the internet (the most official place was actually the Fortran Wikipedia article). So we volunteered that. In a way, I think it could already be treated as official, or if not, hopefully in a few years most people will treat it as the de-facto official. Kind of like the git webpage has migrated from https://repo.or.cz/ to https://git-scm.com/, but both in a way are unofficial, but the second one is a nice modern webpage and most people liked it and accepted it as official. I don't know the people behind either webpage.

certik··on Fortran
If you like both Python and Fortran, checkout LPython: https://lpython.org/, which shares internals with LFortran (https://lfortran.org/), thus delivering exactly the same performance.
certik··on Fortran
Check out LPython as well (also in alpha): https://lpython.org/
certik··on Fortran
It's a great question. I think we answered it in one of our initial blog posts when we started LFortran:

https://lfortran.org/blog/2019/05/why-to-use-fortran-for-new...

Let me know if you have any follow up questions.

certik··on Lfortran: Modern interactive LLVM-based Fortran compiler
A simple example is returning an allocatable array from a function, where the Fortran compiler can decide to allocate on a stack instead, or even inline the function and eliminate completely. While in C the compiler would need to understand the semantics of an allocatable array. If you use raw C pointer and malloc, and use Clang, my understanding is that Clang translates quite directly to LLVM and LLVM is too low level to optimize this out, depending on the details of how you call malloc.

Of course, you can rewrite your C code by hand to generate the same LLVM code from Clang, as LFortran generates for the Fortran code. So in principle I think anything can be done in C, as anything can be done in assembly or machine code. But the advantage of Fortran is that it is higher level, and thus allows you to write code using arrays in a high level way and do not have to do many special things as a programmer, and the compiler can then highly optimize your code. While in C very often you might need to do some of these optimizations by hand as a user.

certik··on Lfortran: Modern interactive LLVM-based Fortran compiler
The original author of LFortran. Great question.

We designed LFortran to first "raise" the AST (Abstract Syntax Tree) to ASR (Abstract Semantic Representation). The ASR keeps all the semantics of the original code, but it is otherwise as abstract/simple as possible. Thus by definition it allows us to do any optimization possible, as ASR->ASR optimization pass. We do some already, we will do many more in the future. This optimizes all the things where you need to know the high level information about Fortran. Then once we can't do any more optimizations, we lower to LLVM. If in the future it turns out we need some representation between ASR and LLVM, such as MLIR, we can add it.

We also have a direct ASR->WASM and WASM->x64 machine code, and even direct ASR->machine code, but the ASR->LLVM backend is the most advanced, after that probably our ASR->C backend and after that our ASR->WASM backend.

certik··on Lfortran: Modern interactive LLVM-based Fortran compiler
I am the original author of LFortran. What kind of comparison would you like to see? If you have any specific questions, I am happy to answer.
certik··on Lfortran: Modern interactive LLVM-based Fortran compiler
Thanks! I know, LLM came much later after LLVM. But if you are interested in LLM that LFortran can compile, check out: https://github.com/certik/fastGPT/.
certik··on Fortran
Very well put. Also a subset of Python: https://lpython.org/, and it shares the internals with https://lfortran.org/, both run at the same speed. If you have any feedback, please let us know.
certik··on Fortran
Yes, I lead the development of these two compilers:

https://lpython.org/ https://lfortran.org/

LPython compiles Python at the Fortran speed, since both compilers share the same internals.

certik··on Fortran
One of the co-founders of the fortran-lang effort. Actually, I would say it is as official as it can be. If you google "fortran", it comes up as first. We provide the home for all of fortran, all of the compilers (commercial and open source), documentation, forum, etc. Wikipedia lists fortran-lang.org as the webpage. Both of the initial co-founders of fortran-lang are part of the standards committee (but the fortran-lang effort is now much bigger than them), and many of the standards committee members are participating in the discourse. I don't think it is the job of the ISO Fortran Standards Committee to maintain a website and a discussion forum, but rather the job is to standardize the language. So it is a complementary effort. To maintain the language has to be an effort of the whole Fortran community, which includes users, compiler writers, the standards committee, and anybody who is able to help.
certik··on Fortran
I am one of the co-founders of the fortran-lang effort. I did it while I was at LANL, where I worked as a scientist for almost 9 years. I think the report is overly pessimistic. Here is my full reply on the report: https://fortran-lang.discourse.group/t/an-evaluation-of-risk....
certik··on LPython: Novel, Fast, Retargetable Python Compiler
It doesn't do garbage collection. Variables and arrays get deallocated when they go out of scope (similar to C++ RAII). When you return from a function, it will be a copy, although I think we elide the copy in most cases. We restrict the Python subset in such a way that this exactly works with both LPython and CPython.

Local variables are on stack. Lists, dicts, sets and arrays use heap, unless the length is known at compile time, where we might use a stack, but I think we need to make it configurable, since one can run out of stack quite easily this way.

certik··on LPython: Novel, Fast, Retargetable Python Compiler
Thank you! Please report all bugs that you find once you try it.
certik··on LPython: Novel, Fast, Retargetable Python Compiler
The direct x86 backend (in fact we have two, one is direct ASR->x86, but it only supports a few things so far; the other one is ASR->WASM->x86, which is a lot more complete and even supports 64bits) seems comparable to LLVM without any optimizations. The main added value of LLVM are all the low level optimizations that it can do. Also our LLVM backend is the most complete, it's our main backend. We will eventually catch up with the other backends.

I think we only require C++11, I don't think we use anything newer than that. If we do, it wouldn't be difficult to change. I think we use `std::filesystem`, which is C++17, but if that is a problem, we could add some emulation of this for C++11.

certik··on LPython: Novel, Fast, Retargetable Python Compiler
Only if the libraries used the subset of Python that LPython can compile. So currently no.
certik··on LPython: Novel, Fast, Retargetable Python Compiler
By being a superset of Fortran and subset of Python. It turns out the features map on each other almost 1:1, and the differences can be taken care of by the respective frontends, so the abstracted ASR maps perfectly.
certik··on LPython: Novel, Fast, Retargetable Python Compiler
That, but also being simpler and higher level, having multidimensional arrays in the language itself and simpler semantics (such as you cannot just take a pointer to an arbitrary variable, it has to be marked with "target"), no exceptions, and so on. What carries over to the IR today are all the language semantic features, such as all the array operations (minloc, maxval, sum, ...) and functions (sin, cos, special functions) as well as all the other features without any lowering, and we then do optimizations at this high level, then only at the end we lower (say to LLVM). Python/NumPy can be optimized in exactly the same way, and that's what LPython does. I think C++ can also be compiled this way, but the frontend would have to understand basic structures like `std::vector`, `std::unordered_map` as well as arrays (say xtensor or Kokkos, whatever library you use for arrays), and lift it to our high level IR. Possibly we would have to restrict some C++ features if they impeded with performance, such as exceptions --- I am not an expert on C++ compilers, I am only a user of C++.
certik··on LPython: Novel, Fast, Retargetable Python Compiler
It's harder to imagine for a Python compiler, so let's just focus on LFortran (since LPython delivers exactly the same performance, due to sharing the middle end and backends). The Fortran compilers traditionally were faster than C++, and almost always (even today) are at least as fast as C++, due to the Fortran language being simpler and higher level, designed to allow good optimizations. LFortran competes with other compilers as well as C++ compilers and our goal indeed has to be to be at least as fast as the competition. We currently are sometimes faster sometimes slower, but we are in the same league, so that's a good start.

Regarding LLVM: my experience so far is that LLVM is indeed amazing what it can do in terms of optimizations. It's very very good. However, it is not all LLVM. As our benchmarks in the blog post show, we compare Numba, Clang and LPython, all three of which use LLVM. But we get vastly different performance for what seems to look like identical initial code. To know exactly why, we would have to meet with the Numba and Clang developers and study this, I suspect Clang lowers to LLVM too soon, and uses C++ to do abstractions (like `std::vector` or `std::unordered_map`) and perhaps it can't quite get the top performance this way. Numba perhaps doesn't get all the types as tight as LPython, or perhaps implements some things not as efficiently, or perhaps doesn't apply as good optimizations before lowering to LLVM. I suspect LLVM gets the best performance if the compiler generates as straightforward LLVM IR code as possible, without layers and layers of abstractions that might not end up being "zero cost" in practice.

certik··on LPython: Novel, Fast, Retargetable Python Compiler
Yes. ASR is as abstract as it can be, but it is still faithful to the original language, no information has been lost, nothing was lowered.
certik··on LPython: Novel, Fast, Retargetable Python Compiler
Yes, I thought about this too. Also LLMs can or will be able to translate from one language to another, so perhaps the fact that LFortran/LPython can translate Fortran/Python to other languages like C++ or Julia might not be useful.

My approach is that it is still unclear to me what exactly will be possible in the future, while I know exactly how to deliver these compilers today. I suspect a traditional compiler will be more robust and also a lot faster than an LLM for tasks like translation to another language or compilation to binary. And speed of compilation is very important for development from the user perspective.

Conclusion: I don't know what the future will bring, but I suspect these compilers will still be very useful.

certik··on LPython: Novel, Fast, Retargetable Python Compiler
Yes, we have Shedskin in the list at the bottom of https://lpython.org/. Note that the Shedskin compiler is written in Python, so the speed of compilation might be lower than other Python compilers written in C++. Unless it compiles itself, that would be interesting. We thought about eventually writing LPython in LPython, but for now we are focusing on delivering, so we are sticking to C++.
certik··on LPython: Novel, Fast, Retargetable Python Compiler
Thank you for the encouragement. I can answer / clarify your comments:

> This is way too similar to Mojo (language) by Modular. At least at some level conceptually. All in a good way.

Yes, the main difference is that Mojo is (or will be) a strict superset of Python, while we are a strict subset (but you can call the rest of Python via a decorator).

> I think - the performance gains are coming from the Python syntax being transliterated to an LFortran intermediate representation (much like Numba converts code to LLVM IR).

Correct. We use the same IR as LFortran and then we lower to LLVM IR.

> Any calls to CPython libraries are made using a special decorator, which might be doing interop using the CPython API. My guess is, this will come with a performance penalty. More so if you’re using Numpy or Scipy, as you’ll be going through several layers of abstractions and hand-offs.

Yes, it calls CPython, so it's slow.

> This is because Numba, Pythran and JAX (in a way) get around this by reimplementing a subset of Numba/Scipy/other core libraries. Any call to a supported function is dynamically rerouted to the native reimplementation during JIT/AOT compilation.

We do as well: we support a subset of NumPy directly (eventually most of NumPy). We also support a very small subset of SymPy. Over time we add more support to more basic libraries. The rest you can call via CPython, but slow. For SymPy we'll experiment building it on top (at least some modules, like limits) and compile using LPython. Given that any LPython code is just Python, this might be a viable way, as long as we support enough of Python directly.

> I’d be interested in seeing how far LPython can tolerate regular Python code, with a ton of CPython interop and class use.

We support structs via `@dataclass`, but not classes yet (although LFortran does to some extent, so we'll add support soon to LPython as well). For regular Python call LPython will give nice error messages suggesting to type things. Once you do and it compiles, it will run fast.

> In any case, glad to see more competition. Not to take anything from the authors - this is a massive effort on their part and achieves some impressive results. Mojo has VP money behind it - AFAIK, this is a pure volunteer driven effort, and I’m grateful to the authors for doing it!

We are supported by my current company (GSI Technology) as well as by NumFOCUS (LFortran), GSoC and other places; we have a very strong team (5 to 10 people). In the past I was supported by Los Alamos National Laboratory to develop LFortran. I have delivered SymPy as a physics student with no institutional support. So I have experience doing a lot with very little. :)

certik··on LPython: Novel, Fast, Retargetable Python Compiler
Thank you! I opened up an issue to do this: https://github.com/lcompilers/lpython/issues/2220.
certik··on LPython: Novel, Fast, Retargetable Python Compiler
See my comment here for all the details regarding implicit typing and why we don't currently do it: https://news.ycombinator.com/item?id=36920963. But we give you nice error messages if types don't match.
certik··on LPython: Novel, Fast, Retargetable Python Compiler
My understanding from Mojo's plans is that they want to compile all of Python via their compiler (eventually), and then extend Python with extra syntax that will compile to high performance. I think right now they might not compile all of Python yet, so you are right they are neither a subset nor a superset, but once they deliver on their plans, they will become a superset.
← PreviousPage 2 of 5Next →