Without a bootstrap process like this one, it could happen that you run git pull and then can't build the compiler because it's using a new feature that the version of the compiler you have doesn't support yet.
The wasm process ensures that you can always build the latest commit.
I maintain what i said, i do not understand how this is a good idea
WASM file is produced by compiling new compiler, one written in Zig.
They are compiling the current compiler to wasm and then using that compiler to build future versions of their compiler.
In other words, they are doing the described rust approach but instead using a platform agnostic target instead of doing a binary. That allows them to build on any platform that has a C compiler and to use current language features without needing manually backport them.
They could directly target C or C++ but that runs a greater risk of accidentally generating UB. Targeting a bytecode decreases that risk.
Imagine a scenario where you are testing out a brand new RISC-V development board. The vendor has only provided a C compiler for this board as is often the case. You want to be able to use the zig language to write programs for your new development board but no zig compiler exists for this board yet. That means you need to compile the zig compiler from source. The latest version of the zig compiler is written in zig. Again you don't have a zig compiler so how will you compile the zig compiler from source? You need a way to go from C compiler to Zig compiler. That's what this is describing. It does not make sense to maintain two completely separate versions of the compiler. The "real" one written in Zig and the "bootstrap" one for compiling from C. So the zig source is compiled into WASM and stored in the repo. On a system with only a C compiler this WASM can be ran instead of a native zig binary. The WASM version can then be used to compile an actual native zig binary.
EDIT: since my performance critical sentence is getting misunderstood: https://news.ycombinator.com/item?id=33914718
As a generalization, the people who are motivated enough to work on a language, want to use that language. It's only natural that they would want to write their compiler in that language too, if practical. Contributors to the Zig project would on average probably be more proficient and productive in Zig than they would be in a language they don't care about so much.
It's also just helpful to have the people who are designing the language working in that language regularly in the context of a sizable and nontrivial project.
Have they written am incremental compiler then? Or just an old-fashioned slow batch one? Compiler architecture matters much more than whether it's runtime has a GC or not.
It's more important than the choice of programming language. Not all C++ is fast.
> You can't implement an incremental compiler in Zig?
I never said this. This is a thinly veiled ad hominem.
I understood your original comment to imply that you disagree with their choice of self-hosting the Zig compiler because they should instead have focused on architectural improvements, in which case I disagree because the two things are completely orthogonal in my mind—the benefits of self-hosting have little to do with performance, and certainly don't come at the cost of it. I apologize if that's not what you were implying.
> This is a thinly veiled ad hominem.
I never personally attacked you, so I'm not sure why you think this is ad hominem. That's a very uncharitable interpretation of my comment.
https://kristoff.it/blog/zig-new-relationship-llvm/
https://mobile.twitter.com/andy_kelley/status/15651098597708...
Is there a reason why keeping a compiler in, say, C would be a bad idea long-term?
Today’s managed languages are very fast. For example, if Java is not fast enough for your HFT algorithm, than nor is C++ or generic CPUs even! You have to go the custom chip route then. Where there is a significant difference between these categories is memory usage and predictability of performance. (In other applications, e.g. video codecs you will have to write assembly by hand in the hot loops, since here low-level languages are not low level enough). Since these concerns not apply to compilers, I don’t think that a significant performance difference would be observable between, say a java and zig implementation of a certain compiler.
The alternative would be to port your Haskell compiler to the new CPU too in order to set up a self hosting toolchain. Much more work involved, because you not only have to be proficient in Haskell, but you need to have Haskell compiler implementation skills in addition to your own compiler.
Thank you!
My full-time job is making a compiler for a high-level language, and I only considered systems languages (e.g. Zig, Rust) as contenders for what to write it in - solely because compiler performance is so critical to the experience of using the compiler.
In our case, since the compiler is for a high-level language, we plan never to self-host, because that would slow it down.
To me, it seems clear that taking performance very seriously, including language choice, is the best path to delivering the fastest feedback loop for a compiler's users.
I honestly fail to see why would a lower level language be faster, especially that compilers are notorious for their non-standard allocation and life cycle patterns, so a GC might actually be faster here.
Didn’t know about that backend though, will check it out, thanks!
The nonstandard allocation and lifecycle patterns are a major part of the reason I want a systems language and not a GC - it means I have strictly more control over when allocations happen, I can do cheap phase-oriented allocations and deallocations with arenas, etc.
Rust's compiler is an interesting example. It was originally implemented in OCaml (which has a reputation for being a GC'd language with good runtime performance), and then rewrote to Rust in order to self-host - and got faster. In contrast, the Go team rewrote from a systems language to Go (which also has a good reputation for runtime performance), again in order to self-host, and it got slower.
Nonstandard lifetimes are not really helped with arena allocators though, and not everything is needed in each phase, or is there that divided phases at all. But you may be right, I honestly can’t tell with certainty.
Rust was originally written in OCaml before being self-hosted, and it wouldn't be as fast (or would be even slower ;) ) today if it was still OCaml.
And remember, low-level =/= poor abstractions. I think there are several novel abstractions available in Zig which the compiler devs probably want to make use of themselves.
They might have good language abstractions, but manual memory management is simply an orthogonal implementation detail to solving a problem — dealing with that is simply more work and more leaky abstractions.
Actually, yes.
A compiler is not just a dumb filter that eats text and spits out machine code. It can provide infrastructure to other tools (like linters, analyzers, LSP servers...), and even allow importing and using parts of it in user programs.
Some of this can't be done in a different language (like Haskell) without some very crazy FFI, and might add an extra runtime dependency for those tools, which might not always be desirable.
The number of ignorant comments in this thread is astounding. All this criticism feels like it was written by back-seat drivers who have no clue about the complexities of language design or compiler implementation.
Why wouldn't you? I understand why not for high level languages like Python or Ruby (since they're interpreted) but not for low level ones. Rust for example is also bootstrapped.
Bootstrapping is actually a very old tradition that any programming language must implement in order to showcase the power of language.
You'd want faster compilation so that you can test your changes without it breaking your flow where having to wait a 1-5 minutes means you'll end up reading HN or checking chat for 10 or so minutes. That's also why there's interest in hot-code reloading and incremental linking to make it faster as it will further reduce compilation to just the changes you've made and nothing more.
I haven't seen much opportunity to improve algorithm performance because which algorithms are applicable is heavily constrained by language design.
In my experience the difference between something like Zig and a language with managed memory is large.
https://media.handmade-seattle.com/practical-data-oriented-d...