Virgil: A fast and lightweight programming language that compiles to WASM
github.com
github.com
One of the things that excites me about Virgil is that is completely self-hosted (the runtime is fully implemented in Virgil, and it can compile itself).
Curiously enough, the person that created Virgil (programming language) and Wizard (Wasm runtime) was Ben L. Titzer, who worked on the V8 team at Google and co-created Wasm. I'm pretty excited with what he can bring to the table :)
Update: I'm trying to have virgil to self-compile into Wasm/WASI, and will try to upload it to WAPM [2] so everyone can use it easily (as the first compilation of v3c currently requires the JVM)... stay tuned!
[1]: https://github.com/titzer/wizard-engine
[2]: https://wapm.io/
It will bootstrap using whichever stable binary appears to work on your platform.
The compiler does have support for specifying arbitrary Wasm imports and I wrote most of the "System" module against WASI. The only thing missing is 'chmod' to flip the permissions on a just-generated executable, but other than that, it should be possible to bootstrap on WASI.
“It goes without saying that any function that throws an exception which isn't caught is wrong and must be punished.”
> Q: Is this serious?
> A: Eternal moral vigilance is no laughing matter.
EDIT: It wasn't me that was downvoting you FYI, but go ahead and down vote me. :)
I'm not a "professor" but as a software engineer with 35 years in this industry I can say that new languages should avoid GC's (as in, generational and related) and stick to either ARC or Rust-like compile-time memory management.
Just because the original comment is by, let's say, a prominent figure, doesn't make it right.
P.S. I rarely downvote out of disagreement, only for comment quality.
With respect, and much less experience than You, I really don’t think so. I believe the majority of languages are better off being managed. Low-level languages do have their place and I am very happy for Rust that does bring some novel idea to the field. But that lower detail is very much not needed for the majority of applications. Also, ARC is much much slower than a decent GC, so from a performance perspective as well, it would make sense to prefer GCd runtimes.
ARC, while in certain cases can elide, will still in most case have to issue atomic increases/decreases that are the slowest thing on modern processors. And on top it doesn’t even solve the problem completely (circular references), mandating a very similar solution than a tracing GC (as ref counting is in fact a form of GC, tracing looking it live edges between objects, ref counting looking at dead edges)
Someone also conducted tests, for the same tasks and on equivalent CPU's Android requires 30% more energy and 2x RAM compared to iOS. Presumably the culprit is the GC.
It is not an accident that on powerful server machines all FAANG companies use managed languages for their critical web services, and there is no change on the horizon.
It is actually quite rare that companies think of their infrastructure costs, it's usually just taken for granted, plus that there aren't many ARC languages around.
Anyway I'm now rewriting one of my server projects from PHP to Swift (on Linux) and there's already a world of difference in terms of performance. For multiple reasons of course, not just ARC vs. GC, but still.
Just because GC can be a bottleneck doesn’t mean it is bad or that alternatives wouldn’t have an analog bottleneck. Of course one should try to decrease the number of allocations (the same way you have to do in case of RC as well), but there are certain allocation types that simply have to be managed. For those a modern GC is the best choice in most use case.
I am also working on a new Wasm engine and I have a neat trick (TM) where the implementation of the proposed Wasm GC features are implemented by just reusing the Virgil GC. So the engine is really a lot simpler, doesn't need a handle mechanism, doesn't have a GC itself, and the one Virgil GC has a complete view of the entire heap, instead of independent collectors that need to cooperate.
Go does have a garbage collector though, so maybe the conflating of GC with safety (not by you, but earlier in thy thread) is a bit misleading.
This makes more sense in light of a comment by the author in a previous discussion about C ABIs [1]:
> Virgil compiles to tiny native binaries and runs in user space on three different platforms without a lick of C code, and runs on Wasm and the JVM to boot. [...] No C ABI considerations over here.
For the subset that Virgil uses to get off the ground, I haven't been broken by the kernel changing system calls. MacOS has been a pain for a number of other reasons though, not the least of which is deprecating 32-bit altogether...with a student's help I finally got around to generating x86-64 Mach-O binaries. That works on x86 macs again. But something is still wonky and they don't run under Rosetta 2.
Linux is rock solid though. I've never been broken by the kernel.
I think Go was the first language that tried to buck that trend and it did not go well (aside from Linux).
This single function is all you need to do literally anything on x86_64 Linux from writing to a file descriptor to graphics card ioctls:
https://github.com/matheusmoreira/liblinux/blob/master/sourc...
In all others, interfacing with the kernel directly eventually leads to breakage because the people in charge change the system calls. Go binaries broke on OS X because of this. We're not meant to bypass their system libraries.
In case it's useful for future readers, here are some samples: https://github.com/titzer/virgil/tree/master/doc/tutorial/ex...
This is an example from itzer/virgil [1]:
def fib(i: int) -> int {
if (i <= 1) return 1;
return fib(i - 1) + fib(i - 2);
}
And an example from munificent/vigil [2]: def fib(n):
if n < 2:
result = n
else:
result = fib(n - 1) + fib(n - 2)
swear result >= 0
return result
[1] https://github.com/titzer/virgil/blob/master/doc/tutorial/Me...Once you have that, you can post your docs in GitHub Pages (or something like Netlify[4] or Cloudflare[5]), they both can run a command to build your website (from markdown to html) every time you push to a branch, and then serve the HTML generated as a static site.
Before this though, your language seems similar enough to others (maybe Java or C#?) that if you tell the converter to use those languages, you'll get decent enough highlighting. I did this to highlight Zig code before it became supported by telling the converter it was typescript code (coincidentally, many keywords seem to have aligned well enough)!
[1] https://github.com/Depado/bfchroma/
[2] https://github.com/alecthomas/chroma#supported-languages
[3] https://pygments.org/docs/lexerdevelopment/
[4] https://www.netlify.com/blog/2016/10/27/a-step-by-step-guide...
Or are there large performance differences between compiled WASM?
Let's say there's a Python to C compiler (there is) and let's say there's a C++ to C compiler (there was, may still be idk). The performance characteristics of programs compiled via both won't be very similar at all. They can't be any faster than C. They can both be much slower than it. They'll differ wildly from each other.
https://github.com/titzer/virgil/tree/master/doc/tutorial/ex...
I'm asking also because if the compiler supports JVM as a target, it shouldn't be hard to add support for a Dalvik target; however, u64 is also a problem in case of this one.
One more question I have is about FFI - I didn't find a mention of it after a quick skim; can I call functions from some thirdparty JARs or JS/WebAPI?
> One more question I have is about FFI - I didn't find a mention of it after a quick skim; can I call functions from some thirdparty JARs or JS/WebAPI?
For the Wasm target, Virgil allows you to write an imported component, so the module that gets generated has the imports with the signatures that you want. You can then load the module in JS and supply Web bindings and such. That latter process is quite clunky, but in theory gives you access to any API that can be expressed in terms of wasm (before externref).
I wonder if they have a bootstrap compiler, so you can build entirely from source without their pre-existing binaries.
Virgil bootstrapped for the first time using an interpreter I wrote in Java, back around 2009. When the new compiler could finally compile itself well enough to be stable, I checked in the first bootstrap binary, a jar file. Since then, every once in a while (41 times so far), when a major set of bugfixes or new features is done, I've rev'd stable by checking in binaries that are generated by the first compiling the existing code with the stable compiler, and then compiling the compiler with that compiler. I generally wait several stable revisions before using new features in the compiler. What that means is that you can always compile the source in the repo with the stable binary in the repo, and that compiler can compile itself again too, and both should behave identically. You can usually even go back a revision, but I've never had to do that.
There is a full interpreter built into the compiler as well, so if there is a bug in the stable compiler's codegen that is a showstopper, it can be fixed in the source and then the new source run in the interpreter of the old compiler in order to get a new stable binary. I've never had that happen, though.
The Bootstrappable Builds folks are working on this for every part of a modern Linux distro. The main approaches are alternative smaller implementations written in other languages (including interpreters) and also compiling with chains of older versions that were written in other languages. They are also working on a full bootstrap from ~512 bytes of machine code plus a ton of source all the way up to a full Linux distro. They have gotten quite far in that and are continually improving the situation everywhere.
--
Invoking methods with tuples. https://github.com/titzer/virgil/blob/master/doc/tutorial/Tu...
Just terrific. Similarly, I've wondered about the symmetry of maps (dictionaries) and named parameters, and the potential to invoke methods with maps.
--
"Functions are contravariant in their parameter type and covariant in their return type" https://github.com/titzer/virgil/blob/master/doc/tutorial/Va...
Terrific. This "hole" in language design has always frustrated me. As a noob, I've been wondering:
Do covariant return types resolve "the expression problem"? Thereby reducing the number of 'instanceof' type checks? https://en.wikipedia.org/wiki/Expression_problem
Do covariant return types moot the need for using double dispatch when implementing Visitor (for statically typed languages)? If so, that'd close at least one ergonomic gap between dynamic and static languages.
--
Whinging:
Virgil is practical. It uses modern techniques to address programmer's actual needs.
I love both functional and imperative programming. Separately. I do not want multiparadigm. I do not want metaprogramming in my bog standard data processing code. I do not want exquisite puzzle boxes (inspired by Haskell and ML).
For most of my code (done in anger), I want data centric, I want composition, I want static typing. To noob me, it appears Virgil is on the right path.
For just one example, Java jumped the shark. Specifically annotations, optionals, and lambdas. (I grudgingly accept the rationale behind type erasure for generics; it was a different time, when backward compatibility reigned supreme.)
We need features, often syntactic sugar, for the 98% of our daily work. String intrinsics (finally!), intrinsic regex expressions, intrinsic null-safe path expressions (not LINQ), tuples, multiple return values (destructuring), etc.
I want concision without magic.
In Java's defense, specifically, I love many of the JEPs of the last decade. Project Loom is a game changer. Switch expressions are great. Ditto values types and records. (There's more, but you get the idea.)
Also, project Zig embraces the practicality vibe. And shout out to D language.
Smalltalk the proverbial Object Oriented language already had first-class functions a.k.a lambdas (called BlockClosures in Smalltalk). So it's not like you should choose between OOP-classes OR lambdas. Having them both is better than having just one of them.
Ironically, Smalltalk syntax for lambdas/closures remains my favorite. No “trailing closure” hack that doesn’t scale and looks ambiguous (looking at you swift). No one line limitation (looking at you Python). No ambiguity between function/method bodies and closures (looking at all of you members of the curly brace Algol descendents). No “you have to reference all of the arguments or you can’t use it” (looking at you Elixir).
There were two capital ironies with Smalltalk’s free functions (closures). Smalltalk USED them. It may have been “objects all the way down”, bit they drank closures the whole way. Both the extensive class library and your own code. I coded and invoked more closures in Smalltalk for a given unit of functionality than I have in any other language to date.
What’s even more ironic to me is that closures/functions in most compiled languages are a semantic facade. You see them in your mental model, but the language/library doesn’t model them for you to interact with. In Smalltalk, you could send messages to closures, AND you could add your own.
I have always found it ironic that in Smalltalk, the “only objects ultimate”, I was actually exposed to closures/functional patterns more than I have been in many other systems that supposedly were more about just that.
Another innovation of Smalltalk which may not be familiar to everybody is that all control structures in Smalltalk are implemented by passing closure-objects as arguments to methods like ifTrue:ifFalse .
"… compiler optimises ifTrue:ifFalse: and friends using special opcodes, so as to avoid having to create closures…"
https://stackoverflow.com/questions/32662354/optimising-iftr...
It's like saying that functional languages implement tail-recursion as iteration in order to "get rid of" recursion.
But the programmer can still think in terms of recursive calls in their program, and can reason about the correctness of the program by assuming it works by recursion.
Except when they do!
When they expect to be able to change those methods, because they have been led to believe they are just methods.
g := [:y :z |
|f|
f := [:x | y * x].
f
].
(g value:3 value:2) value:4.
(g value:2 value:3) value:4.
https://wiki.c2.com/?SmalltalkBlocksAndClosuresSomehow, like always, progress is full of left turns, and we could have been much better if the IT world wasn't busy reinventing them.
My own style has evolved as I've learned over the years. I use classes less and make them smaller. I use enums and ADTs a lot more these days. But I rarely go whole-hog functional unless it is in tests.
From OO to FPGA: Fitting round objects into square hardware?: https://web.cs.ucla.edu/~palsberg/paper/oopsla10.pdf
Vertical Object Layout and Compression for Fixed Heaps: https://web.cs.ucla.edu/~palsberg/paper/cases07.pdf
Virgil: Objects on the Head of a Pin: https://escholarship.org/content/qt13r0q4fc/qt13r0q4fc.pdf
I found these with https://scholar.google.com/.