APL Compiler Based on Tail (Typed Array Intermediate Language)
github.com
github.com
I designed VVM to have extremely fast dispatch. It already exists on top of a vector-aware runtime, which means that a single dispatch works over an entire array. But grouped aggregations require multiple dispatches, one per unique key in a Dataframe. For example:
from trades select volume = sum(size) by symbol
I wanted to hit the runtime repeatedly with as little overhead as possible. So VVM has no type look-up, multiple operands per instruction, and a cache-efficient IR. The sum() operand for the above example is invoked directly in a loop almost as fast as hand-written C++.VVM has its own assembly language [2]. I have blog post that explains some of the design choices [3].
[1] https://github.com/empirical-soft/empirical-lang/tree/master...
[2] https://github.com/empirical-soft/empirical-lang/tree/master...
[3] https://www.empirical-soft.com/2020/09/03/a-tour-of-the-vect...
This is because, in high-performance interpreters, the arrays have element types that are tracked automatically, providing many of the advantages of C-like compiled languages at a fixed cost per array operation. In fact, these element types can be more accurate by dynamically adjusting to the data: with a table of unknown size the C programmer must declare indices as a 64-bit type, while APL could detect at runtime that only 16 bits are needed. APL also benefits from the higher-level notation, in that it's easier for a human to write a fast implementation of a particular operation than a compiler that can generally optimize code including that pattern (with TAIL this might not be so bad). I discussed the Replicate (like filter) operation in my talk The Interpretive Advantage[0]. Another page I wrote where you can read about array performance is [1].
APL compilers are needed in order to run code on a GPU (but if you have code that's well-suited to a GPU, it's going to be pretty fast on a SIMD CPU too). And for ML and scientific applications that work with floating-point types, the drawbacks of ahead-of-time compilation are much smaller. I think for truly general-purpose computing the best approach will be to unify interpretation and compilation, combining small groups of primitives and deciding how to run them as data becomes available. Both interpreted and compiled approaches are important in reaching this peak.
My more detailed notes are at https://mlochbaum.github.io/BQN/implementation/compile/dynam....
I suggest connecting directly to DIKU for academic matters.
They are doing other interesting work on array based programming:
Futhark, a related data-parallel functional array language that compiles to GPU code and easily integrates into Python code is another interesting project from there.
Dyalog has some introductory documentation on how to set up the mappings. https://www.dyalog.com/apl-font-keyboard.htm
If you want to type something very small, you could also go to tryapl.org and use the on-screen keyboard and copy and paste.
If you just want to play around with it, you might try https://tryapl.org/