The Austral Programming Language
austral-lang.org
austral-lang.org
module body Fib is
function fib(n: Nat64): Nat64 is
if n < 2 then
else
end if;
end;
end module body.
We here all understand BNF, Ada, Modula, etc and parsing but imagine explaining to the first day student:
Why is there no "end function" like for the other contexts? When do I use a semicolon vs a period to close a context? You shouldn't need the "railroad diagram" to understand the syntax.Statements need an `end if`, `end for` etc. because it lets you find your way in nested code. The rationale for the syntax explains it a bit: https://austral-lang.org/spec/spec.html#rationale-syntax
FWIW I will probably get rid of the `module is ... end module.` bit because it adds unnecessary nesting.
Dare I say, the (important) bit of Lisp syntax fits in a sentence! Yeah they hide the complexity in the library instead...
Correct
> Why do you need a semicolon at the end of "end"?
Per the rationale[1], "The purpose of the semicolon is to provide redundancy, which aids both reading and parser error recovery." Also, "For many people, semicolons represent the distinction between an old and crusty language and a modern one, in which case the semicolon serves a function similar to the use of Comic Sans by the OpenBSD project."
[1]: https://austral-lang.org/spec/spec.html#rationale-syntax
From the link:
>> }
>> }
>> }
>> }
>> }
>> Which one of these corresponds to the second for loop? Unless we have an editor with folding support, we have to find the column where the second for loop begins, scroll down to the closing curly brace at that column position, and insert the code there. This is manual and error-prone.
I'm not fully convinced of that argument. Here's a devil's advocate take...
The keywords do indeed help when the corresponding nested control structures are off-screen, but if the code you are reading is not refactored to move the control structures into their own function so that the indentation doesn't get that much out of hand, you likely have bigger problems with the code than determining which `end` corresponds to which control structure.
IOW, having this sort of identification is moot: if it is needed, then the code itself is in such poor condition that it's likely not very readable anyway. In many cases it won't make a difference anyway (nested 'if's, for example - seeing multiple `end if` doesn't help) and the developer is still going to place a comment specifying which particular `if ()` is being ended.
Having the unadorned closing braces (`}`) leaves the developer one of three options:
1. Refactor that code just to be able to read it, or
2. As you point out, add in comments like `// end if`, etc, or
3. Leave it as it is.
If it's left as is, there's bigger problems in the code anyway.
With regard to 'no arithmetic precedence', I tried
printLn((1 + 2) + 3);
and printLn(1 + 2 + 3);
Sure enough, the first one compiles, but the second doesn't.Also, (n-1) is a parse error unless you put a space after the minus.
I got curious if recursion was properly handled, given it wasn't in the anti-features list, but no luck:
module body Foo is
function go(acc: Nat64, n: Nat64): Nat64 is
if n = 0 then
return acc;
else
return go(acc + n, n - 1);
end if;
end;
function main(): ExitCode is
printLn(go(0, 135000));
return ExitSuccess();
end;
end module body.
yields Segmentation fault (core dumped)How deep you recurse :D
Maybe it's worth saying "they're close enough to the same that parentheses should be optional", but I can definitely see the argument for just requiring them regardless.
Ideally I'd like stack overflow to be a clean abort rather than a stack overflow (just to make the error message more explicit) but I haven't got around to adding that.
Austral’s module system is inspired by those of Ada, Modula-2, and Standard ML,
with the restriction that there are no generic modules (as in Ada or Modula-3)
or functors (as in Standard ML or OCaml), that is: all modules are first-order.
Modules are given explicit names and are not tied to any particular file system
structure. Modules are split in two textual parts (effectively two files), a
module interface and a module body, with strict separation between the two. The
declarations in the module interface file are accessible from without, and the
declarations in the module body file are private.
Crucially, a module A that depends on a module B can be typechecked when the
compiler only has access to the interface file of module B. That is: modules
can be typechecked against each other before being implemented. This allows
system interfaces to be designed up-front, and implemented in parallel.
It was a mistake how C++, Java and other languages forgot to split interface declaration from implementation definition, IMHO. Good to see that Austral learned from Modula-2.Before I can form an opinion regarding Austral, though, I would need to see some larger programs implemented in it, for instance some low-level systems code, a generic data structure, some high-level business logic.
Huh? C++ is split into header files (interface) and cpp files (implementation)...
It does enable header-only libraries though.
It is assumed this would raise eyebrows from the user of this function. Furthermore if you were to take a “safe” function and replace it with a dodgy one in a later version, the function signature would change and users would need to update their code. So nothing quite so brazen would get past.
Of course if you are mixing in arbitrary assembly/machine code in your binary via linking that might make a syscall and that could potentially be unsafe.
On July 3, 1940, as part of Operation Catapult, Royal Air Force pilots bombed the ships of the French Navy stationed off Mers-el-Kébir to prevent them falling into the hands of the Third Reich.
This is Austral’s approach to error handling: scuttle the ship without delay.[1]: https://austral-lang.org/spec/spec.html
[3]: https://vale.dev/
`austral compile hello.aum --entrypoint=Hello:main --output=hello`
vs
`go build`
Etc etc.
What does `go build` do? What files does it implicitly rely on? I have no idea. But I have a pretty good idea of what that `austral` command is going to do without having read any documentation about it.
Essentially like `cargo` vs. `rustc`. I have a little prototype of the build system in Python but haven't pushed it up yet.
To me, having them separate forces you to keep things simple, because the build system can't communicate with the compiler except through compiler-provided interfaces.
Also, I think I like about languages like C, Rust, is that: if I wanted to, I could implement the build system without forking the compiler. In C specially because Make will print all the compiler invocations for you. It lets people build tooling that is not part of the compiler.
I think it's good from a simplicity perspective that language users can figure out what set of compiler invocations a build file "compiles down" to.
But, when you get big enough to where this sort of provided tool isn't enough, chances are tooling is now somebody's actual problem anyway, if you're a for-profit somebody's job is to look after the tooling, you can invest in learning a specialized tool or even writing one because that's a proportionate effort, it makes sense.
On the phone now so I only read the page on linear types, but will look at this closer when back at my desk.
In my own language I am considering destructors purely so that early returns are viable. I'd like to see if there is any alternative to destructors that aren't 'defer' or similar.
There's "no garbage collection" because Austral lets you have manual memory management without the danger, like Rust.
There are no destructors in the sense of special destructor functions which are called implicity at the end of scope, or when the stack unwinds. Rather, you have to call the destructors yourself, explicitly, and if you forget the compiler will complain.
This sounds verbose until you start paying attention to all the mistakes you make all the time that involve, in some way, forgetting to use a value. The language makes it impossible to forget to do something.
The big downside is the verbosity of covering every branch of your code with your explicit close calls unless another mechanism is provided.
And it doesn't seem like succinctness is a top priority for this language.
Some questions from my side:
- 1. As far as I understand, there are multiple models for a linear type system. Which one does Austral implement? Is it verified to be correct?
- 2. Since there is a static checker: What are the limits on 1. expressivity and 2. scalability?
- 3. What is the intended memory model (pointer and synchronisation)?
My suspicion is that this language would be unusable for real time applications, which is ironically what it would be most useful for.
Interesting. I've wondered about this when making an expression parser. Obviously it makes parsing way easier and mistaken precedence is often a cause of bugs (especially in C where some of the operator precedence is plain wrong). But on the other hand that's got to be quite annoying surely?
What is more annoying to me is looking at an expression that mixes arithmetic and logical/comparison operators and mentally trying to recover the parentheses. Because precedence is not just PEMDAS: it involves all binary operators in the language, including logical and bitwise ones.
On the other hand, subtyping interacts very, very badly with both type inference (it almost instantly becomes undecideable, in practice as well as in theory) and with typeclasses (again, decidability issues).
Adding it open a can of worms. Your type-inference/checking must account for "similar but..." types, you need to follow hierarchies/graphs/trees,etc, you need a way to "re-import" code that "belongs to the super type" (and probably recheck it?), it not always mesh well with other features (or make it harder).
Aside: One of the big reasons to make a bit list of "NO" is to avoid the temptation of "add some sugar to make this easier" but without fully understanding the consequences until becomes later.
1. Suppose Civic is a subtype of Car
2. Suppose you have a List<Civic>
3. Suppose a method asks for a List<Car>
4. You can "clearly" provide that original list of civics because every single thing in the list is a Car, and it's a list of those things, so it adheres to what the method seems to want -- maybe the method computes average cylinder count or something.
5. Everything we just described is fine and dandy so long as the list itself is immutable (the individual cars could still be mutable safely), but running some of the mutable list methods will cause runtime crashes and segfaults. E.g., if you append a toyota camry to the list then you've somehow snuck a camry into the original List<Civic>.
In that example, some of the methods would be safe if you interpreted a List<Civic> as a List<Car> (like grabbing the car at a particular index and finding that the particular car is a civic), but others require the subtyping relationship to go the other direction (e.g., if you interpreted the List<Civic> as a List<BlueCivic> and appended a BlueCivic then the invariants expected by `append` work in both cases).
That sort of thing just scratches the tip of the iceberg, and as a rule of thumb all your type systems are unsound, _especially_ if they involve subtyping. Things you would hope would be caught at compile-time are punted off to scary runtime heisenbugs that might not be detected for ages. The type system is helpful at reducing errors but woefully incomplete even for the things it's supposed to catch.
and the main problem in those os are that there are 21 languages and thats before you even start talking about build systems, config files, and query languages and serialization. how would adding a new language to fix memory leaks make anything better? i see much worse problems here like you still have to build an sql query out of a string. the whole rust having this was just to make c people more comfortable to using a post 90s language (i.e memory safe)
And stack overflow is a memory allocation failure, so why is the discrepancy? I.e. for the language focusing on correctness this is an unfortunate omission.
On the other hand none of popular or semi-popular system languages allows to explicitly control stack consumption. Zig has some ideas, but I am not sure if those will be implemented.
The programmer can always statically ensure that the program doesn't experience a trapped overflow, and that the stack size is not exceeded. All the information to do that is available when the programmer runs the compiler.
But there is no way to prevent a memory allocation failure when using `calloc`, since the information required to do that is not available when the programmer writes the code. In fact, when running on POSIX systems it's not even possible to check _in advance_ at runtime whether a memory allocation will succeed.
This is why Austral's allocateBuffer(count: Index): Address[T] returns an Address[T], which you have to explicitly null-check at run-time (and the type system ensures that you can't forget to do this).
Of course, on some non-POSIX systems such as seL4, the programmer can know at compile-time that memory allocations (untypedRetype) will not fail. When you use Austral on such systems, you don't have to use calloc/allocateBuffer at all.
It's a pain, and the type system rejecting any recursion is certainly simpler, but that's not a strict requirement.
For a system language I would like to see that when the compiler cannot infer the bound on the stack size or when that static bound exceeds some static limit, a function call is treated as fallible.
This is fantastic. Clearly, bureaucracy is what we needed all along for memory safety.
What's the logic behind this? It's nice to have a way to pair resource usage with disposal. Or does the linear type system allows not to forget about the resources which have to be closed.
The 'no destructor' rule follows the 'no hidden flow' design, similar to Zig. Some doesn't like it, but personally I prefer it.
Linear types enable manual memory management without memory leaks, use-after-free, double free errors, garbage collection, or any runtime overhead in either time or space other than having an allocator available. More generally, it enables us to manage any resource (file handles, socket handles, etc.) that has a lifecycle without letting us forget to dispose of the resource (e.g. leaving a file handle open), dispose of it twice, or use it after disposal (e.g. reading from a closed socket), all without runtime overhead.
The Austral tutorial's chapter on linear types (https://austral-lang.org/tutorial/linear-types) explains how this works in a fairly clear way.
Interesting that this was mentioned. Are there languages that are not simple enough to be understood by a single person?
> Rust does something very practical: it says what properties it will enforce, but doesn’t say how. That is, you know references have to uphold the law of exclusivity, but how the compiler achieves this is subject to change. So the borrow checker is allowed to evolve over time, in the direction of accepting more programs and becoming more ergonomic while retaining safety.
> The upside is that most of the time you can use Rust without thinking about lifetimes or the borrow chchecker. The downside is that the borrow checker is hard to spec (it’s basically “whatever rustc does now”) which makes is hard to have multiple implementations of Rust. Some people argue that’s fine or a good thing because multiple implementations waste effort. This argument has merit, but I think languages being specification-defined is a good thing from a stability perspective. It’s what they call a tradeoff.
[1]: https://borretti.me/article/type-systems-memory-safety#rust
I’m not looking for APL levels of terseness, but I also don’t want to have my code mistaken for an essay filled with what amounts to scaffolding.
Ooh! What about [0]? Gambit Scheme has what they call `six-script` which is essentially a refactoring of Scheme grammar to fit into an Algol-esque syntax.
[0] http://gambitscheme.org/latest/manual/#GSI, see section 2.6.1
In my view, the job of the C committee should be to eliminate undefined behaviour and any remaining specification ambiguities (or errors) and otherwise leave the damned language alone.
If you want a new language every 3 years, you can always use C++.
Personally I prefer more terse syntax, but I’ve heard some people defend the style of more verbose languages like Pascal and Java when writing large programs.
By most people’s account, the actual coding part of programming isn’t really a majority of their time. Domain research, speccing, debugging, testing, planning, etc. the microscopic savings in key presses are just really such a strange thing to even debate about when you think about it.
Major tersness pushes also tend to cause “symbol soup” and, personally, I find symbol soup very hard to digest.
It's about readability
Do you remember which languages they looked at?
* "Shorter [C#] identifier names take longer to comprehend" (2019) https://link.springer.com/article/10.1007/s10664-018-9621-x
In this paper, we investigate the effect of different identifier naming styles (single letters, abbreviations, and words) on program comprehension. We conducted an experimental study with 72 professional C# developers who had to locate defects in source code snippets. ... We found that word identifiers led to a 19% increase in speed to find defects compared to meaningless single letters and abbreviations, but we did not find a difference between letters and abbreviations.
* "[Java] Identifier length and limited programmer memory" (2009) https://www.sciencedirect.com/science/article/pii/S016764230...
names used in existing production code are long enough to crowd programmer’s short-term memory. This provides evidence that software engineers need to consider shorter, more concise names. As the study considers individual names extracted from production code, it tends to underestimate the demand on memory because there is no need to remember context as well.
* "Evaluation of Rust code verbosity, understandability and complexity" (2021) https://www.ncbi.nlm.nih.gov/pmc/articles/PMC7959618/
[1]Me
`foo(){}` is just as clear as `fn foo(){}`
There is no need to distinguish between functions, modules classes, lambdas, or whatever.
Hence, no need to distinguish their start or ending with keywords, as long as you can determine their scope.
Brackets determine scope, and unlike indentation or words, that is all they are used for.
https://www.youtube.com/watch?v=5kj5ApnhPAE
And the quote to go along with it:
"I'm always delighted by the light touch and stillness of early programming languages. Not much text; a lot gets done. Old programs read like quiet conversations between a well-spoken research worker and a well-studied mechanical colleague, not as a debate with a compiler. Who'd have guessed sophistication bought such noise?"
Except for COBOL, of course, which is one of the oldest.
It made me realize that people naturally want to be more verbose when they are less comfortable with the concepts and less verbose when they are very familiar with the concepts. It also made me realize it is totally subjective and there may not be a right answer, that it depends on someone's background and familiarities.
If so, what?
Work on Austral was initiated by Borretti in 2021, but its earnest development really only commenced in January this year. So we're talking about a language that's, for most intents and purposes, less than a year old. Even if the January release was stable, there would not have been time for anyone to develop and deploy significant projects in it. But it is not stable: it still under construction, with certain inherent instabilities. Notably, recent updates changed the borrow syntax and FFI pragmas. Similarly, the surrounding infrastructure (compiler and standard library) is far from production-ready yet.
I built an Austral interface for the seL4 Core Platform (https://github.com/zaklogician/libmantle/tree/main), and ran some Austral apps on hardware. The purpose of this was experimental (we wanted to know how Austral can help us design fail-safe seL4 Core Platform APIs, and it was a success), but probably among the closest anybody came to "production": I wrote more Austral code than currently included in the Austral standard library.
It revealed several bugs, including a typo in Standard.Buffer which leads to the invariant check for the Buffer type always failing, and issues related to the compiler's handling of large unsigned literals.
Austral shows great promise, and has already provided us with valuable insights about linear API design, and exciting glimpses into its future capabilities. Using Austral in production code, however, might require a few more years of maturation and stabilization. So if you want something that people use for writing production code, check back in a couple of years.