The Vale Programming Language
vale.dev
vale.dev
For those interested in the PL space, here are some of the shenanigans we've been up to in Vale:
* We added "Higher RAII", a form of linear typing that allows destructors to have return types and parameters. [0]
* We just yesterday finished the first prototype of deterministic replayability, which will allow us to capture all inputs to a program (including all network data, user input, file input, etc, and eventually, even thread orderings). No article yet, but our docs on it are pretty approachable. [1]
* Last month, we finished the first milestone of "Fearless FFI", which lets us call into our C code without fearing it corrupting our Vale data. Later on, we'll be adding automatic sandboxing, either via subprocesses (using IPC for FFI) or wasm2c (which should be a lot faster).
We've also got some interesting plans for concurrency. We've found a possible way to make memory-safe and data-race-safe structured concurrency even easier, [2] which I hope to get prototyped before the year's end.
As a fan of C++ and Rust, I'm most excited about having the RAII and flexibility of C++ with the memory safety and data-race safety of Rust, while being easier than either of them. We seem to be succeeding, hopefully that continues!
I also want to emphasize that Vale is a work in progress, and we endeavor to be very clear on the site about which parts are implemented and which parts aren't.
We couldn't have gotten this far this without our sponsors' support, so big thanks to all of them! [3]
[0] https://verdagon.dev/blog/higher-raii-7drl
[1] https://github.com/ValeLang/Vale/blob/master/docs/PerfectRep...
[2] https://verdagon.dev/blog/seamless-fearless-structured-concu...
IIUC it's basically a form of memory tagging in userspace, but designed to be safe (64 bits, very low risk of collision) and fast (language is designed to avoid tagging whenever possible: owning refs don't need them; inline objects reuse the gen of the parent; static analysis).
I buy that this can be almost as safe as reference counting or GC, and as easy to use, but faster. But not sure I understand how inline objects work, and how effective static analysis will be is important.
Comparisons to other languages (C++/Rust/JS) are here: https://vale.dev/comparisons
Generational references were less than half the overhead of RC as of last benchmark. [0]
Re inline data: A piece of inline data will (most often) share the generation of its containing object. A generational reference can include an offset to the parent's generation, which we can use for generation checks.
Inline data should make it even faster, because then we can control our objects' layouts and use the cache more efficiently, and not be chasing pointers all the time. I think that's the biggest advantage of generational references over RC and GC.
We're also building regions into the type system, which can be used to statically temporarily freeze areas of memory (similar to the Rust borrow checker) to eliminate generation checks. [1]
Another very experimental aspect we're prototyping is "hybrid generational memory" which can temporarily lock an object (similar to a RefCell) to elide a lot more generation checks [2] but it's too early to promise it will work.
[0] https://verdagon.dev/blog/generational-references
I see, thanks for the clarification. Then interior pointers will be larger than normal pointers, like all generational pointers? While the interior object itself doesn't have a generation, avoiding that overhead. Makes sense I think.
All the larger pointers do make me worry about increased memory overhead, though (kind of the reverse of the x32 ABI which has half-sized pointers; I think 5-8% is the quoted perf difference there). Do you have benchmarks of memory overhead - looks like the link has throughput?
Regardless, it sounds like the other features you mention should help with both forms of overhead, so it will be interesting to see how much.
Cool project btw!
The increased memory overhead would seem to be a problem, but in the programs I work on at least (games, web servers, compilers), the vast majority of references in a program are owning references which aren't fat pointers. Non-owning references are very short-lived and on the stack, so the memory overhead shouldn be pretty minor in practice.
Also, if we want to save a little more space, we can reduce the generations to 32 bits (we're leaning this way in fact, as it opens the door to some interesting other features).
1. This technique prevents undefined behaviour, by halting whenever an invalid dereference is detected. While Rc is a garbage collection algorithm. So you are comparing apples and oranges, no?
2. Rc is a bad gc algorithm. How does this compare to a good quality, generational, incremental gc?
One central issue is that anything touching concurrency and Go is a tangle of nastiness with implicit effects. But apparently Go is "natively concurrent" and everyone simply prays that the race detector is good enough.
Parallelism does need some design to avoid making a mess (minimize shared mutable state and lock what you can’t minimize and you’re fine). Note also that Rust and other safer languages don’t help much because most shared mutable state is remote and accessed by multiple application processes (e.g., an object in an S3 bucket, a file in an NFS volume, etc). With Rust or with Go, you need to test for these kinds of bugs, but with Go you’ll be able to start your testing sooner (with Rust you’d still be writing the app code).
I'm aware of the potential for data races to lead to data corruption in Go, although as an exploitable issue it seems a bit theoretical. There are also soundness bugs in Rust from time to time (as for example with the whole Pin saga).
https://news.ycombinator.com/item?id=31703732
> I'm aware of the potential for data races to lead to data corruption in Go, although as an exploitable issue it seems a bit theoretical.
Go's language design says it doesn't care about this problem, Rust's language says it eliminates this problem. If you see these as basically the same because you also don't care about the problem, it seems like your original claim (that they're both safe) was simply wrong.
> There are also soundness bugs in Rust from time to time
To be absolutely clear here: Data races in Go are not "bugs". Go is specifically designed not to even be safe if you have a data race which touches compound types. Undefined behaviour, you lose, game over, the language designers have nothing further to say on the matter. If you raise a ticket saying "I had this race and now everything is on fire" in Golang it will get a WONTFIX or whatever the equivalent is.
In short, both languages let you write memory unsafe code if you want to, but both discourage it and make it easy not to do so in most cases. Rust discourages it in a more bondage and discipline kind of a way. But it still falls to the programmer to verify the unsafe kernel of their application to their satisfaction. That is, Rust provides a 'here be dragons' warning by forcing the use of an 'unsafe' block, but the language itself doesn't offer any assurances about the correctness of an application's kernel of unsafe code. Modern async Rust doesn't uniformly discourage the use of idioms that make use of unsafe code. See for example the discussion of stack pinning here: https://rust-lang.github.io/async-book/04_pinning/01_chapter... ("A mistake that is easy to make is forgetting to shadow the original variable since you could drop the Pin and move the data ... (which violates the Pin contract).")
I'd advise against inferring that people "don't care" about these problems, etc. etc. This kind of personal stuff just makes it harder to focus on the technical details.
I wish JavaScript had the memory layout guarantees that Go has, that would make it a lot easier to work around memory-based bottlenecks.
And no, the fact that I can use low-levels bindings in NodeJS is not at all similar.
The worst thing is that this travesty of a language will win and lead to decades of stagnation.
How does this work with threading (use after check)? Or are allocations always limited to a single thread?
> This will safely halt the program
So it’s not statically memory-safe. That’s not hugely attractive I have to say.
Ref-counting with mutable data (which this seems to be replacing?) works the same way in Rust. You need to take an Rc/Arc of RefCell, and calling borrow/borrow_mut on RefCell can fail at runtime.
Looks very interesting, and the blog has very nice writeups, I found the article on mutable/constant variables syntax[1] particularly clever. Tho I'm unconvinced of the final decision it's certainly a breath of fresh air.
[0] https://en.m.wikipedia.org/wiki/Vala_(programming_language)
"Vale is powerful enough to use, and it feels really good to finally use Higher RAII in a real-life program. It's an incredibly versatile and valuable pattern, one that I hope Vale will bring into the mainstream!"
Bring [Higher RAII] to the mainstream. It's a sandbox to experiment with a composition of ideas in the hope that some of these ideas influence the 'mainstream' including, perhaps, D.
The number of new, fast, safe (for some value of safe) native languages that are appearing is amazing. I think this is down to powerful tools for creating native languages. We're in a new era of language development; the yawning gulf between C/C++ and the world of managed/scripting languages is being rapidly filled. This shouldn't be discouraged with "why not just make my preferred thing better?" Eventually it will.
https://verdagon.dev/blog/on-removing-let-let-mut
Originally there was let and let mut like Rust. But that was changed to simple assignments, followed by the compiler checking for subsequent reassignment with set.
While I can see that might make for cleaner code while still allowing the compiler to get on with its job, surely one of the benefits of explicitly immutability is to stop the coder accidentally letting off foot guns. If simply setting a variable after declaration makes it mutable, it feels like it could open the door to data in a code block being allowed to change behind the scenes when the assumption in that block is that it is fixed.
I'd be interested to know how this works in every day practice and possible unintended effects.
Let's say it is a mental burden. I'd rather have the burden when writing the code over when reading the code. (https://news.ycombinator.com/item?id=31823991)
The solution in the link adds to cognitive complexity. Imagine you had bad code (Worst case scenario; which always happens). A function with 400+ lines of code. You would have to scan the entire function just to figure out if a variable is mutable then you would have to force yourself to remember _all_ the variables that are mutable when evaluating behavior.
The Vale Programming Language - https://news.ycombinator.com/item?id=25160202 - Nov 2020 (171 comments)
The Next Steps for Single Ownership and RAII - https://news.ycombinator.com/item?id=23865674 - July 2020 (38 comments)
Otherwise, it looks very promising!
[1] https://github.com/project-everest/vale [2] https://en.m.wikipedia.org/wiki/Vala_(programming_language)
But it could have some implications for software stability (compile time being obviously better) and productivity (run time being obviously better).
It does in practice. With a compile time check the issue simply cannot happen, and there is no need to deal with it in the code or at system level.
With a runtime check the issue may happen, but will be detected when it does. Still, the choice then is either to crash or to deal with it with some runtime recovery action. Either way has some cost: the system around the executable must deal with more crashes, or the code gets more complex (and sometimes more brittle, even if the intention is the opposite).
Only the compile time check makes an issue really go away. This is to me the attraction of strong typing and any form of compilation time check. There's a price to pay too in accepting the related constraints: type checking must pass. So there's a cost here too. For complex or sensitive applications I personally much prefer this upfront cost.
But it's definitely not the only way: Erlang has runtime type checks, but a very good runtime error handling framework (the reference?), and it works well. Still, most environment relying on runtime checks are not at this level.
The issue that goes away in both cases is the issue of unsafe memory access. There are disadvantages to runtime checks, as you mention, but there is no difference in the level of safety achieved.
With a compile time check, there's a development cost but absolutely no runtime consequence.
With a dynamic check, sure the memory access is detected and blocked. But if you stop there the application crashes, which may be completely unacceptable. In an embedded system such a crash may be as bad as the incorrect access itself.
More generally, the issue is not so much the incorrect access than its possible adverse consequences. With a compile time check, there are no runtime consequences. With a runtime check, there are still runtime consequences: either the impact of an application crash, or the extra error handling code and behavior to deal with the detected wrong access and mitigate it at runtime. Whether such consequences are acceptable or not depends on the context, but it's there and do make a difference with a compile time or static analysis check.
I don't agree on this point. An incorrect access on an embedded system has the potential to cause all kinds of horribly subtle bugs involving memory corruption. A simple crash is generally much better.
I think there's a lot of possibility in runtime assertions that fail as fast as possible, ie the instant the program enters an invalid state rather than the later time you get to the part of the code the assertion is written in. Don't know if there's any systems like that, it's just something I thought of.
In practice, the borrow checker has to reject a lot of perfectly fine patterns, such as observers, dependency injection (the pattern, not the frameworks), delegates, backreferences, many forms of RAII [0], graphs, etc. Sometimes, the workarounds lead to more complexity, less flexibility and decoupling, and more refactoring which would be unnecessary in other languages. This is likely why GUI is difficult with the borrow checker, and why one has to bring in frameworks to compensate.
The borrow checker can prevent certain kinds of logic bugs at compile-time, which might mean a release lets 9 bugs into production instead of 10 or 11 (comparing to a safe language with a strong type system). However, that tradeoff might not be worth it, depending on the domain. Flexibility and decoupling can be more important, at least in the domains I've worked in (roguelike games, web servers, apps) and the size of the program. There are domains where it's better to add more complexity to detect even more bugs at compile-time, that's where I'd choose Rust (or perhaps GC'd FP languages, which prevent even more bugs). Just my two cents!
Note that this is only a problem if a programmer is a bit too religious with borrow checking; in practice Rust offers reference counting which can nicely avoids these problems.
One of Vale's principles is to move checks up to compile time, but prefer not to when it causes too many architectural problems or "infectious leaky abstractions" so to speak. This is also why Vale will be using coroutines (similar to Go) instead of async/await.
It's also why we're adding a region-based borrow checker, which is opt-in and doesn't impose constraints on its callers. [1] If we do it right, it should give Vale a lot of the performance benefits of Rust's borrow checker, but without the complexity and architectural constraints.
Also, IIUC this is not actually UB in Vale: it's a guaranteed error.
Yes, it's not as good as a static guarantee. But there are tradeoffs where it makes sense. Again, RefCell in Rust does the same - it's a useful technique.
C adds a third state undefined. When you reach this state you know nothing about what is happening in the program.
Now Vale adds a third state called "memory error" which is just a refined error and well defined. This means that if you have handled the error case, you already handled the memory error case even if not in a satisfactory way.
What is strange to me is that you consider the former okay and the latter undefined behaviour in the application logic when it just means that an additional exit state has been added which a highly defined behaviour.
[0] https://verdagon.dev/blog/hybrid-generational-memory#afterwo...
This is an interesting claim since JavaScript runtimes are written in C++. I suppose we do trust them, but only after wrapping them in all the other layers of safety we can get our hands on.
It's not a problem for your own programs because you're not trying to hack yourself, but if you're an aggressive enough developer you will find bugs in your compiler. In fact, this is a great reason to have full test coverage.
If one defines safety as a lack of UB or vulnerabilities, then I'd say Vale's a bit safer than other languages like Rust.
I say this because:
1. There are no unsafe blocks in Vale.
2. Vale memory is decoupled from native memory, it only passes messages and handles between, and uses a different stack. [0] This prevents accidental bugs in unsafe code from corrupting data from safe code. We just finished our proof-of-concept of this last month!
3. Building on that, we could automatically sandbox any native code for a module in theory, with either subprocesses or wasm2c. [0] (Note we've only just started on this part.)
I'd say that's safer than Rust, where unsafe code can undermine the code around it.
I'm particularly excited about how this might make it safer to use dependencies; #1 and #2 protect from accidental corruption, and #3 could help protect against malicious corruption and certain kinds of supply chain attacks. We're even tossing around a potential #4 to add permissions, so dependencies must be whitelisted to be able to access network, files, etc.
We're also thinking about relaxing the above 3 restrictions on a per-module basis if we can do it in a way that doesn't compromise the safety of the ecosystem (note that these relaxations are not implemented yet):
* Instead of full sandboxing (3), we can rely on the decoupling (2), if we trust a dependency's intentions.
* Have operators to skip generation checks for extra speed. [1] These operators would be by default ignored for dependencies. It's unclear if this will help, as the combination of regions [2] and HGM [3] might combine to eliminate the vast majority of generation checks. We'll see!
* Perhaps add a keyword for blocks, to ignore all generation checks within (maybe `unsafe`?). This would also be ignored by default for dependencies.
Rust must always allow unsafe blocks, because libraries often believe it's necessary for their performance, or sometimes for just working around the borrow checker. Vale would default to disallowing it (and sandbox FFI), and it would be much more noticeable (and suspicious) if a dependency said it wouldn't work without unsafe; they can't just sneak it in there like they can in other languages. These adjustments are still under consideration (thoughts are welcome!) but as of today, there's no unsafety in any Vale code.
To summarize, Vale has stronger protections against memory unsafety and UB, so I'd say it's a bit safer. Though, if one wants to prevent all bugs at compile time, then one shouldn't be looking at Rust or Vale which often panic, but instead look at languages like Pony (which doesn't even have panics, hence its amazing uptime) or proof languages like Coq.
Hope that helps!
[0] https://verdagon.dev/blog/next-fearless-ffi (still a draft, read generously!)
[1] https://vale.dev/guide/unsafe
It seems to me that generational reference isn't replacing reference counting as it is not deciding when to free the object? Perhaps this is more like a weak reference or something?
Wondering how does this deal with multithread though. For example, the object A might be dropped in thread 1 (increments the generation number in thread 1) and accessed in thread 2 (reads the generation number). How do they guarantee an error in this case without using atomics? (although this might not be a valid question because you can't really say A is dropped before thread 2 accessing A without some sort of synchronization...)
Does deference means accessing a field of object ? Does alias/dealias mean creating/destroying another variable reference ?
If answer of above is true, isn't programs deference more than alias/dealias ? i mean, in:
var alias = shared_ptr_of_some_object; for (int i = 0; i < 1000; ++i) { alias.foo += bar_fn(alias.baz) }
is just creating 1 alias but defrences 1k times and above looks like very typical code.
That statement is referring to our sample program, a roguelike game that we use for benchmarking. As for why this is the case, I'm not sure! I suspect its because a lot of classes will alias/dealias objects without ever dereferencing them, such as List<T>, HashMap<T>, etc.
That's just untrue right? You would have more room to optimize writing assembly directly?
Maybe it's a nitpick, but it gives me a little bit of pause to see a statement like this on a language website's page, since it seems to indicate a bit of a weak or sloppy understanding of computer science fundamentals.
OTOH Fortran (and, I suppose, something like APL / K / Q) can generate faster code than C++ because Fortran guarantees the absence of aliasing. In C or C++, aliasing precludes certain kinds of optimizations, and proving the absence of aliasing at a particular spot may be too time-consuming (if tractable at all) for the compiler to try.
1. Vale: it's worth it, fine, in Spanish
2. Vale: an area of lowland between hills or mountains, in English
Might be an apt connotation for a language that tries to make code less error-prone and less verbose.
In my experience it's generally easy that gets dropped, reading the site I don't think there was anything there that made me think this is easier to learn than other mainstream programming languages so I think that's the case here as well.
We also just introduced variadic generics in version 0.2, which is the first step towards a unified IFunction<R, ...> interface.
Long term, we'll have syntactic sugar for it: func(...)R
However, we can always change our minds later and make them optional, so it seems wise to just require them for now and revisit later.
Litterally anything is better than dragging conventions from the 1960s forward; from a time where ASCII was all that was available.
Use Whitespace, Try unicode glyphs for different syntactic functions, try to approximate natural language, TRY NEW THINGS FFS.
> And even ; to end the line.
Why do people even hate ; ? It gives you better error messages and it's no different from using . to end a sentence in spoken languages.
> Use Whitespace
When you make whitespace important the only thing you gain is that it forces people to format their code (see Python) but it also often leads to idiotic decisions (e.g. Nim not allowing tabs which are simply superior to spaces, there is literally not one reason to ever indent with spaces). C++ and co have the right idea, the only thing whitespace should do is separate tokens then people can format the code as they see fit.
> Try unicode glyphs for different syntactic functions
Now sure Unicode stuff can make things more readable in some cases but the big problem of Unicode is how you input those things. Your keyboard is basically ASCII, that leaves you with some workarounds like ALT codes or Julia's LaTeX conversion stuff. And those things are not supported on every platform/editor. Not to mention that for some things you need to install a new font or whatever.
> try to approximate natural language
What would you like to see? I don't see what you'd gain from that except mostly making things more verbose. Spoken languages are quite different from programming languages. As soon as you have a large block of instructions (aka a program) you naturally resort to some kind of structuring (e.g. making a list that someone else has to check from the top) and that will pretty much look like pseudo code already. Sure you'd do something like `list.add(foo)` instead of "Add foo to the list" but that's just because the former is much easier to parse and encode into rules for the computer.