Vale's first prototype for immutable region borrowing
verdagon.dev
verdagon.dev
The very first impression of running 'valec' compiler is that it "panics" when no arguments are given. It literally writes so: "(panic)". I want to point out that "panic" is a very strong word and should be avoided in scenarios where a normal error handling takes place. Any kind of panic is always a sign of an uncontrolled situation, and if a program ever "panics" it leaves a bad taste in the mouth.
The next struggle was to get the basic help for the command-line parameters. But it's currently non-existent.
The next and final stuggle: trying a hello world sample. I copied the code from the website:
import stdlib.*;
exported func main() {
println("Hello world!");
}
and saved it to 'hello.vl' file. Then, I tried to build it: > valec hello.vl
However, no luck for me this time: Unknown subcommand, specify `build`, `run`, etc. Use `help` for more.
It looks like I should specify 'build' command. Let's try: > valec build hello.vl
Well, here is the result: Unrecognized input: hello.vl
(panic)
Hm. Let's try to get some help: > valec help
The result of the help command is: <nothing>
Not very helpful. At this point, I gave up. How does this thing work?Then again, their README doesn't really indiciate this and does say "Try Vale" so idk. But it seems very R&D/POC at this point.
(The inverse is "VC 1.0", which means pre-release but we had to meet a deadline, and "Programmer 1.0", which means it's the first stable version.)
I thought "Programmer 1.0" is more like +Infinity and never ever reached? ;)
Conventionally 1.0 is just the inflection point where things are featureful and stable enough that you start making guarantees about the stability of whatever interface/contract/API/whatever that matter to the consumer of whatever you are building.
The compiler is very rough around the edges right now. August-May was spent being 100% focused on prototyping regions, and you're experiencing the tech debt I accrued on the way there (including the lack of an integration test for the help system). I've been paying that debt down for the last 1-2 months and we're still not back up to where we were at the 0.2 release.
If you need any more help, let me know, or swing by the discord server where there are many helpful folks. Cheers!
...more predictable latency than tracing garbage collection.
...better performance and cache friendliness than reference counting.
...prototype and iterate more easily than with borrow checking.
Ok, you had my curiosity, but now you have my attention.Just started following your RSS feed: https://verdagon.dev/rss.xml
https://verdagon.dev/blog/generational-references
Statically eliminating memory operations does seem to be a win though.
Copying references is assumed to be more frequent than allocating and freeing, so this is a win.
(Or possibly the whole region could be discarded once it was running full, and the physical memory recycled at a new virtial address. There's plenty of virtual address space to burn through when not limited to 32 bits)
With 64-bit counters, that's never going to happen. Alloc/free costs more than a nanosecond, and there are lot more than one element that you will be allocating, but even if you somehow managed to reallocate the same object a billion times a second, it would take over 500 years to run out of indexes.
Practically speaking you could never allocate memory that fast, a memory allocation is going to be well over 1000ns on average.
Then there's the little matter of address space. Pointers on x64 are limited to 47 bits, meaning that if even you had a magical memory allocator with no book-keeping overhead, and all your allocations were 1 byte, you'd run out of pointers first. The actual virtual memory space is limited further on many operating systems, but you're still always going to be well short of 64bits.
Except when it is: https://threatpost.com/another-linux-kernel-bug-surfaces-all..., fix at https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...
If you have a multithreaded app doing a lot of communication, that's going to be a lot of cheap allocations happening very fast.
Reducing GC and allocation overhead results in more allocations being done, and pushback against ever-expanding allocation behavior is more of a challenge. Instead of ten other things being a higher priority than judicious data architecture, it's dozens or more.
Reference counts are not re-used when memory is re-used. And again, even if for some reason you had a global 64bit counter that you incremented on every allocation and never decremented, and you could somehow handle a billion allocations per second, you'd have 585 years before that counter overflowed back to 0. No computer or program can run for that long.
And even if there was a single counter, multithreading cannot make incrementing a counter faster. Two cores cannot write to the same cache line at the same time. Instead, cache lines need to bounce across cores when you write to them, and this takes such a long time that it turns the time it takes to roll over from centuries to millennia.
But I digress. Overhead like increments only matters on hot paths, which are very few. The Python + C stack for ML is a manifestation of this truth.
Having ergonomics of a “regular language” (affects all code) and the ability to optimize for performance (hot paths only) and stay in the same language is what I’m excited about.
https://github.com/sponsors/ValeLang
Let's use this post's time on the front page to help the project meet its $3,000/month goal.
I'd love to help Evan work on this full time (I'm a sponsor). A fast and safe language that's also fun to prototype with is worth supporting.
Patreon: Varies a bit more. Patreon takes 8%, unless they have been on the platform since before the 2019 change and are still on the 5% plan. And payment processing depends on size. Under $3 is 5% and $0.10 per transaction. Over is 2.9% and $0.30 per transaction. And more if PayPal or Venmo in not USD. [2]
So the split seems much better on GitHub. But the conditions are a bit different for using the platforms, and you can get perks on Patreon which you may not be able to get on GitHub. I can't remember who / which project but I believe I saw one that said something about a difference in taxes / VAT and not being able to give some of the perks on GitHub because of it. Cannot find it right now though.
[1] https://docs.github.com/en/sponsors/sponsoring-open-source-c... [2] https://support.patreon.com/hc/en-us/articles/11111747095181...
I think this might instead be MIR, mid-level IR, there’s a good blog post here: https://blog.rust-lang.org/2016/04/19/MIR.html
Cranelift is a compiler backend, mainly focused on JIT, but theoretically could replace LLVM, there’s an alternative backend being worked on but has limitations: https://github.com/bjorn3/rustc_codegen_cranelift
I don't use smart pointers since shared ownership is a bad concept.
The problem of memory management is largely trivial.
Google et al have been working on sanitisers etc because, even in well kept codebases with strict coding standards that are rigorously applied in reviews, memory bugs do actually creep in.
This applies regardless of programming language.
Of course the web people and their "frameworks" is just another demonstration of how bad relying on third party code is.
def f(x)
return ...
in Python or Java, those functions will work regardless of what other modules I import. (modulo monkey patching in Python, though you can defend against that)In C you don't have these guarantees -- foreign code can stomp on your code.
This is probably why C does not have an NPM-like culture of importing thousands of transitive dependencies -- because that style would just fall down.
Also a minor issue, but C doesn't have namespaces (for vars or macros), so it's a bit harder to compose foreign code.
Also compiler flags don't necessarily compose. Different pieces of C code make different portability assumptions, and different tradeoffs for speed.
That is just plain false, since a Python module can trivially be tainted by what you import before, and the Python environment is widely known for its dependency hell.
Meanwhile C modules, once compiled, can be fully isolated from what you link against them, depending on build and link settings.
Non-local bugs is just a matter of sharing state across a module boundary. Memory errors is just a very small subset of the possible bugs you can have in a program, and preventing them doesn't magically solve all the other more important bugs.
This is actually an important point. I think all codebases can (and should) be split into small, opinionated, privately owned sub-codebases. This is why developing large scale projects can work even in languages like C. After all this is what that whole 'modularity' thing is about ;)
(it also implies that external dependencies need to be managed the same way you handle internal dependencies, as soon as you use an external dependency you also need to be ready to take ownership of that dependency)
Arguably that's even a good idea in memory safe languages, it avoids tricky borrow checker issues, and also prevents the outside world to directly manipulate objects. Everything happens under control of the system.
Or you could give the message in fragments to the user, but that immediately becomes a very inconvenient API.
I think you mean you don't use shared smart pointers? Or do you avoid unique_ptr too?
Of course I know unique_ptr.
If you witness the endless amount of bugs, many security related, which stems from the idea that people can handle memory, I'd say it is far from a trivial problem.
If you witness any modern language, a common design principle is to eliminate memory management. Which argues it is far from a trivial problem.
There is a reason C++ still reigns supreme even though it was built in the 90s.
More manual memory management methods still have their place, because there are problems where you can't afford to use a garbage collector, or where it gets into the way.
C++ will be relevant for many years to come. It has way too much momentum as a language and too much software has been written in C++ to ignore it. I personally think Rust will eventually carve up a large part of its niche though, because I think it has a far better approach to managing memory.
If I understand clearly, this prevents use-after-free and double-free? Thus, the program can still fail on a memory access when the expected and actual generations don't match? In this regard, this seems less "safe" than reference counting, tracing garbage collector, or borrow checking?
How would it even stop use-after-free and double-free?
The "check" function accesses the allocation because it needs the generation number of the allocation. So basically, the reference needs to access the allocation to check if it can access the allocation. Right.
(That doesn't work of course, because if the allocation was freed, access to the allocation and so its generation number is undefined).
This seems obvious so maybe I am missing something big here?
Or something entirely different is meant or targeted here with "memory safety".
> so its generation number is undefined
With e.g. a random C compiler and a random malloc, that's true. But why couldn't the language and runtime cooperate to ensure it is defined?
For example deallocation can write a predictable value to that slot, which is never used as a legit generation index. The memory allocator can make sure that a memory address that ever contained a generation can never contain anything else than generation ids for the entire runtime of the program (e.g. by ensuring that for a given page, all objects are the same size and the allocations are aligned to that size). The language can make sure that nothing else can get written to such a memory address by enforcing bounds checks.
We created two newer approaches since then, which let any memory be reused for any purpose:
* Random generational references, where it's fine if generations overlap with other data.
* Side-table generations, which is slower but we keep the generations in a side-table. It's can be seen in old 0.1 versions as the "resilient-v2" mode, and I plan on resurrecting it for unrelated reasons.
The former will be the default, and the latter we'll be adding back in as an option. Hope that helps!
We have a future improvement planned here too: for unrelated reasons (to support generation pre-checking) the random generational references implementation will soon not even unmap any virtual address space, instead remapping it to a central page, so we won't even get any segmentation faults, just assertion failures.
void __check(GenerationalReference genRef) {
uint64_t currentGeneration = *(uint64_t*)((char*)genRef.alloc - 8);
assert(genRef.rememberedGeneration == currentGeneration);
}
So indeed you it allows you to check for a match, as long as the alloc pointer is valid. The alloc pointer is invalid after a free, because it maybe be in a region no longer accessible to the program (it was returned to the os by free's implementation) or it was given out as part of an other allocation, in which case it can hold arbitrary data.Also, the blog post talks about 'generational indices', not pointers. This seems to indicate that items of the same type (or at least same size) are grouped into arrays (and since it's an index anyway, the metadata could be stored in one or multiple separate arrays at the same index).
PS: I already linked it elsewhere, but here's how the same can be achieved without language/compiler support (at least it's the same general idea): https://floooh.github.io/2018/06/17/handles-vs-pointers.html
The big step forward by Vale is that the compiler can elide most of the 'dangling checks' on memory accesses, the method outlined in the blog post requires a few rules-of-thumb the coder must follow when using a pointer that's been looked up from a generational-index.
EDIT: oh, borrow checking.
...wouldn't that also be prevented by the generation-check even if there is no single-ownership? Because once the referenced item is destroyed (and thus bumping that "memory slot's" generation counter) that item reference becomes invalid because the generation no longer matches, so the next attempt to release the item with that same reference should also fail?
One nice property of generational-indices is that they can be shared without compromising memory safety. As soon as the item is destroyed, all shared references in the wild automatically become invalid. But I guess single-ownership still makes a lot of sense for thread-safety :)
Nothing in this article is relevant, but it still stays up, the only article in the blog.
1) The creator of it used a disposable GitHub account, launched the review/attack for the drama, then disappeared.
2) The only thing they ever published on their blog, was the hit job on V. No other reviews ever made.
3) Anything, which had any kind of substance, is already fixed[1].
4) A search of mawfig.github, shows how it is spammed on HN, and usually used for smearing.
[1]: https://github.com/vlang/v/issues/14803
Now we have Evan with Vale, and Adobe’s Software Technology Lab wit Val, it will be great searching for related stuff.
Anyone have an explanation of what is going on here? I'm finding the article impenetrable.
TL;DR: Vale is like a cleaner C++, and it uses generational references [0] which are similar in spirit to running with ASan [1] turned on. Generational references have a bit of overhead, but it can be removed by regions [2] or more specifically, immutable region borrowing [3]. This helps Vale achieve its goal of being a high-performance language while still remaining memory safe.
Hope that helps, happy to answer any other questions =)
[0] https://verdagon.dev/blog/generational-references
[1] https://github.com/google/sanitizers/wiki/AddressSanitizer
[3] https://verdagon.dev/blog/zero-cost-borrowing-regions-overvi...
[4] https://verdagon.dev/blog/zero-cost-borrowing-regions-part-1...
I'm still going through the guide but one thing I find curious is the module naming when building your application. It seems like you clone a library to disk and then "import" it as a command line argument with the name you choose. I'm trying to wrap my head around how dependencies would work if you have the following situation:
- parse library
- http library (requires parse) -> import parse=~/parse/src
- my_app (requires parse and http) -> import http=~/http/src parse_with_different_name=~/parse/src
Note how my_app uses a different name for the parse library than http.If the http library uses "parse" in the source code when referencing the module (import parse) and my application uses "parse_with_different_name" when referencing the module (import parse_with_different_name), does that mean to compile my app I would have the following...
valec build mymodule=~/my_app.vale parse=~/parse/src http=~/http/src parse_with_different_name=~/parse/src
Maybe I'm missing something and maybe it's too early to worry about things like this. Regardless I am loving this language and very excited about it.Edit: trying to fix my example list
Making C++ safe without borrow checking, reference counting, or tracing GC - https://news.ycombinator.com/item?id=36448759 - June 2023 (214 comments)
Memory safety without borrow checking, reference counting, or garbage collection - https://news.ycombinator.com/item?id=36351415 - June 2023 (93 comments)
What Vale taught me about linear types, borrowing, and memory safety - https://news.ycombinator.com/item?id=36156790 - June 2023 (12 comments)
How Memory safety approaches speed up and slow down development velocity - https://news.ycombinator.com/item?id=34410187 - Jan 2023 (137 comments)
The Vale Programming Language - https://news.ycombinator.com/item?id=31786487 - June 2022 (90 comments)
The Vale Programming Language - https://news.ycombinator.com/item?id=25160202 - Nov 2020 (171 comments)
The Next Steps for Single Ownership and RAII - https://news.ycombinator.com/item?id=23865674 - July 2020 (38 comments)
* I like mythical birds, so I wrote an article about memory safety and mythical birds: https://verdagon.dev/blog/myth-zero-overhead-memory-safety
* The Rosetta stone fascinates me, so I wrote an article on linear types and the Rosetta stone: https://verdagon.dev/blog/linear-types-borrowing
* I heard about a pigeon named G. I. Joe so I added a side note to about it on the C++ article at https://web.archive.org/web/20230629052606/https://verdagon....
* And now I had to find a way to spread the word of Brigadier Sir Nils Olav III, so I used a side note in this one. It's embarrassing but I was giggling with glee all day yesterday at the thought of putting that note in!
I suspect this is a curse that a lot of bloggers can empathize with, but they don't have the proper lack of professionalism that I do.
Once I had these little side notes, I figured I'd give some sort of prize to the first person who told me they saw them, which evolved into "comment somewhere mentioning it!" which I guess is a social hack? Maybe? I'll allow it!
I remember seeing Vale many years ago (~10). Back then, it was something revolving around Gnome project, but now it still has a pre-prelease version 0.2-alpha. This means that the project's progress is relatively slow, but the language is very interesting to me.
Update: I confused Vale with Vala! Vale is the new project, Vala is 10+ years old, but they have an intersection of syntax, goals and ideas. That's why I was misled by seeing a similar name. Val-vale is almost the same!
[1]: https://vala.dev/
I was slightly confused when I first read the title as well :)
https://floooh.github.io/2018/06/17/handles-vs-pointers.html
(disclaimer: I only wrote a blog post about it, that idea is much older and probably has been re-invented many times over since the first computers were built)
Essentially "non-owning weak references with spatial and temporal memory safety".
What is similar to ARC though is that moving that stuff into the language lets the compiler remove redundant handle-to-pointer conversions, similar to how with ARC the compiler can remove redundant refcounting operations.
If you store them inline with the program data for max speed(tm) you need to ensure that e.g. after 2 2kb chunks are deleted, you don't overwrite them with a 4kb chunk, because that would trample over a generation.
If you do keep the generations inline and rely on a statistical approach, you have to be very careful to never generate "common numbers" like 0 as a generation because then it's extremely likely there will be a collision.
It'd a hard problem and I'm quite curious how all the edge cases are handled.
iOS is getting a 'typed allocator' which seems to work similar:
https://security.apple.com/blog/towards-the-next-generation-...
[0]: https://docs.rs/generational-arena/latest/generational_arena...