That's still a clear statement, because it's trivial to tell if you used "unsafe" or not.
That's still a clear statement, because it's trivial to tell if you used "unsafe" or not.
FWIW, Fil-C’s guarantees are literally what you want. There’s no escape.
I do think the ideal kind of memory safe language either has no "unsafe", or has an "unsafe" feature that only needs to be used in super rare an obscure cases (Java is like that, sort of).
Fil-C has no "unsafe", so in that sense Fil-C is safer than Rust. You don't need an escape hatch if the memory safety guarantees are dialed in just right.
Someday, man
I am trying out a couple of new directions, e.g. generating more of the tiers from a more abstract description, constantly shrinking the amount of hand-written compiler/interpreter code. My hard requirement is the end result has to be pretty darn close to what I'd write by hand.
One thing I am thinking about now is how to make more use of the implementation language's (e.g. Virgil) compiler to be able to paste together machine code templates gotten from writing in the implementation language. Think copy-and-patch compilation, but as language primitive. E.g. "please emit an inlined copy of the machine code for this function (first-class ref to said function) into memory here, under this ABI".
You say that and then you describe exactly what I would have used as a solution: copy and patch.
Just have the checker check the templates that the baseline JIT is stitching together and then have a safe way to ask for the prechecked templates to be stitched together.
Bunch of details in getting that right obviously, but it doesn’t seem impossible.
> Just have the checker check the templates that the baseline JIT is stitching together
Sure, from my second paragraph, the templates it's using could be opaque things it got from requesting the static compiler generate a template from a first class function ref (at compile time), in which case the verification has already been done.
If I told you that I have a snippet of machine code that:
- obeys the ABI of your safe language (ie it has exactly the calling convention that safe language uses)
- corresponds exactly to a function body whose signature is T->U (or whatever, different safe languages have different function type syntax)
- obeys the language’s type system.
Then you could run an abstract interpreter to check that the machine code follows that type system. Simple example: given the above claims, if we further assume that the host language impl puts argument one into register 5, and the first argument’s type is “pointer to an array of bytes”, and we know that arrays have a 64-bit length prefixed to the start, then the abstract interpreter would just need to check that any deref of register 5 is preceded by a bounds check on whatever was loaded at offset -8 from register 5. And so on, for every possible thing you can do in the language.
Then the JIT would just have to make sure it puts checks in all of the places that the absint expects them. If the absint fails, then the machine code is rejected.
But Fil-C objects (at least for now?) only seem to allow one single capability type, and that capability grants unrestricted read/write access to the object’s bytes.
I wonder if one could build a handle system in Fil-C that would allow this to be extended. Or if a different variant of a Fil-C-like system could distinguish between pointers with different access levels to an object and could allow only the correct piece of trusted code to increase the permission of a pointer.
What I mean by that is: the memory safety issues of C are a total dumpster fire, while whether a number is even or not (and whether you can prove that) is maybe like icing on the dumpster fire. It just doesn’t matter by comparison.
So I want to decisively fix the memory safety issues and not lose focus.
C got its performance fame thanks to optimizing compilers that abuse UB semantics.
Microsoft team on .NET, especially the great Stephen Toub blog posts, has been showing off how much performance can be squizzed out of a managed language compiler toolchain when people actually care.
Also lets not forget Apple only moved away from Object Pascal due to an internal team doing MPW initially as kind of submarine project, due to their UNIX roots, and still their focus was C++, not C.
Toy systems, yes. Hey, I too, think CircuitPython is really neat. But I'm skeptical someone would base a PLC (or similar) on it.
The James Webb Space Telescope runs JavaScript, apparently [1].
[1]: https://www.theverge.com/2022/8/18/23206110/james-webb-space...
Fil-C doesn't necessarily have to run in production. It just needs to catch the bugs, e.g. by making it easy to fuzz C code compiled via Fil-C.
(This seems like one of those "throw the baby out with the bathwater" cases that people relitigate around Rust -- there's ample empirical evidence that building safe abstractions around unsafe primitives works well.)
Not sure the data is clean enough to draw meaningful conclusions because of confounding factors.
The biggest confounding factor is that Rust is relatively new, code written in it is even newer, and folks who research vulns may not have applied the same level of anger to Rust as to C.
That said, your point about "throwing the baby out with the bathwater" is well taken. I would expect that Rust has much fewer vulns than C/C++. My point is only that it's an unproven expectation.
I'm thinking of things like the Windows user- and kernel-mode font parsers; these have a pretty long and steady public history of exploitation that seems to have mostly stopped with the Rust rewrite 1-2 years ago. I don't think that's because vuln researches have stopped looking at them!
But yeah, I would like it if Google and Microsoft (among others) would put more hard data out there. I don't think of the Windows kernel teams as typically suffering from hype-driven development, so my abductive conclusion is that they have strong supporting data internally.
Edit: here's a hard data source from Google, showing that Rust has contributed to a marked decline in memory unsafety in Android[1].
[1]: https://security.googleblog.com/2022/12/memory-safe-language...
It's hard to say.
That's the whole point.
There's tons of trivial unsafe in the Rust ecosystem, and a little bit of nontrivial unsafe, because crates.io is full of libraries doing interesting things (high-performance data structures, synchronization primitives, FFI bindings, etc.) while providing a safe API, so you can do all of that without writing any unsafe yourself.
The point of Rust isn't that you can implement low-level data structures in safe code, but that you can use them without fear.
The operation is completely safe, as after the call the original mutable slice is no longer live, but the borrow checker won't let you write such a function yourself unless you tag it as unsafe, so that's what the implementation must do.
The same thing happens in the implementation of Vec: there is low-level code that is unsafe, used to provide safe abstractions.
Rust can split up an array without unsafe if you use Rc. That has overhead, but is it more than Fil-C in this situation?
That capability pointer is always stored and loaded using monotonic 64-bit accesses.
Therefore, in the worst case you'll get a pointer that is torn from its capability, and then you'll trap accessing that pointer. But you'll never get a corrupt capability.
If you don't want the pointer to tear from its capability, just use `_Atomic`, `volatile`, or `std::atomic`. Then Fil-C uses lock-free shenanigans to make sure that the capability and pointer travel together and don't tear from one another.
You seem to like (a), Linus Torvalds seems to like (b), and Rust targets (c).
The truth is: Rust’s approach to memory safety only works if you also have no data races. This makes it a strictly inferior approach to memory safety, since programs sometimes do have to race (sometimes it’s just the best solution). So, Rust has data race prevention not because it’s a good idea but because the whole language falls apart without it.
In terms of the design goals and evolution of the Rust language, this is exactly backwards. Rust was originally conceived as a garbage-collected, green-threaded language designed for concurrent programming -- think something similar to Go but with a stronger type system and no data races. The ownership model was created, first as foremost, not for memory safety but to prevent data races; and, more broadly, to help programmers reason about the correctness of their code.
Midway through development, as people started using the language & getting a feel for it, the language designers realized that the ownership model was powerful enough to express not just data race freedom but full memory safety without needing a GC (some writing from this stage in Rust's development: https://smallcultfollowing.com/babysteps/blog/2013/06/11/on-..., https://pcwalton.github.io/_posts/2013-06-02-removing-garbag...). So they removed the GC, and that decision positioned Rust where it is today: a nicer, memory-safe "C/C++ competitor", popular for systems code where the performance overhead or runtime complexity of a garbage collector is considered unacceptable.
Pfft, whatever. Rusters gonna rust, I guess.
But in this case, it's not just the memory safety I'm interested in, it's the data races. If we have multiple threads but can guarantee that any object either has only read-only references, or one mutable references and no readers, we don't have data race issues.