How Safe Is Zig?
scattered-thoughts.net
scattered-thoughts.net
> null pointer dereference ... [C] none; [Zig] runtime; [Rust] runtime
Assuming this is talking normal "safe" Rust, I think I disagree with this. The Rust analogue of a pointer in safe code is a reference, not a Rust pointer, and these can't null at all. You could use an Option<> of a reference, and Rust will internally use null to represent the None (empty) case, but an attept to use the option without checking for None will result in an error at compile time, not runtime. Yes you could convert that into a runtime error, but if it was an error condition for that variable to be None then (depending on the context) you could choose not to use an Option at all and then it would be a compile-time error at the call site to attempt to put None into it.
I don't think I understand what is meant by "type confusion". Surely this would also cause compile-time errors? Even C++ would give compile time errors for this unless you use a cast! (C, unlike C++, lets you implicitly convert from void* to any other pointer type so you don't need a cast to get pointer confusion.) Could someone think of an example of what might be meant here, and how it would cause a runtime error?
So I agree with your disagreement.
This is checked at runtime in ReleaseSafe.
jamie@machine:~$ cat test.zig
pub fn main() void {
var x: usize = 3;
_ = @intToPtr(*u32, x - 3).*;
}
jamie@machine:~$ zig run test.zig -OReleaseSafe
cast causes pointer to be null
Aborted (core dumped)
(The x-3 is because if you use a literal 0 it gets caught at compile time)> Java-style wrapping integers should never be the default, this is arguably even worse than C and C++’s UB-on-overflow which at least permits an implementation to trap.
Either way there are explicit methods for doing wrapped, checked or saturating operations in every mode.
I imagine throw-on-overflow is slower than wrap-on-overflow.
C# can be configured to throw an exception when an int is overflowed, [0] but this behaviour isn't the default and is rarely used (typically it uses wrap-on-overflow). I imagine it might have a significant performance impact, but I'm not sure.
In a language like SPARK Ada, intended for formal verification, you can insist upon a rigorous proof that unintended overflow can never occur. That isn't an option for Safe Rust, at least not without significant breakthroughs in tooling.
[0] https://docs.microsoft.com/en-us/dotnet/csharp/language-refe...
The push for safe languages is motivated by pragmatism, not theoretical purity.
It depends on the CPU: MIPS has(had if you consider MIPS dead) some integer operations which trapped in case of integer overflow AFAIK the 'trapping operations' were as fast as the non-trapping operations (when no overflow occurred of course).
Unfortunately even though RISC-V is MIPS successor, it doesn't have those trapping operations :-(
I must admit I don't understand why no modern CPU ISA has these trapping instructions: explicit checks have an 'instruction cache cost', so users won't enable them which reduces security..
An obvious example: does this code result in a divide-by-zero? (I'll use C syntax.)
int myInt = INT_MAX;
++myInt;
int myOtherInt = 1000 / myInt;
If signed overflow is permitted to result in myInt holding zero, then we have a divide-by-zero. Not the kind of thing that should be left up to the particular platform.The behaviour of your Java code does not change when you move it from a 32-bit x86 machine to a 64-bit ARM machine. That's part of the appeal of Java. The same should be true of Safe Rust.
To put that another way: Safe Rust is remarkable because of its ambition: to be a truly safe language, while also having excellent real-world performance. It seems to be succeeding in doing both, without trading off on performance (Java, Go, C#) or safety (C++, and even Ada). If it starts compromising on either dimension, it becomes 'just another language'.
Integer overflow is a "program error." This case is handled by either "default" or "enabled" overflow checking:
* If checks are "enabled", then overflow must panic
* If checks are "default", then you'll get two's compliment wrapping
For implementations, if debug_assertions are enabled, then so must overflow checking be, unless the user specifically requests otherwise.
According to these rules, rustc today has "enabled" checking when debug_assertions is on, or when the user requests it via a flag. Otherwise, it leaves it to "default." If these checks ever become cheap enough, rustc may move to "enabled" in all cases by default. We'll see if that ever happens.
Integer division in Java never results in undefined behaviour. Divide-by-zero results in an exception, as does (Integer.MIN_VALUE / -1). In C/C++, both of those operations result in undefined behaviour. Modern JVMs have to generate some additional instructions to implement this [0], but I wonder if the real-world performance penalty is that substantial.
Another example is reading an uninitialized variable, which is undefined behaviour in C/C++. Does this footgun really improve performance, with modern compilers? I don't have a solid answer here but I suspect not.
> IMHO 'integer overflow is UB' should be scrapped
The C committee is opposed to radical change, they never want to step on the toes of exotic compilers and exotic hardware architectures, so I doubt they'll ever change it. I think it would be more realistic to ask the major compiler vendors to commit to never doing anything unsafe on signed integer overflow. I believe GCC has an 'opt-in' flag for this. I wonder what the performance cost is, if any. Perhaps it breaks some optimisations, but as you say, other fast languages like Rust seem to get by fine.
Still, I agree, I prefer consistent semantics between dev/prod as much as possible. Especially since there's methods to have checked/wrapping/etc arithmetic, so I can always go to those if I want the other behavior
Thankfully, this behavior is a flag which you can personally configure to be consistent across dev/prod: https://doc.rust-lang.org/cargo/reference/profiles.html
It does make sense when interpreting wrapping checks in dev as debug assertions
More like:
> "integer overflow checks are a painful/unacceptable performance degradation for some use-cases, but we still want to cough over-/under-flow bugs during testing"
Anyway luckily you can just enable integer overflow checks in release builds, which is not a uncommon setup in use-cases like server code.
(I only have limited C experience, and only for hobby projects)
So even if you knew your specific C implementation handled overflow properly there is no way to check the flag afterwards anyway.
That said, most of the time you end up counting objects in the current address space. If you assume that there can exist no more than `SIZE_MAX` objects in memory, you can avoid many overflow checks.
With regard to C compilers, there are a few cases where a compiler performs optimizations because it assumes signed integer overflow cannot happen. This is bad but typically, compiled C behaves like the underlying platform which means signed integers wrap around. With GCC, you can enforce this behavior and make signed integer overflow a defined operation with `-fwrapv`. You can also compile your code with UBSan to get runtime checks during testing. UBSan can also check for unsigned integer overflow which is defined behavior in C. So with modern C compilers, the situation is basically the same as with Rust, Zig or other safer languages.
I would say that's very rare, only for special cases or defensive programming. Usually you either know/assert the operation will not overflow because the inputs are bounded, or you use wider types (e.g. use a int32_t when adding two int16_t).
GCC says:
> This option generates traps for signed overflow on addition, subtraction, multiplication operations.
In my experience with C (which is biased towards some specific use cases), most numbers that are likely to overflow are things like sizes and counts, which cannot meaningfully be negative anyway, so you may as well use unsigned integers for them. Other cases really do require signed integers, but for most arithmetic operations you can 'just' convert to unsigned before doing the arithmetic and then convert the result back to signed.
(Some may disagree. For instance, the Google C++ style guide [1] specifically says not to "use unsigned types to say a number will never be negative", because they want the undefined overflow behavior of signed types, in order to allow the compiler to diagnose bugs and to avoid "imped[ing] optimization". I think this is mostly nonsense; the drawbacks far outweigh the benefits, and tools for detecting overflow like UBSan can be told to check unsigned overflow as well.)
That said, even if you avoid the UB cases, checking for overflow correctly is hard; I've found many security vulnerabilities caused by missing or incorrect overflow checks. __builtin_add_overflow and friends are very nice if you have them, though unergonomic. I wish a more ergonomic version were standardized as part of the language.
Yes, I really disagree. unsigned integers mean one thing, which is "modular arithmetic". Unless you are in the very uncommon case of actually needing modular arithmetic, for instance, when implementing a crypto or hash algorithm, you want normal integers. As soon as you have anything that may have any chance of introducing a substraction somewhere, unsigned will cause bugs.
I don't know how many times I had to debug broken code such as
for(int i = 0; i < some_size - 1; i++) { ... }
because some_size was unsigned.If you really want a "number that cannot be negative", you don't wan't some_size - 1 to silently give you UINT_MAX, you want a type that will give you a compile-time or at worst run-time error.
Could you supply a non-Google link to example code? Perhaps in some git repository or a code snippet on a non-Google snippet website?
It's however zero-overhead wrt a+b plus the manual error checking you'd have to do if you wanted to be correct
Parts of postgres just heavily document the allowed range and expected domain of such functions (effectively encoding a type system into the comments and relying on a human to enforce it).
I've seen people drop down into assembly to prevent UB. Note that checking the overflow flag doesn't necessarily suffice; the problem happens during compilation -- e.g. if you write `x = (x-MIN_INT)+MAX_INT;` the compiler might be able to reason that the only value for x which wouldn't have UB is MIN_INT. Consequently the result must be MAX_INT, and the compiler can inline that constant anywhere else it's used without ever issuing the instructions so that you could check an overflow flag if you wanted to. If you do those calculations in assembly then you have a lot more freedom in that regard.
Some projects do manually check every arithmetic call (or more commonly they'll lean on the pre-processor to do that kind of busywork for them), or at least they'll do so out of some small, core kernel which is more heavily vetted and can't afford the overhead.
Given the foot-notes, I think they are referring to unions - probably writing to one union member but reading through another (e.g. uni.intVariant = 19; float a = uni.floatVariant).
I don't personally know how Rust and Zig handle this,I believe it is UB in C.
EDIT: Having looked into this more I think I was getting confused with C++ (where the only supported way to type pun is to use memcpy or similar, which despite the name might not have to actually copy the bytes at runtime). Here [2] is a StackOverflow answer (/discussion) on the matter.
[1] https://www.yodaiken.com/2018/06/07/torvalds-on-aliasing/
[2] https://stackoverflow.com/questions/25664848/unions-and-type...
Does the Linux kernel still only support compilation using gcc? That’s about the only situation in which the standard could be considered “not important”. I wonder how many other C projects are in the same situation and only support the use of a single compiler?
There was an effort to build the kernel with Clang [1], I'm not sure about its current status. It helps that Clang tracks GCC's features quite closely, including some of its idiosyncracies.
> If the member used to read the contents of a union object is not the same as the member last used to store a value in the object, the appropriate part of the object representation of the value is reinterpreted as an object representation in the new type as described in 6.2.6 (a process sometimes called "type punning"). This might be a trap representation.
This verbiage has existed in a footnote of the standard since a defect report was filed against C99: http://www.open-std.org/jtc1/sc22/wg14/www/docs/dr_283.htm. This happens to be one of the few instances where Linus is right about the C standard, but without knowing why ;)
Thanks, you could be right. In that case, I think it ought to be listed as a compile-time rather than a runtime error too (just as I already argued for null pointer dereference).
It depends a bit on use of course: if you have a function that takes an enum and considers all cases using pattern matching, then the compiler will stop you from accessing the wrong member within a given case. If you pass an enum to a function that expects one particular case to be active, then yes you can convert the compile time error into a run time one, but this is your own choice, not something that is naturally a run time error.
Two random Rust projects:
https://github.com/tokio-rs/tokio:
$ grep -irn 'unwrap' | wc -l
1398
https://github.com/denoland/deno: $ grep -irn 'unwrap' | wc -l
2050
I dunno, but all Rust projects I have encountered had unwraps everywhere.The uses of unwrap in the projects the GP cited are overwhelmingly in tests and examples, which are not expected to handle errors in the same way as application code. The remainder are mainly lock poisoning unwrapping, which is a completely endorsed idiom.
But for options it's trickier. Unless the surrounding function is structured just right, the equivalent to Typescript's and Kotlin's ? operator is .map()/.and_then(), and that's pretty ugly. .unwrap() is easier.
Try blocks might help.
Additionally that metric of yours is faulty.
A) Looks at tests/examples
B) Probably counts 'unwrap_or' which is fine.
> Heavily discouraged by who?
I'd ask to see however encouraged it?
$ find . -name '*.rs' | grep -v -e test -e example | xargs -n1 sed '/mod test/q' | grep -v '^\s*//' | grep -cF -e '.unwrap(' -e '.expect('
224
I see a lot of unwrapping of locks. IIRC that only fails if another thread crashed while holding the lock, in which case crashing the current thread is often unobjectionable.In rust an unwrap of a None is defined and memory safe.
In C++, you'll be lucky if you throw a nullptr exception. It's undefined behaviour, and it's often exploited by compilers for optimization. Here's a super simple example [0] showing how the compiler makes assumptions and generates a very unexpected result.
My understanding is that Zig doesn't allow pointers to be null in the first place (regardless of release mode) unless you're 1) manually creating a pointer from an integer or 2) interfacing with C, and in both of those cases all bets are already off anyway (as they would be in Rust). The only "supported" options outside of that would be a non-null pointer or None.
Well I did link to a blog post that explains this, so I feel like it's sort of on you at this point. But I guess the short answer is that in Rust you know exactly what 'unwrap' will do on a None, and in C/C++ you can't.
> can you imagine your long running application (e.g a server) stopping in production..
Yes. It happens all the time, and in fact it is inevitable. Far better than the program misbehaving, which is what undefined behavior leads to. In fact if you're building a serious, production service, you might want to skip panicking altogether and just kill the process.
So however bad it might be for your production server to crash, imagine how much worse it might be for it to appear to continue working while it actually corrupts memory and database entries, or an attacker uses it as an exploit to read other users' information, or whatever else. At least when it crashes you can detect that with your health monitoring and potentially start it back up (if the same problem doesn't cause it to immediately crash again in a loop, of course).
[1] A secondary difference is that it is possible to catch Rust panics in an analogous way to catching C++ exceptions. This isn't encouraged and might not even work (if they've been turned into calls to abort() at compile time). The fact that they're not undefined behaviour is the big one.
This wording still suggests that it is a null pointer error reference, which is technically u.b., but in practice on any modern operating system a segmentation fault, but that is not what happens in Rust.
Rather, an optional reference cannot be dereferenced at all ere it be converted to an actual reference, typically with appropriate guards for the None-case, it is true that there is a way to convert an optional value of type `T` to type `T` without handling the None-case, but that is simply a trivial library function that handles the None-case by panicking the thread. The None-case must technically always be handled in some way in safe code.
The distinction is rather meaningful, I would say, as it is not the null pointer dereference stage where the error occurs, but the panicking of the thread of attempting to convert a None-case, which not only happens before that point, but provides for far cleaner debug information.
In safe Rust, one simply cannot dereference an optional reference, one can only coerce it to a reference, with the understanding that this operation panics the thread in the None-case. — an optional reference in Rust is not a “reference” at all as it is in other languages with special support for it.
Accessing a memory location with one type as if it's another. In C or C++ it's usually because you accessed a union without checking a tag somewhere else.
One classic way this happens is in an interpreter where you have a big enum for all the different possible data types:
https://github.com/MaterializeInc/materialize/blob/1b9b3cfab...
And then in various builtin functions you expect particular types:
https://github.com/MaterializeInc/materialize/blob/main/src/...
In rust and zig this code will produce a runtime error if you screw up, but in c it's easy to forget to check the tag and then you get UB. Similarly unwrap is checked in rust and zig but the equivalent in c - dereferencing a pointer that you are pretty sure is not null - is not.
Of course rust and zig both have support for c-style unions too, but they're not the first thing people reach for.
Fundamentally in the language, mismatching type is a compile time error. In many situations it's feasible to exhaustively match and deal with all possible cases of an enum (which then causes a compile time error if you try to reference one of the other inner types from within the "wrong" case). Where one particular case is expected, you can often ensure through the type system that this case is the only possibility by using that individual type directly rather than passing that enum around. In some situations its not feasible or it is feasible but the extra faff isn't worth the reward - but even then I don't think it justifies saying that the type confusion is detected in runtime in a comparison of languages.
As pointed out elsewhere, somewhat surprisingly, this is not UB in C, though it is in C++. In C it is merely unspecified behavior, but perfectly safe if intended.
One thing I found where Zig is currently worse than C compilers: returning a pointer to a stack variable generates a warning in "modern" C and C++ compilers, while Zig lets this slip through. I hope that "trivial" things like this will be fixed on the way to 1.0
Ensuring safety from important classes of bugs with sound guarantees is one way to help write correct programs, and both Zig and Rust use it; ensuring safety from from important classes of bugs with sound guarantees based on runtime checks is another way, and both Zig and Rust do it, too; a simple language that's easy to understand and analyse is another very important way to help write correct programs, and both Zig and Rust try to be simpler than their predecessors; making it easy to write tests and run them frequently is another way to get more correct programs that both languages try to employ. Both languages drastically differ in the use of those techniques from either C or C++ because they are both languages that put a very strong emphasise writing correct programs, but they also differ a lot from each other in how they balance those techniques.
It is impossible to tell without careful empirical research which helps write correct programs more than the other, and it is also possible that different people find it easier to write correct programs in either Zig or Rust. Rust certainly provides stronger guarantees that prevent temporal memory bugs than Zig, so let's assume Rust programs will contain zero, and a Zig program will contain more than zero but much less than C or C++. But that delta is insufficient to determine that Rust's balance of techniques reduces more bugs overall.
Also, Zig already has decent checks for use-after-free, and they'll get better, and not having uninitialised memory is also very easy to do (and verify) in Zig, despite there not being any checks. Even if, like other runtime checks, it is turned off in production, it still helps catch errors in that category.
While the thought as merits, the empirical evidence we have indicates that yes it is possible to achieve good software with faulty languages with strong rules and tooling but it is certainly not as straightforward as you make it seem.
In the end though, for the Zig case I agree that the jury is still out. But if I was a betting man, my money would not be on it, even though personnaly prefer Zig.
Whether static checking vs faster iteration time is more important depends entirely on the context, but rust isn't going to help you when you accidentally did front-face culling instead of back-face culling.
Even C++ can have relatively fast build times, depending on how everything is structured, and the use of binary libraries.
It is a matter of tooling, as an anecdote all my UWP C++ applications compile faster than most of my Rust experiments.
I never claimed it is straightforward; it is anything but. As a practitioner and advocate of formal methods and verification, I've been following research in software correctness for many years (and have written much about it, e.g. https://pron.github.io/posts/correctness-and-complexity), I've come to realise how complex the problem is, and there's more we don't know than we know, and even the things we know are problems, we don't know what the best solution is, because often solutions carry with them more problems.
Nonetheless, there are certain principles. We know that we can eliminate certain bugs with compile time guarantees; we also know that code reviews catch many (many!) bugs, and so making them easier helps. But what if these two are in opposition? It's not easy to tell which wins in which circumstances.
> In the end though, for the Zig case I agree that the jury is still out.
True, but the jury is still out on Rust, too. In fact, for most languages. However, there is no clear argument that we should assume, a priori, that Rust results in more correct programs than Zig. Many such arguments in the past have failed to yield positive empirical results (e.g. https://youtu.be/ePCpq0AMyVk). In fact, given empirical research, the safest bet is to assume the null hypothesis -- that there is no difference. Out of an abundance of caution, I'll assume that languages whose designers place a strong emphasis on correctness might achieve it more easily than languages whose designers put no emphasis on it at all, but Zig and Rust are in the same category here. Both are designed with correctness as a primary goal. But as their design and means of achieving correctness is so different, I think it's impossible to make an educated guess as to which of them, if any, yields more correctness more easily.
If we want some bottom line, it is this: software correctness is so complex, and solutions are often so non-obvious (i.e. many work in theory but not in practice), that we cannot say anything with certainty until we have actual empirical results, and even then we need to be careful not to be careful not to extrapolate from one study to other circumstances with different conditions (i.e. that TypeScript seems to have fewer bugs than JavaScript does not seem to extrapolate to the general claim that typing always reduces bugs compared to no typing in the same amount or at all, when other languages are concerned).
Wow, thanks for that link. I only made it through the first part for the moment but it is an incredible read. You clearly thought about this more deeply and carefully than I did.
Edit: I'm not entirely sure how that came across so I want to explicitly say that this is not a dry ironic statement (communication is hard, and I am a poor writer).
I understand the argument, but I'm not sure on what basis you consider that Rust's type system harms code review. Do you have specific examples in mind? (And because the discussion is about Zig, this is a pretty strange argument to make, because Zig's ubiquitous usage of metaprogramming is in fact a hindrance to code review).
Calling Zig's comptime "metaprogramming" is a little misleading when compared to other low-level languages. It is used for the same purpose as metaprogramming in other low-level languages (like macros in C++ and Rust, or templates in C++), but doesn't have any quoting mechanism [1] and doesn't operate at any "higher-level." In fact, Zig's semantics would be unchanged if comptime were executed at runtime. It is more similar to meaprogramming in dynamic language with reflection, with the benefit that related "runtime" errors are actually reported at compile-time. So comptime doesn't increase Zig's complexity. It can be thought of as a pure optimisation.
[1]: Unlike metaprogramming in Rust or C++, Zig's comptime is referentially transparent, i.e. if two terms, x and y, have the same meaning, then, unlike in C++ or Rust, one cannot write a unit e in Zig, such that e(x) and e(y) have different meanings. So the metaprogramming features in C++/Rust are trickier than Zig's.
You said that already[1], this is unsubstantiated and you declined to answer to my rebuttal.
> So comptime doesn't increase Zig's complexity. It can be thought of as a pure optimisation.
I'll grant you that it doesn't increase Zig's implementation complexity and also have a smaller learning-curve cost than other mecanisms. But when reading a piece of Zig code, you constantly have to wonder at which time the given code is gonna run. And there's much, much, more comptime in use in any piece of Zig code, than you'll uncounter macros in Rust or C++. So yes, it adds its share of friction when reading Zig code.
How would you propose to measure the concept of "programming language complexity"? One metric could be "how difficult is it to write programs that do not contain certain classes of bugs"? By that metric, C is indeed incredibly complex. An alternate metric might be "how long does it take the average developer to learn the language well enough to write reasonably effective programs"?
In the absence of formal studies we just have to go by our intuition. Personally, I kinda hate the "I'm not smart enough to write C, so I write Haskell/Rust" argument. It comes across as incredibly condescending to me. What I can tell you from my experience is that I spent a month trying to learn Rust on nights and weekends, and by the end of that was able to write some extremely simple programs with a lot of effort. On the other hand I was making nontrivial contributions to Zig itself within a week of learning the language. So to me, Rust is much more complex than Zig.
> Complexity characterises the behaviour of a system or model whose components interact in multiple ways and follow local rules, meaning there is no reasonable higher instruction to define the various possible interactions.
This is in fact the most antithetical possible description of Rust, which, thanks to its strong type system and compile-time rules, keep the interactions between different components or features as clear and specified as possible.
Yes Rust is hard to learn, but learning curve and complexity are orthogonal concerns.
Sorry, didn't see your response. I can answer it in two ways, subjective and objective. The subjective is "I know it when I see it," which roughly corresponds to the difficulty in determining what an unfamiliar piece of code does as well as how many language rules I need to know to figure that out. The objective one is literally language complexity, i.e. the computational complexity of determining whether a string belongs is in the language or not (i.e. whether or not it is well-formed).[1]
> you constantly have to wonder at which time the given code is gonna run
You really don't. The semantics of Zig are the same as those of Zig', which would be the language that runs comptime at runtime. The whole point of comptime is that as far as semantics -- not performance -- is concerned, you do not have to care when code would run.
[1]: There's a complex theoretical caveat here, because I believe both Zig and Rust are undecidable. So we can exclude degenerate cases from Rust, and look at the complexity of Zig' , the language I introduce in the second paragraph, which is semantically the same as Zig.
Anyway, complexity can come from many factors:
- feature bloat: C++ is way more complex now than it was in 1990, because features where added on top of features. In that regard, the older a language gets, the more complex it becomes. C++ is the most cited example, but I think PHP is even worse in that regard: it's probably the one and only most feature bloated PL ever, probably because there is not even a standardization committee to add frictions to the feature additions process. By that metric, Rust is slowly becoming more complex every year, like every other language (but the growth of its complexity isn't particularly concerning compared to others, Go for instance has recently been on a much steeper track).
- platform fragmentation: when Internet Explorer was still a thing, JavaScript development was made incredibly complex by the huge implementations differences between browsers. Code that worked somewhere failed somewhere else more often than not, and you had to keep work-around for old versions or IE for years. IE is mostly dead, Safari is less shitty every year, and google killed Android Browser and replaced it with Chrome, so it's a much smaller issue than before, but problems remain.
- cultural factors: Haskellers love for obscure mathematical terms or the fetishism of OOP's design patterns in Java in the late-90 and 2000 are good examples of culturally-induced complexity.
- ecosystem churn: JavaScript between 2013 and 2018 or something, with new framework or libraries or tools replacing the old ones every six months, before getting replaced themselves in the following month was a massive source of complexity, fortunately it seems to have settled a bit and the churn rate is lower than before. In Rust's early days, when many useful features were still unstable and feature-gated in the nightly version of the compiler, this phenomenon also existed (though at a much smaller scale). By that metric, Rust's complexity decreased quite a bit since 1.0, as many libraries have been adopted as de facto standard way of solving a bunch of problems (a few domains remain prone to this though, like error handling helpers, and ECS for game engines apparently) and Rust is now roughly in the same situation as most languages.
- counter-intuitive semantics: c.f. pre-ES6 JavaScript, how `this` and `var` bindings worked, which was simply the opposite of what people wanted in 95% of the cases.
- obscure control flow: `with` statement in non “strict mode” JavaScript, languages relying on a lot of `goto`, or even languages with exceptions.
- too much responsibility: manual memory management in C (or Zig for that matter) which we now have significant evidence after half a century that no human is able to do it consistently right of the time.
- poor interactions between features: see C++, how modern features interact poorly with older (more C-like) ones.
Rust is less complex than many mainstream languages on a least one of these dimensions, and less complex than JavaScript on most of these…
> The objective one is literally language complexity, i.e. the computational complexity of determining whether a string belongs is in the language or not (i.e. whether or not it is well-formed).[1]
This is a stupid metric, because it confuses implementation complexity with user-facing complexity (brainfuck wins this benchmark, yet good luck building anything with it). But from a theoretical perspective, this is a fun one because there's not only one but two classes of indecidability involved:
First, with most language with type polymorphism, it is undecidable to know whether a given program will successfully compile. But there's also a second level: when a language has Undefined Behaviors, a program compiling successfully isn't enough: it can still be invalid, and whether or not it is valid is also undecidable. C is not in the former situation but is in the later, C++ and Zig are in both, safe Rust is in the first only, but unsafe Rust is also in both. So in that regard, safe Rust is strictly less complex than Zig, but the whole Rust is equivalent.
> You really don't. The semantics of Zig are the same as those of Zig', which would be the language that runs comptime at runtime. The whole point of comptime is that as far as semantics -- not performance -- is concerned, you do not have to care when code would run.
This argument is pretty similar to the Rust point of “when you get used to it, ownership doesn't adds any cognitive burden”, maybe when gaining enough familiarity with Zig you can gloss over it without hassle, but I'm clearly not in this case yet so you really better not assume that it's gonna be straightforward and instantaneous for everybody, it is not.
I don't think so, because I don't think I'm claiming what you think I am.
> Anyway, complexity can come from many factors:
I completely agree, but I'm only talking about language complexity, in the strict syntactic, linguistic sense. I am not saying that all things considered, Rust makes maintaining programs harder than other languages -- nobody knows that until we have some empirical study -- but linguistic complexity is one very prominent property of Rust, as is, say, the memory safety of safe Rust, which, similarly, does not mean that Rust programs are overall safer than those written in, say, Zig, when all things are considered, because correctness also has many contributors, not just sound syntactic guarantees. There, too, only empirical study can settle the issue.
But you can't have it both ways, focusing on one specific piece when it comes to correctness yet insist on looking at the full picture only when it comes to complexity. All you can say is that, linguistically/syntactically, Rust offers some sound guarantees re memory safety and that it is complex, and that overall, both subjects are complex, involve many aspects, and require empirical study to make any definitive claim about.
> This is a stupid metric, because it confuses implementation complexity with user-facing complexity
It is obviously not stupid because it is commonly and usefully employed in computer science. But as with any precise definitions, it focuses on some aspects and not others. It captures the intrinsic difficulty of answering a question about a program. You are correct that it does not take into account human ergonomics and psychological aspects, but it is one more useful metric, even if not comprehensive.
> This argument is pretty similar to the Rust point of “when you get used to it, ownership doesn't adds any cognitive burden”,
Absolutely not (and, BTW, I was not referring just to ownership and lifetime when I spoke of Rust's complexity). It is a very precise and well defined property of Zig. The semantics of a Zig program, i.e. what it does in terms of what action it computes, is completely independent of comptime. It is not an ergonomic or psychological argument. comptime does not change the meaning of anything, and not only do you not need to figure out what happens at compile time and what happens at runtime -- unless you want to reason about efficiency -- but that knowledge contributes nothing. It's a meaningless distinction when it comes to semantics. It's a very powerful, well thought-out, theoretical and practical aspect of Zig's design.
Then again, on the strict syntactic sense, Rust is even less complex than C, because of the “most vexing parse” issue. If you wanted a rigorous analysis of the syntactic complexity, you could attempt to measure how difficult it is to write a lexer and a parser for every popular languages, and see how Rust performs. But given that the language grammar has been designed with parsing complexity in mind and have benefited from the hindsight of others before it, you'd be terribly disappointed.
From this discussion, and many of your previous interventions on this forum, it's pretty clear, even though the reason isn't, that you have developed a resentment towards Rust and you can't help bashing it.
Rust isn't a silver bullet, it has a fairly tough learning curve and as it tries to push the frontier of system programming language forward, it will take a few decisions that will ultimately be regarded as dead ends, and I have no doubt that future languages will avoids these pitfalls and provide improvement over the state of the art.
In the meantime, spreading your hate with unsubstantiated judgements like “Rust is one of the 5 most complex programming language ever” or “Rust harm code reviews” isn't really constructive for anyone.
Zig is a cool motorbike, Rust is a SUV. Arguing that your bike can indeed be safer than a SUV because you have more visibility and agility to avoid the danger is beyond childish.
Super easy cross compilation and incredible development velocity on small-medium projects are super cool features of Zig, and Rust can't beat that. No need to downplay the importance of Rust for the software industry (and as a friendly reminder, Rust is making its way to the Linux kernel, with the approval of Linus because unlike C++ or Ada isn't too complex to his taste ;).
Perhaps this may disappoint you, but I -- like many and perhaps most developers -- don't get such emotional responses, positive or negative to any programming language [1], which might appear as resentment to the emotionally attached. I am very impressed with some aspects of Rust, less impressed with others, and overall, my feelings toward it overall are shaped just like yours: by personal aesthetics. I don't find Rust's aesthetics very appealing and so Rust isn't my cup of tea (although I would't resent working in it because I'm not emotional toward languages and I currently program mostly in a language whose aesthetics I like even less than Rust's [2]), while you find them appealing and so you do like Rust. It's all just a matter of taste, and I fully accept that not everyone shares mine. I think your approach is too coloured by emotion, and is therefore unconstructive. You're a zealot, and you project that attitude on others, so "unconvinced" appears to you as a personal attack, and scepticism or dislike seems to you like bashing.
> Rust is even less complex than C, because of the “most vexing parse” issue. If you wanted a rigorous analysis of the syntactic complexity, you could attempt to measure how difficult it is to write a lexer and a parser for every popular languages, and see how Rust performs
No. The complexity of a formal language, like that of any set, is the computational complexity of deciding whether a string is in the language (so, including type-checking), not as the complexity of the parsing phase (https://en.wikipedia.org/wiki/Computational_complexity_theor...). I'm not saying this is the most useful way to talk about language complexity in this context (and caveats are needed, anyway, to make a more fine distinction between languages), but that is certainly one well-known way to talk about the intrinsic complexity of a language.
> Zig is a cool motorbike, Rust is a SUV. Arguing that your bike can indeed be safer than a SUV because you have more visibility and agility to avoid the danger is beyond childish.
It is beyond childish to make such inane statements about software correctness when you're clearly not very familiar with the subject, and are drawn to arguments like "more correctness => more soundness". The effect of language design on correctness is a complex subject with mostly unsatisfying answers, and even in software verification research, the debate over the value of soundness is far from settled (and not currently leaning toward more soundness). An equally inane statement would be, "Zig is like a modern aeroplane, relying on multiple levels of safety, some mechanical and some human, while Rust is like an old train that breaks down and kills everyone once there's a problem with the tracks." If we've learned anything about software correctness in the past decade it's that there is not much we can assume in advance, and that we don't really know one best way to improve it. It is true that some researchers think that the best answer to any correctness issue is more soundness in the language, but not only is this not a consensus opinion, I doubt it's even a majority opinion.
[1]: I would say I'm a "language sceptic." I'm generally sceptical toward any empirical claim about the bottom-line effectiveness of linguistic features without empirical support, and overall think that whatever empirical studies we do have show little impact overall to language design (comparing "reasonable" alternatives, at least), certainly compared to what all language fans claim. I would never, say, make a definitive claim like, "Zig yields more correct programs than Rust", or "Rust yields more correct programs than Zig," without clear empirical support (and my guess based on prior results would be that they're about the same).
[2]: C++
Your little “I'm a rational agent, you are too emotional” is pretty cute. But it would work better if your whole attitude in this thread didn't contradicted it, don't you think? “Rust is among the five most complex language” is not a rational argument, it's a personal feeling. Why you feel the need to spread your feelings over the internet while pretending you're not an “emotional” person is quite intriguing. If you want to look more like a rational person (no human really is), try to keep as much personal and unsubstantiated judgement out of your writings.
“Rust marketing makes safety claims that we should not take at face values” is alright, “Rust is one of the five most complex language ever” doesn't pass this test.
> is the computational complexity of deciding whether a string is in the language (so, including type-checking)
But for both Rust, C, Zig, and many others, it's undecidable, so by this definition of complexity, these language are definitely too complex (and equally so).[1] In fact, your desperate attempt to save your initial argument about complexity, by narrowing it to a tiny technical corner makes me cringe a bit.
> I would never, say, make a definitive claim like, "Zig yields more correct programs than Rust", or "Rust yields more correct programs than Zig," without clear empirical support (and my guess based on prior results would be that they're about the same).
The technicality of what constitute a “definitive claim” in a human to human conversation is an interesting question, but in practice truisms like “we can't conclude whether Rust is safer than C” or “we can't conclude that Rust doesn't bring more bugs than Zig does” aren't neutral: what such claim attempts to do is insinuate the opposite. And when combined with gratuitous judgements like “Rust is one of the 5 most complex language ever”, it looks a lot like an attempt to deter people from using this language you don't like. (And now I have a clue on what the root cause of your bad feelings can be).
[1] And as I said earlier, in the case of Zig and unsafe Rust and it's not just about type-checking: because of UB, even after compilation whether the compiled binary is the binary of a valid program or not is also undecidable.
My comments on this tried to be as careful and as precise as possible, and touched on this very issue.
> In fact, your desperate attempt to save your initial argument about complexity, by narrowing it to a tiny technical corner makes me cringe a bit.
Your emotional response here is so powerful that I think we're conversing on entirely different levels. I am sorry if my mild and careful statements have touched on something that you clearly see as essential to your identity.
> aren't neutral: what such claim attempts to do is insinuate the opposite
No. It is the most precise and careful statement that I can say, having followed the research for years and being a practitioner and advocate of formal formal methods. Anything I say, the more careful I try to be, it just seems to send you into further rage (and abuse). I think you're in a middle of a tantrum that's clouded your judgment, or perhaps, being a zealot, you cannot imagine any other attitude. Anyway, this conversation is making me very uncomfortable, as I sense you're in a very agitated emotional state, and I want no part of that.
I have to admit, this is cool rage-quit punchline!
Edit:
> I am sorry if my mild and careful statements have touched on something that you clearly see as essential to your identity, but I simply see no way to discuss this subject with you.
> Anyway, this conversation is making me very uncomfortable, as I sense you're in a very agitated emotional state, and I want no part of that.
Wow, it got even better as I refreshed the page!
At first order not even that; what people want is programs that behave correctly (or correctly enough for their purposes, which may not be very correct at all) for their particular inputs and execution environment.
Conversely those of us who want the industry to advance the state of the art generally don't want to just produce correct programs at a particular point in time, but programs whose correctness can be easily maintained even as implementations and requirements change. More than that, we want to produce libraries and frameworks that will lead as-yet-unknown programs to be correct.
Ah, yes. This raises an interesting philosophical question with real ramifications for software quality assurance: is a bug in the algorithm that never manifests in the system really a bug? Something like that happened in two well-used pieces of code: There was a bug in the TimSort algorithm used in both Java and Python, whose probability of actual failure is similar to the probability of failure due to a bit flip caused by cosmic rays. Because hardware can only be correct with probability, no running system can be soundly verified, i.e. with certainty, anyway, so while the correctness of algorithms can be absolute, the correctness of system cannot. And since soundness has a big cost in verification, many in software correctness research now focus on unsound techniques that are cheaper.
> Conversely those of us who want the industry to advance the state of the art generally don't want to just produce correct programs at a particular point in time, but programs whose correctness can be easily maintained even as implementations and requirements change. More than that, we want to produce libraries and frameworks that will lead as-yet-unknown programs to be correct.
True, but that is not a winning argument for soundness. The cost of soundness manifests even at maintenance. It's therefore an equally strong argument that a language that compiles quickly and more easily allows running, say, concolic tests, mutation tests etc., serves that goal, too.
A language that makes code reviews easier also works toward that goal of maintaining program correctness over time. The point is, there are many different paths to correctness, all of them state-of-the-art yet are often in conflict with one another, and we don't have any mechanism other than empirical research to compare them. For example, is it beneficial to increase soundness at the expense of making code reviews harder? Not only do we not have an answer to that question, it is likely that there is no general answer (I say it's likely because whatever empirical research we do have shows messy results with large variance).
So much this. And also keep in mind the way that we typically do code reviews, we typically are looking at github diffs. So if you are in a situation where changing code is ill-composable, for example, if something looks safe in place A and something looks safe in place B but when you put them together it's not unsafe... Then you could be in deep trouble with the async way that we do reviews.
I can already use C and C++ to suffer that in production, use VC++ static analysers to mitigate them, while languages like Ada, D, Rust, Nim, Swift take care of them not happening at all.
C17 is also quite different from K&R C, specially in what optimizers do with UB.
The only way OS vendors have to fix issues that Zig also shares with C, C++ and Objective-C, as per the article, is to adopt hardware memory tagging, something already available on Solaris SPARC, Azure Sphere and iOS (yes PAC is a bit different), with ongoing work for ARM.
So I really don't see the benefit, but lets see how Zig 1.0 actually looks like, and I might be wrong by then.
But I've long ago learned that language preference is mostly a matter of personal aesthetics, so all I can say is that I find Zig very appealing. Its design is certainly radical, and it doesn't feel like any other low-level language I've ever seen (it is about equidistant from C, C++, Rust, D, Nim, Ada; even when pushed I don't think I'd be able to say which of those Zig is most like, because it is so different). Like it or not, it offers a fresh vision on how low-level programming can be done.
To this list then I would also want to add and compare: OOM-safety under overload conditions, and fine-grained error handling safety, in particular because error handling tends to be one of the leading causes of faults in distributed systems [1].
To be fair, I was surprised that Rust did not have checked arithmetic on by default and that this needs to be turned on via compiler setting or linted against. The presence of integer overflow in a program can facilitate a whole range of exploits, even with memory safety.
[1] - https://www.eecg.utoronto.ca/~yuan/papers/failure_analysis_o...
Do you have examples that do not involve array indexing? (which are uncommon in Rust because iterators exist and are faster).
* Rate limits for an OTP, which could be trivially reset through overflow.
* Parsing zip files, where the hostile content is self-referencing, to sneak in cyclic references for a DoS, or to change file extension type for code execution after bypassing a content filter, or to change output destination (e.g. as part of a symbolic link or directory traversal) to overwrite system files.
Unchecked arithmetic for me is by far one of the scariest exploit vectors because it's so easy to do, and one of the first things a trained attacker would look for.
Programs in any language are always exploitable, in some way or another. Memory safety is no guarantee that a program is "safe" let alone correct.
A good example of this is the SQLite documentation page "Why Is SQLite Coded In C". Among other things, it describes those memory safety issues as "the easy problems" compared to "the rather more difficult problem of computing a correct answer to an SQL statement".
Not all of our programs are like SQLite, of course (and not all of us mere mortals are like its developers). But I would certainly say that just because you've eliminated memory safety bugs doesn't mean you've eliminated all bugs. Depending on the program, you might not even have eliminated most bugs.
https://research.checkpoint.com/2019/select-code_execution-f...
> We established that simply querying a database may not be as safe as you expect. Using our innovative techniques of Query Hijacking and Query Oriented Programming, we proved that memory corruption issues in SQLite can now be reliably exploited. As our permissions hierarchies become more segmented than ever, it is clear that we must rethink the boundaries of trusted/untrusted SQL input. To demonstrate these concepts, we achieved remote code execution on a password stealer backend running PHP7 and gained persistency with higher privileges on iOS. We believe that these are just a couple of use cases in the endless landscape of SQLite.
> All that said, it is possible that SQLite might one day be recoded in Rust. Recoding SQLite in Go is unlikely since Go hates assert(). But Rust is a possibility. Some preconditions that must occur before SQLite is recoded in Rust include:
> Rust needs to mature a little more, stop changing so fast, and move further toward being old and boring.
> A. Rust needs to demonstrate that it can be used to create general-purpose libraries that are callable from all other programming languages.
> B. Rust needs to demonstrate that it can produce object code that works on obscure embedded devices, including devices that lack an operating system.
> C. Rust needs to pick up the necessary tooling that enables one to do 100% branch coverage testing of the compiled binaries.
> D. Rust needs a mechanism to recover gracefully from OOM errors.
> E. Rust needs to demonstrate that it can do the kinds of work that C does in SQLite without a significant speed penalty.
Except it does, in debug mode, which is the default compilation mode. And you don't have to tweak compiler settings or linters to get checked arithmetic in release: you simply call the checked arithmetic functions, which give the added benefit of giving you complete freedom in how to handle overflow.
Wherever you 'learned' this misinformation from, please stop considering it a trustworthy source :(
https://uploads.peterme.net/nimsafe.html
(Source: https://www.reddit.com/r/nim/comments/maj1lz/nim_safety_in_c...)
Nim now has ARC/ORC, which is a reference counting scheme similar to Swift. This is not a "generational GC" like Java, Go, or D.
The entire standard library and most average libraries, "just work" with it, and it will be the default near future.
In contrast, I have yet to be able to build any RISC-V binaries with Rust. It just doesn't work. Sure, I could see some potential things like writing a custom JSON to describe the env and maybe build using a cross-compilation toolchain. But after a certain amount of time and no answers, it was not worth my time anymore.
https://stackoverflow.com/questions/64308644/rust-unable-to-...
If you think you have the answer ^
I kind of like this philosophy, because in a sly way it's a carrot to get you to write tests. Come for the memory safety, stay for the robustness.
As a bonus, the beginning 2 hours of the video is a fantastic and honest discussion about the role of emotional empathy in tech communities and tech employment (while also acknowledging that it is possible to be an asshole and deliver good tech).
If the GPA is practical to use in production, that will be a different story. But it doesn't sound like it's there yet.
The more I read about Zig I feel like it's made for C programmers stuck in the Stockholm Syndrome of C and don't want out. (Speaking as a former C programmer.)
But if you have a clever language that navigates this tradeoff and lets you build powerful zero-cost abstractions (C++, Rust, Zig), then it attracts significantly more talent that compounds.
Everyone praises UNIX for pipes in the shell and forgets about OS composition APIs.
Regarding talent, it depends on the kind of talent you’re talking about. Yes, systems programmers like Rust. On the other hand, needing to deal with a strict borrow checker excludes a very large number of people with numerical computing, data science and machine learning expertise (not all, but definitely most). So it cuts both ways.
Go has a GC, which helps with memory, but you're entirely on your own for files, sockets, database connections, mutexes, channels, etc.
The features that Rust uses for memory management are fully general-purpose, and also help you with safely handling all other kinds of resources.
Consider how many ways you can misbehave with a Mutex in Go that are all caught by rustc at compile-time. This has nothing to do with memory management, but the same thing that prevents use-after-free is what also prevents using a mutex-protected value after releasing the lock.
I really like zig's approach of explicitness and fast iteration cycles. Fast compile times and the very flexible build system makes me hopeful for a really slick workflow for embedded development,where zig code can be used to deploy and test as well. For my own use I think it's a clear win.
On the hand, the amount of damage poorly architected zig code can cause is about as large as for poor c code. For typical enterprise code the rust compiler will make sure that many bad decisions will not even compile. There's still a risk of towering abstractions, but at least I could avoid spending as much time debugging hideous race conditions.
The borrow checker is not about memory safety but about aliasing guaranteed.
It just happen that combining this with deterministic destructors (/RAII) happend to enable reliable "automatic-manual-memory-management" (or however to callit).
And combining it with some clever auto traits (Send/Sync) happen to prevent data races (if no unsafe is used, like always).
But the benefits are not limited to just that. Not just memory-resource management but also other resource management profits from this design.
Similar while Send/Sync is about multi threaded data race prevention there are also problems in single threaded patterns which are quite similar to that e.g. "racing" between iterating a collection and changing it in the body of the iteration, and the aliasing guarantees make sure you don't have such problems either.
Similar rusts main pointer type (`&`) does not only provide compiler time non-null guarantees but also provides compiler time guarantees about how the data can be accessed (dereferencable, writeable etc.).
And then there is the choice to use the type system to prevent application logic bugs in many ways.
So the bullet points in the table miss many dimensions .
But then zig is still a grate language, but trying to convince people that it's good enough by telling them that not reusing allocations seem to not be the best way.
Instead look at arguments why people still use C today (not C++!) what they conceptually like about it and you might realize many of the parts still apply to Zig.
Honestly Zig seems to be a grate choice for webasm or similar sandboxed systems where the potential damage of use-after free or double frees can be massively reduced.
I think the biggest thing was that university curriculums and mainstream app development platforms (like Microsoft) stopped pushing it as hard when the level of horror got past a certain point. It used to be pretty bad. Business apps being written using MS "Active Template Library" in C++ and then used as signed ActiveX plugins on IE6-only web pages etc.
Though I believe there were some languages with features to that end, at least research languages, they weren't that well represented. I think Rust's presence brought attention to the possibilities there, and an increasing number of people see the value of investigating and developing that niche.
So basically you have the DevTools and Azure teams pushing for .NET, Java and other safer languages, while Azure Sphere has a C only SDK and WinUI/UWP push C++ above anything else, with some C++ only APIs.
Politics.
Deeply embedded code doesn't use malloc, doesn't use threading.
I could use a better type system, ala Ada, being able to say "this variable is of type distance in meters, this variable is of type time in milliseconds", that'd have cut the # of bugs by a huge amount.
But simple, unsexy, type system changes like that aren't what language designers are focused on.
Who here has never confused Milliseconds and Seconds when passing a variable around? Trivial for a compiler to catch with a half decent type system, but few modern languages bother to try.
Even when writing modern code in newer languages, I rarely directly use threads, and if I need to pass data between them 95% of the time I can get away just doing a deep copy to avoid the hassles of sharing data between threads!
Obviously Rust is meant solving different problems than the ones I face, I have friends who frequently write highly threaded code, but in my day to day, Rust doesn't offer much more safety.
(However, Zig does look super cool and interesting!)
Rust seems to be a very complex language. Is all that complexity essential to providing memory safety without GC? Or would it be possible to have a significantly simpler language that is equally safe? A language that’s “safe & C-like” compared to Rust’s “safe & C++-like”?
I'm not a rust developer, but I've ported over a handful of python or golang projects to see how it works. I managed to write the code without understanding much about things like the borrow checker.
I'm certain my code is not as performant or elegant as it could be using some of the more complex tools and concepts in the language, but it is possible.
It’s also true that team using languages with more feature then they need can just take the parts that they need. It’s not quite as ideal as having a language that’s perfectly suited to your use cases, but it works well enough.
For example, I’ve been writing JavaScript for 3+ years. I have yet to use the prototype chain directly, only through the use of the `class` keyword, and I’ve only reviewed code using it once. I hear C++ is similar in that teams use a slice of the available language features.
If you don’t need maximum-possible performance, though, why use Rust/C/C++ at all? Wouldn’t a better choice be Go/Java/C#?
Of course, I know who the author of that code is. I just wanted to point out that it's not such a trivial comparison to make.
I tried Go, I wanted generics and errors. I like C#, but not the ecosystem that comes along with it. And so on. So for me personally, Rust is a a valid choice even when performance isn’t a first concern.
Technically true, but only really true if you don't use many dependencies. At some point you're going to use some dependency that uses async/await all over the place or really goes wild with generics and then it is definitely not simple.
Two examples:
* Heim (https://docs.rs/heim/0.0.11/heim/) is a great crate for getting system info, but it only uses async/await so you are thrown into that rather painful world even if you don't need it.
* Plotters (https://github.com/38/plotters) is a pretty great graph plotting library for Rust (the only one as far as I know), but they have definitely gone a bit overboard with the generics. Want to draw a scatter graph?
I tried simply calling `PointSeries::new()` and got a basically impossible-to-follow error about Rust not being able to infer the type `E` here:
https://github.com/38/plotters/blob/master/src/series/point_...
Very simple it is not!
In reality it isn't nearly as simple as that. There's been a lot of discussion about the pitfalls of async Rust recently, highlighting the issues.
Here's a good one: https://tomaka.medium.com/a-look-back-at-asynchronous-rust-d...
The language features specific to memory safety i.e. the borrow checker, are essentially irreducible. It is also Rust's biggest piece of complexity, and the one that is hardest to learn. There is no simpler language inside Rust that has the same safety guarantees, unless you strip out other useful features (traits, async, etc.).
This sounds like “there is no simpler language inside C++, unless you strip out useful features (classes, templates, etc.)”. Yet C exists.
So there could be a simpler language with Rust’s safety guarantees if you were willing to strip out traits, async etc?
Well yes, there exists a hypothetical C + borrow-checker language. But that language wouldn't really be significantly simpler, because the borrow-checker is the largest contributor to Rust's complexity. The only things you would have taken away are the more well-understood features, as they already occur in other languages.
The C/C++ comparison doesn't really work, because there is (to my knowledge) no single C++ feature which makes up the majority of its complexity over C. You could strip out independent features of C++ one at a time to return to a simpler language. Rust doesn't have the same property.
Maybe template programming?
I’d argue there is: there’s reference counting. Rather than using references and fussing with lifetimes you could sprinkle Rc<> wherever it’s necessary. You’d take a performance hit but the code would be simpler to write.
Then there is Ada, but it is in the same complexity level as C++, although much easier than Rust.
I hope something like Zig gets widespread adoption, including in embedded/IoT/automotive environments. Especially automotive. We're moving more and more life-and-death scenario-type tools into software.
There definitely is a design space for a simpler language than Rust that is easier to write, but Zig is too far on the side of C and has lots of trivially introducable unsafety. It's an improvement over C , but imo not enough.
I have tried to like the language, but sadly having to think about types and lifetimes robs precious energy which should be devoted to thinking about business rules and what am I actually trying to achieve.
In some niches Rust is perfect, but in every language thread on HN there's often someone that suggests to use Rust whatever the use case. C, in that respect, is more flexible and gets out of the way much more, of course while being unsafer, but it's easier to keep your mind on the goal and not figuring out the best memory safe approach for this piece of logic.
Which is why I'm very excited for Zig. I don't want another C++. Give me safer C, thanks.
To be honest I haven't used C in a long time, but I've been looking for a low level language that sparks as much joy as C does. Go, Rust ain't it, IMO.
Now there's an argument the front loading these decisions may be beneficial overall but I don't find that compelling either. After the "make it work" stage there's thinking about performance, comments, logging, etc and you usually need to think about lifetimes as they change at this point as well, so front loading it has only added work overall.
During the exploratory phase of my projects, I very rarely run into nontrivial lifetime issues, and when I do, I can just put whatever data into an Arc and then it just works. Most of my exploratory code is just objects that own their data.
Later in a project, while doing a bunch of refactoring to handle performance, logging, etc. , I have a much easier time letting the compiler tell me when I've made mistakes with a value's lifetime rather than trying to keep the whole program in my head for the duration of the refactor.
I don't spend a lot of time thinking about lifetime issues during either phase of my projects. Mostly I just write the same sort of thing that I'd have written in other languages, and the compiler tells me when I've made a mistake.
I've always wanted a "shut up about memory safety for a while, just don't free anything, I want to find out if my code produces the right answer" mode in rustc.
But especially in all those domains where memory safety is an issue , in my view it is currently the best option.
Rust forces you to think about memory safet and ownership, which is hard to adjust to for many. But it does so for a good reason.
In C you also have to think about lifetimes all the time, but the compiler let's you do whatever you want , and the issues instead have to get fixed when bugs pop up, or with static analysis tooling, etc.
"I don't want to think about lifetimes" is exactly how we end up with vulnet and buggy software.
After the initial learning curve, Rust is a very productive language, exactly thanks to the powerful type system.
Like I said, I do wish for a simpler language that can provide similar guarantees, and I do think the design space is in reach, but Zig is (currently) not it.
For example, if an attacker could arbitrarily inject overload to restart rate-limiting processes and then abuse this to trivially brute-force OTP logins.
The definition of safety with respect to a resource in general always needs to include the safety of the system as it crosses the threshold i.e. in this case into out-of-memory, so if a system claims memory safety, the first thing I would want to ask is, what about OOMs?
Anyway, for Windows and other plateform where this is a reasonable goal, there is work in progress to add this to Rust. See this RFC[1] which has been merged and whose implementation progress can be followed here [2]
And taking this further, since Zig's convention around allocators is for them to be an explicit argument of all functions needing to allocate, it's trivial to write tests specifically to validate correct behavior in OOM conditions. There's even a custom allocator in the standard library for exactly this purpose.
Beyond the correctness argument, also because the GC can really come back to bite you when you least expect, following the sudden "knee" of Little's Law. I've seen multi-minute pauses every few seconds even with V8's GC in production and it was not a pleasant experience. It cropped up, out of the blue, and in the end required a V8 core team member to advise and help comment out a few lines of C++ GC code that were overzealous.
GCed is orthogonal (as in: doesn't have anything to do) to type safety.
Java is GCed, so are Scala, Kotlin, C#, F#.
Even dynamic GCed languages towards the scripting side of things are moving to static typing: Typescript, Python Mypy, Ruby types (I forgot the name of the project).
That's a really strange argument, because as soon as you're not in a GCed language, you need to think about the lifetime of your objects. The big difference with Rust is that you can't make mistakes when doing so, because the compiler will catch it.
You don't have the mental burden that if you make a mistake everything will blow up and you can focus on your business rules instead.
JetBrains went out of their way to make migrating to Kotlin as easy as possible: you can literally upgrade your Java project file by file.
Yes, offering a good path to switching (and conversely keeping the old voodoo part of the code no one wants to touch) is the way to go. And as I understand it, Zig offers this possibility as well.
https://andrewkelley.me/post/zig-cc-powerful-drop-in-replace...
As Java evolves and Kotlin needs to cater to Mountain View masters, upgrading the Java file won't be enough as many modern features don't exist on ART.
I think the first Kotlin version to use post-Java 9 bytecode is the Kotlin 1.4.20 using invokedynamic for string concatenation: https://blog.jetbrains.com/kotlin/2020/11/kotlin-1-4-20-rele...
Kotlin sealed classes are planned to be rendered as JVM sealed classes when JVM sealed classes go out of preview (probably Java 17), and the same is planned for mapping Kotlin value classes as Project Valhalla user-defined primitive types.
All these features obviously won't be available if you target Android, but they will still be there for you if you don't.
Faking JVM features on other platforms means that the performance is not the expected one when moving across them, and some surprises might happen when linking to libraries that use modern features.
Copied here [from lobsters](https://lobste.rs/s/v5y4jb/how_safe_is_zig#c_vddk9j):
With regards to use after free and double free, this is solved in practice for heap allocations. The basic building block of heap allocation, page_allocator, uses a global, atomic, monotonically increasing address hint to mmap, ensuring that virtual address pages are not reused until the entire virtual address space has been exhausted. In practice, this is a very long time for 64-bit applications. The standard library GeneralPurposeAllocator in safe build modes follows a similar strategy for large allocations and for small allocations, does not re-use slots. Similarly, an ArenaAllocator backed by page_allocator does not re-use any virtual addresses.
This covers all the use cases of heap allocation, and so I think it’s worth making the safety table a bit more detailed to take this scenario into account. Additionally, as far as stack allocations go, there is a plan to do escape analysis and add this (optional) safety for stack allocations as well.
As far as initialized memory goes, zig forces you to initialize all variable declarations. So an uninitialized memory has the word undefined in it. And in safe build modes, this writes 0xaa bytes to the memory. This is not enough to be considered “safe” but I think it’s enough that the word “none” is not quite correct.
As for data races, this is an area where Rust completely crushes Zig in terms of safety, hands down.
I do want to note that the safety story in zig is [under active development](https://github.com/ziglang/zig/projects/3) and will be worth checking back in on in a year or two and see what has changed :)
out-of-bounds heap read/write: runtime, some cases at compile time
null pointer dereference: relies on hardware protection
type confusion : compile time
integer overflow: wrap-around semantics
use after free: prototype protection in @live functions, not a problem when GC is used
double free: prototype protection in @live functions, not a problem when GC is used
invalid stack read/write: compile time
uninitialized memory: compile time
data race: read/write to shared memory can only be done via library functions
> integer overflow: wrap-around semantics
Interesting choice of words to not say "none".
Back in the 16 bit real mode DOS days, writing to the first K of memory would trash the interrupt vector table. But that was a 1980 design.
But hey, it sells better than Flash or PNaCL.
The null pointer thing is true however all the other mechanisms mentioned do help eliminate bumping into them, and there will be more on the way.
If you want to provide the null checks yourself D has Ada-style contract programming too.
Not really rusts pointers are the `&`/`&mut` references which are compiler time proven to be not just non null but actually differentiable and potentially writable in the given context. Which are MUCH stronger guarantees then just "not null".
> 2. only when using tagged unions
Which in rust are the default types, unions didn't exist for quite a while and require the use of unsafe making them heavily discouraged to be used.
Besides that it's not that rust enums are tagged unions where you at runtime check a tag and then access them, or where you "panic/throw an exception" when you try to access the wrong type, but incoperated into the type system and language given quite a different experience to classical tagged unions.
Lastly type confusions applies to more then just "tagged union style access" but also subtype stile access in which case rust can use trait objects instead of sum types.
Besides that many of the ways listed zig can archive more safety (weather applicable or not) are also applicable to C. And some of the checks Zig do can be "somewhat" archived in C too by combining non standard compiler options and code analysis tools.
Don't get me wrong. Zig is a very interesting language and I would argue the spiritual successor of C in how it's designed.
Still I guess the main ways to add more safety (and similar) to languages like Zig (or C) is to compile them to WebAsm. The module isolation while still being able to call other functions without to much overhead which can be archived with WebAsm might lead to quite interesting trends in the future.
I mostly have experience building things in GC languages. But with Rust I managed to safely use [1]:
- stack references in threads
- kept mmap references alive until threads finish work
- zero copy xml parsing (from mmaped data!)
- SSE/AVX enabled searching
The Rust language empowered me to do these things with a high degree of confidence. Not one segfault or core dump, just lots of compiler errors.
I played with Zig. Admittedly, the small ecosystem aspect is something all languages go thru, and it would be a better experience with a Zig specific libraries. But Zig doesn't empower library authors to make a large category of bugs impossible, and leaves it to documentation. This is like C, I don't have enough confidence in myself to use it.
Brilliant people are building powerful, safe-ish, reusable libraries in Rust. For mere mortals like me, this is Awesome.
[1]: https://gist.github.com/daaku/58557e2545612df8f40b13b66b7d3b...
I believe your use of `unsafe` on this line is unsound: https://gist.github.com/daaku/58557e2545612df8f40b13b66b7d3b...
Namely, there is no guarantee that the bytes between `<page>` and `</page>` will be valid UTF-8. It may be the case that you only run this program with UTF-8 input, in which case, UB is never triggered. But it's worth pointing out here since there is nothing actually stopping your program from hitting UB.
Also, as long as you're bringing in the twoway crate, you might as well use it on lines 43 and 48 since you're just searching for a single needle.
I brought in `twoway` when I couldn't find a way to `rfind` using `aho-corasick`. I'll switch the use over for consistency.
Thanks for the quick code review!
PS: Thanks for ripgrep too!
Either way, my point here is to be a counter-balance. To be fair, you did say, "But with Rust I managed to safely use." But the code you posted is technically unsound. It's not a huge deal if you know you'll always be feeding the program valid UTF-8. But it is worth mentioning here in this HN thread that is specifically comparing the safety properties of competing programming languages. :-)
For example, let’s say you have `a = b + c` with b/c being i8 and a i32. This calculation is first performed as an 8 bit add, then extended to 32 bits. This is true for both Rust and Zig, but Rust requires an explicit cast to widen the result of `b + c`, making it obvious that an extension happens and that `b + c` is not performed in 32 bits. In Zig there is no such indication- you need to look up the definition of b and c. Other problems occur as well, that both C and Rust avoid in their own ways. Hopefully Zig can improve this situation.
See more here: https://c3.handmade.network/blogs/p/7651-overflow_trapping_i...
Zig is apparently a PL for folks with perfect memory who never make mistakes like those described in that ticket. "Lesser" programmers like me can choose another language.
But I think this page may be overstated? Again: I don't know anything about Zig, but I sure know how a UAF bug works. :) And it doesn't look like Zig is meaningfully susceptible to them? You an crash a Zig program with a UAF, but the actual vulnerability wants more than the crash: it wants the program making uncontrolled writes to live memory used elsewhere in the program, which is a condition I don't think is present in Zig as it's being described.
If that's the case, that bodes poorly for the claim that Zig is susceptible to C/C++-style double free vulnerabilities, too.
It would be genuinely weird to see a new language rolling out that had C/C++'s UAF problem.
(As was pointed out elsewhere: if you're using an external allocator, or the `c_allocator`, all bets are off. But so is unsafe code in Rust, I guess?)
> The standard library includes a set of allocators which don't reuse allocations, preventing use-after-free, and which catch double-free. I'm not clear yet on how high the runtime and memory overhead are though, which will dictate when it is practical to use these.
I didn't include it in the table because I'm not yet convinced that the overhead will be low enough that people will actually ship software using those allocators. (All the zig programs I've written so far use the libc allocator and are definitely susceptible to UAF)
Perhaps I'll spend some time measuring it this week and post an update.
And more from https://github.com/google/sanitizers/wiki/AddressSanitizerCo...
Zig vs Rust vs C aside, this cannot be a serious position for a software developer to take in 2021 CE.
Suppose you just document "please take this into account" because it's a low-level FFI library, that's being called by a high-level PL, which will make sure that your memory management is sane.
An analogy: "non-threadsafe" code; typically you just tag 'non-threadsafe' and if someone misuses it, it's on them.
> For various common safety issues, we can look at protections that are present in software as it is typically shipped (ie excluding tools like AddressSanitizer that are not recommended for production use):
...with a link to a five year old opinion post to oss-security as a reference for "not recommended".
To wit: "There is other stuff in this space that might be relevant, but I don't want to talk about it so I'll just make up a reason. Moving on..."
That kind of logic tells me instantly that this is spin and not a serious analysis. I know next to nothing about Zig, but I know I shouldn't trust this post to tell me about it.
I will be staying away from Zig exactly for that purpose. Great idea but I can't get behind a maintainer that adamantly refuses to even discuss proper, safe standard library design.
EDIT: Yep, the HN crowd tends to be the same. Downvote me all you want please :) We'll see over time.
The bulk of the conversation happened in Discord around that time. Initial attempts to bring this up were met with "zig is perfect"-type conversation, none of which was very technical.
Finally, the conversation grew to be so large and fiery that Andrew had to step in and say "everyone play nice, now!" and then head back out.
Then this PR was filed. Invalid user input should not be classed as undefined behavior and concluding that the standard library's UTF decoder shouldn't be used if you want a safe execution is just absurd.
There were a few other run-ins on discord in the same vein. It made a few people at the time leave, including myself.
Andrew's smart. Zig is a cool idea. But I don't like when this laisse-faire attitude is taken when designing a programming language that places so much emphasis on being safe.
Just to copy Andrew's final words in here:
> I think the entire std.unicode needs an audit both in terms of API design and performance. This module is not yet what it will become before stabilization. But this commit is not where this is going.
The PR wanted to make a function that takes runtime values only take comptime values. I read Andrew's response as saying "this all needs to be looked at before 1.0 but this isn't the way to fix this", which seems to me to be an entirely reasonable thing to reject a PR with.
I think you're getting downvoted because "citation needed", while accusing the Zig community of "cult-like" behavior without justification.
Where has Andrew "shot down discussions about DOS vulnerabilities in the standard library"? As is, your comment isn't helpful.