Failing to Learn Zig via Advent of Code
forrestthewoods.com
forrestthewoods.com
Every year, I see people quitting their AoC run with frustration, because they don't measure their effort and try to complete each puzzle daily. The usual scenario is that, at some point, they block on a problem and spend longer than they should, or they can't solve the problem because they have some other thing to do. The next day, they try to solve that puzzle and another one, and as work piles up, they start doing it joylessly to catch up. After a while, it doesn't work anymore, because people are stressed, because they code poorly just to finish rather than taking pleasure in their craft. The strict calendar can become deadlines that remind of work woes.
For a person not used to sport code, AoC can take an hour to solve, sometimes more. Over the whole advent, people can easily work for 30 to 40 hours more than they are used to. In terms of work hygiene, this is not good. Not everyone has a week's worth of work to use just before the end of year.
I would advise people doing AoC to time their effort, and not hesitate to stop a puzzle and write it down for later in the year.
Zig is basically C with a fancy type system, so you should not expect things like special String types, overloading of index based access etc.
I think the author was thinking that Zig was very close to Rust or C++, when in reality it is much closer to C. I had to keep reminding myself of that many times as I was learning Zig.
I had my own struggled with Zig, but not quite as much as the author. I think will probably have a much better experience if you don't try to jump and code right away but read some articles or listen to some videos to get a sense of the overall philosophy of Zig.
I am normally against having to look at source code, but with Zig that is kind of needed but also not quite as bad as it sounds. Zig code base is not that large and it is relatively easy to search. You can lookup a Zig function signature very easily. You need to do this if you are going to use any of the standard library apart from the most basic stuff.
I had the impression Zig proponents were praising it because it wasn't as low-level as Rust.
In Rust communities, it's often pitched as alternates to both, but closer to C.
I suppose it's all relative, but the comparison of rust to c++ seems external
Though I believe the mozilla code it's replaced was all -very- C++.
However it can occupy some niches that C can but C++ does poorly because of zero cost abstractions, I think.
(Very handwavy)
Rust is already approaching C++ in complexity, surpassing it in some places; and also in expressive power, but not surpassing it anywhere yet.
If Rust does not end up fizzling (which is still very possible!), Rust programmers will generally be drawn from the same population as C++ programmers. They will be people who want and can use a powerful language to make themselves more productive and able to manage bigger projects, without need to worry that they are taking a performance penalty, or losing control of details that matter.
Users of Zig, like of Nim and C, will be those uncomfortable with language power, disinclined to automate. Their attention is not on software and what they can build of it, but on problems where a thin veneer of software can add something useful. When there is not much for the software to do, you don't need much power to get it doing that.
A subject that now has become even regular presence at C++ conferences and considered a must have in static analysers roadmap by all major vendors.
Rust might fizzle out in a decade, and still leave such a mark in the industry.
Chapel, HPC language mostly sponsored by Intel and HPC
D programming language,
https://dlang.org/blog/2019/07/15/ownership-and-borrowing-in...
Ada/SPARK,
https://docs.adacore.com/spark2014-docs/html/ug/en/source/la...
Swift,
https://github.com/apple/swift/blob/main/docs/OwnershipManif...
ParaSail
Project Verona from Microsoft Research
https://www.microsoft.com/en-us/research/project/project-ver...
Project Snowflake from Microsoft Research
https://www.microsoft.com/en-us/research/publication/project...
And finally your favourite C++
"Implementing the C++ Core Guidelines’ Lifetime Safety Profile in Clang"
https://llvm.org/devmtg/2019-04/slides/TechTalk-Horvath-Impl...
Also the "Clang Static Analyzer - A Tryst with Smart Pointers" talk at 2021 LLVM Developers Meeting.
For the Visual C++ part of the story
https://devblogs.microsoft.com/cppblog/lifetime-profile-upda...
And GCC as well, although they are late to the party
https://gcc.gnu.org/wiki/DavidMalcolm/StaticAnalyzer
Finally a couple of CppCon 2021 talks that touch on the subject in various ways,
Type-and-resource safety in modern C++
Code Analysis++
Static Analysis and Program Safety in C++: Making it Real
Finding Bugs Using Path-Sensitive Static Analysis
Minor correction: Intel hasn't traditionally been a sponsor of Chapel (though we'd love to see that change). Chapel was pioneered at Cray Inc. and continues on even stronger at HPE after its acquisition of Cray.
-Brad
All the best.
Also, why even draw a stark contrast between "a language" and "its tooling"? As a dev, you get to use both.
What is even the line..? Almost every compiler for anything provides options. Does gcc -fsanitize=.. not count because it's not "standardized" or only because its not activated in "typical" deployments like Rust integer overflow checks?
C++ type system is impossible to fix while keeping backwards compatibility, so static analysis tooling is the only possible solution.
I also think users of C (not sure about Zig) are quite happy to automate things. Linus Torvalds is a big user of C. He wrote a little C-like compiler to check Linux kernel code called Sparse [1]. You seem to be trying to discuss maybe larger (but not very well articulated) subpopulations of "Users" than Apex Programmers like Linus. It is definitely easier to do this with C than giant languages like C++.
Why, the 1980s & 1990s were littered with maybe dozens of hacked C compilers doing "this or that" automation in a way you do not see for C++ (and will probably never see for Rust). In point of fact, C++ itself (C with classes) was an early example of such! The idea was to automate/codify the object-oriented style of Simula in C.
pjmlp's sibling & child comments are also some good color on the history/context of all this. { Of course, partly it all depends on what you meant by "language power" and "automate" - I am just going by what that seemed like. }
Obviously, C and Rust are both low-level languages (in terms of control and overhead), but there are quite a few of those.
>>There is no hidden control flow, no hidden memory allocations, no preprocessor, and no macros.
https://ziglang.org/learn/overview/
It try's (achieves?) to be a C without the flaws and historic ballast.
I never read anything about Zig outside of HN.
What I remember was mostly that Zig should be easier to use than Rust, because Zig doesn't have lifetimes.
If I can't have nice strings, what should I be expecting from a fancy type system?
- non-null pointers, and distinct types for single-item pointers and multi-item pointers (multi-item pointers are rarely used except indirectly via slices, so unchecked pointer arithmetic errors are largely banished)
- builtin tagged unions (AKA algebraic data types) with very pleasant to use switch logic -- it can't be overstated how nice the "handle all the cases" logic is in Zig in general (catch, orelse, if-else/switch expressions)
- a decent proposition for errors (the error union, and try keyword), but I haven't decided if I really like it yet
There's lots of non-type stuff there too. I was writing personal projects in C without libc, but found there to be a lot of annoying work involved -- happy to do it, but it's not earned/fruitful annoyance, more like a long list of incidental historical annoyances. Zig seems to cater to the same level of the stack, but with all that boring stuff taken care of.
- an array [3]u8
- a single item non-nullable pointer *u8
- a single item nullable pointer ?*u8
- a multi-item non-nullable pointer [*]u8
- a multi-item nullable pointer ?[*]u8
- a slice []u8
Typically your API is just made up of slices and non-nullable single item pointers. Arrays are just the typical backing store for a slice, that you might define in main or for small scratch buffers. Here's a typical example of an OS read (where the fd has been set to non-blocking already):
fn readUpTo32Bytes(fd: std.os.fd_t) ![]u8 {
var array: [32]u8 = undefined;
var data = array[0..]; // slice of entire array
data.len = std.os.read(fd, data) catch |err| switch (err) {
error.WouldBlock => 0,
else => return err,
};
handle(data);
}
fn handle(data: []const u8) void { ... }
If you don't need to handle EAGAIN/error.WouldBlock, then just do `data.len = try std.os.read(fd, data);`I didn't explain the `!` in the `![]u8` return type, but it's basically saying "an error or a []u8", where it's compile time known what the full set of errors is (in this case every error std.os.read can return, minus WouldBlock).
Coming from C, I find that Zig's memory alignment options are easier and more powerful.
AFAICT the answer is "inject as many runtime checks as needed" although the docs seem to go way out of their way to avoid making this explicit.... or deal with the fact that these checks are now runtime failures rather than compile-time failures, and therefore need code to handle them.
It seems like it would be the same as writing Rust code using std::cell instead of references, except that Rust would force you to insert handlers for all the new failure modes this would create (of course you could just panic!(), but at least the compiler would force you to insert those panics...).
However, Zig treats memory safety not as an extreme-at-all-costs but as a spectrum (there are reasons for wanting to think like this, at least when writing low-level code if you want to make more use of the hardware), getting 100% on the spatial memory safety front and reaching to 50-75% on the temporal memory safety front through the GPA. That's already an order of magnitude more safety than C, at which point memory exploits have dropped in ranking, and you should be more concerned about things like explicit control flow, error handling and checked arithmetic, not to mention the orthogonality of the language.
Furthermore, in the systems world, there are many safety critical systems where dynamic allocation and multi-threaded control planes are simply off the table to begin with because they're dangerous in some domains and not as safe as static allocation and single-threaded control planes, which are less dimensional and easier to reason about. And in those cases, UAFs and multi-threaded races are less of a concern (still a concern, but less).
Also, Rust won't protect you from all undefined behavior, and Zig often helps more than you think. For example, you might be surprised to hear that Rust has checked arithmetic off by default in safe builds, whereas Zig has this enabled. I've done a little security work on some large systems and the decision to disable checked arithmetic always blows my mind. Integer overflow and underflow are right up there as threat vectors when writing anything that's touching hostile data.
I'm waiting for the day when Rust changes direction on this, and I think there's a chance this will happen because the alternative status quo of not checking arithmetic (at all) is just not tenable, at least not if we care about safety and security holistically, and not only memory safety.
You are correct; I should have written "memory safety".
> Zig treats memory safety not as an extreme-at-all-costs but as a spectrum
Zig needs to be more forthright about this.
When I first heard about Zig, I googled "zig vs rust" and found an article on the Ziglang website addressing that very topic:
https://ziglang.org/learn/why_zig_rust_d_cpp/
It completely fails to mention memory safety at all. That seems extremely dishonest, since memory safety is basically the "headline feature" of Rust (well, one of two or three at most). I wasted a lot of time digging through the Zig language manual ("so then how do they...") before concluding that something didn't add up. It definitely left a bad taste in my mouth.
> Rust won't protect you from all undefined behavior ... Rust has checked arithmetic off by default in safe builds
That didn't surprise me at all, nor will it surprise anybody who knows Java. Modular arithmetic is perfectly well-defined.
It's only C/C++ that picked the crazysauce option of decreeing that signed overflow is totally equivalent to scribbling all over random pieces of memory. It isn't overflow that's a security risk; it's languages that define overflow to be undefined in order to squeeze out a few piddly loop micro-optimizations. This becomes increasingly less beneficial in languages with iterators and no backward-compatible-with-C burden. Details (scroll to "Myth: overflow is undefined"):
https://huonw.github.io/blog/2016/04/myths-and-legends-about...
Yes (and thanks for the link!), I was in fact thinking more of this non-UB case (not signed overflow UB) as an example of where it's clearly defined as wraparound but can be chained into an exploit nevertheless, not technically UB but a vulnerability nevertheless. Not all exploits bother to go as far as a UAF. Unchecked arithmetic can be low hanging fruit.
> That didn't surprise me at all
It surprises me that Rust doesn't just enable checked arithmetic by default with an opt-out for performance, rather than enabling it by default for performance with an opt-out for safety. Zig's choice here is the safer choice from a security perspective.
It means that every arithmetic operation is potentially a branch/jump instruction. This wrecks a lot of pipelining/out-of-order-execution schemes.
I once worked on an exotic architecture where the integer types had a "NaN" value just like floating point numbers do; it had both modular and checked arithmetic, but the checked versions would return NaN instead of branching.
It also had 37-bit integers. Yes, 37-bit. Fun times.
You're right about the branching cost. I believe there's a better way to solve that than disabling checked arithmetic everywhere.
This comes out of something I learned working on TigerBeetle [1], a new distributed database that can process a million financial transactions per second.
We differentiate between the control plane (where we want crystal clear control flow and literally thousands of assertions, see NASA's "The Power of 10: Rules for Developing Safety-Critical Code") and the data plane (where the loops are hot).
There are few places where we wouldn't want checked arithmetic in TigerBeetle enabled by default. However, where the branch mispredict cost relative to the amount of data being checked is too high, Zig enables us to mark the block scope as ReleaseFast to disable checked arithmetic.
> It also had 37-bit integers. Yes, 37-bit. Fun times.
Wow, fun times indeed! We just disabled 32-bit support for TigerBeetle because it was getting too hard to reason about padding. I can't imagine 37-bit, LOL!
Suppose we have been given a 32KB data structure with some "step" bytes - in a conforming input these should always sum to less then 32768 and thus the total will easily fit in a 16-bit unsigned integer, so that's what our naive program does. Unfortunately attackers provided a structure whose step bytes sum to more than 65535...
Zig will panic here if using default arithmetic with default release builds. If the attacker wanted to cause a Denial of Service, job done already.
Rust will panic if explicitly told to enable checked arithmetic on release but it also provides explicit checked, wrapping, saturating and so on variants of the arithmetic operators if you want them for this part of your software (perhaps anticipating the risk) you can just have that without touching the behaviour of all other arithmetic in the program. 65530u16.checked_add(255u16) is None even in a default release build of Rust, what you do with that None (silently abandon this input? log the error?) is up to you and of course may not be adequately tested.
However, in WUFFS we simply can't write the erroneous program. It doesn't compile because WUFFS can't see why it's safe. Because it isn't safe. WUFFS requires the programmer to spell out what's going on, and so either you have to realise what might happen ("Oh, it can overflow, I should handle that") or choose a strategy that can't suffer the problem, ("Let's not sum up those steps, I see a different way to handle valid input").
Thanks! Great recommendation on WUFFS! And completely agreed, it's also easy to turn on checked arithmetic for Rust (if you know about it, but Rust definitely has an unsafe default there for those that don't, which is surprising to me).
At the same time, WUFFS is not always applicable, for example to writing something like a distributed system where you do still want safety, often the flip side of security. I'm sure you'll also agree it's good to balance out that security is more nuanced than just a rant about memory safety to the extreme. It's great to have positive discussions about languages, to evaluate trade-offs positively.
Counter-intuitively, I do feel also that Zig's explicitness as a language as a whole fits a security mindset well. For example, in `std/mem.zig` there's a very careful divExact assertion around underflow when calling `bytesAsSlice()`. This is just a fantastic way to prevent buffer bleeds, i.e. HeartBleed or CloudBleed, but it's probably uncommon to see in many libraries, and something like a borrow checker wouldn't provide this aspect of memory safety automatically. You can easily get lulled into a false sense of security.
From a security angle, I also like Zig's philosophy around very simple control flow and avoiding unnecessary abstractions, no matter if they're zero-cost. I think this is going to lead to a healthier package ecosystem when it arrives, compared to say NPM, where you get these dependency explosions that are a real headache for supply chain attacks. Attackers always go one level deeper, they attack through the basement, and there's often more low-hanging fruit at hand than a UAF (especially considering that many embedded systems that Zig targets probably do static allocation anyway, so bleeds might often be the worst that can happen). It will be interesting to see how Zig's philosophy around explicitness and avoiding bloat makes a difference here.
> Zig will panic here if using default arithmetic with default release builds. If the attacker wanted to cause a Denial of Service, job done already.
In the security world, a DoS is usually not treated as a P1. Perhaps a P3 at best (if you're lucky as a researcher!). For example, I've submitted one or two DoS MIME bomb samples that can shutdown Gmail servers and got very much an "okay, we'll just not bother about it because we're Gmail and our fleet is so massive". The DoS is probably still out in the wild for Gmail. Even ProtonMail, which has experienced numerous outages, didn't classify it as a P1, although they awarded it.
However, for a read/write exploit (running with the email example, perhaps a directory traversal in Apple Mail), having checked arithmetic convert what could have been a P1 into a P3 is actually exactly what you want because it prevents the exploit from going further (these things are almost always chained).
It also surfaces the bug visibly, you get a crash, you investigate, you fix. So from an attacker's perspective, they're actually less likely in fact to try and trigger it, because then they reveal they're in your system.
That recent Apple bug where they render PNGs incorrectly can (in principle) happen in WUFFS. The other recent Apple bug where bad guys seize control of your iPhone by sending a malicious image file cannot. One of these things is not like the other.
I think you're missing the point if you expect the borrow checker to care about buffer underflow. Rust has a runtime bounds check to check bounds, the borrow checker is, as its name suggests, checking the borrow rules. The trick (compared to arithmetic overflow) is that the optimiser can often push a bounds check outside a fast loop or eliminate it altogether, so you really can afford to do this in all or almost all your release code unlike checked arithmetic. WUFFS shows that you can do away with both of these runtime checks and be entirely safe if you're not interested in being a general purpose programming language. Which is (part of) why WUFFS gets to be both safer and faster. Both Zig and Rust are intended as general purpose languages.
I don't buy the "surfaces the bug" thing because I have too much experience of real world systems where there's so much noise and mayhem that you are focused on stuff that's causing your real users pain. Even if the DoS means the server falls over and must be manually restarted, the ticket in my queue says "Urgent: Auto-restart server. Watchdog maybe?" not "OMG bad guys are trying to break into our system somehow, find out how ASAP"
No, I was saying earlier that there are limits to WUFFS. The example I gave was that you can't write something like a distributed system (think consensus protocol like Viewstamped Replication, Raft or Paxos) in WUFFS, but where safety is nevertheless still critical, and where you reach that through crystal clear control flow and explicitness. In other words, safety is the other side of the coin to security. Hope it's a little more clear now.
> That recent Apple bug where they render PNGs incorrectly can (in principle) happen in WUFFS. The other recent Apple bug where bad guys seize control of your iPhone by sending a malicious image file cannot. One of these things is not like the other.
Of course.
> I think you're missing the point if you expect the borrow checker to care about buffer underflow.
No, I was stating the obvious, that it can't (or at least not always, but in some cases it can), not that it should.
> I don't buy the "surfaces the bug" thing
I was just trying to convey a little bit about how security works and how hackers (or at least red teamers) think, especially when blue teams are involved. I've found that the more I get into this, it becomes much less about preventing the breach and more about "assume breach, okay, now how do we detect it?". And a software DoS is also really just bottom-of-the-rung, you'll find almost no programs paying out for any findings. You shouldn't worry about them. Asserts are the safe thing to do. They close semantic gaps and make your code much more secure. It's like putting in a thousand trip wires, any thing off and an attacker can't get further. It completely shuts down exploit chaining.
Of course there are limits to WUFFS, that's why it isn't a general purpose language. You shouldn't implement these distributed protocols in it for the same reason toothpaste isn't a good engine lubricant, you deliberately can't even write "Hello, world" in WUFFS.
And yet, if you find yourself, in your distributed system, Wrangling Untrusted File Formats, you should reach for WUFFS to do that safely. Somewhere between "The device has a single button, it's green, press it" and "We process any PDF, HTML or XML documents sent to this email address" you will realise you need all the help you can get to Wrangle the data safely, and that's why WUFFS.
LOL, I would never have thought to do that till now! :)
I think we've always been on the same page regarding WUFFS and file format sanitizers. For me the question here really is, how do we improve the status quo when WUFFS is not an option? i.e. What are sane defaults for general purpose programming languages?
I still maintain that checked arithmetic should be enabled by default in general purpose programming languages, and that's because I believe in the principles behind WUFFS, having worked exactly on these kinds of tools myself.
(Neither does Rust.)
Your unsafe Rust is supposed to provide suitable constraint/ guarantees that you, the programmer, conclude it does not have any Undefined Behaviour. The language can't force you to do this, and at some point it becomes a social contract not a programming language feature.
I wrote the misfortunate crate to explore Rust's promise here. The crate provides legal but obviously inappropriate implementations of lots of safe Rust traits, and sure enough nothing blows up, there is no undefined behaviour.
The defined behaviour can be undesirable for example if you insist on putting a bunch of misfortunate::Maxwells in a HashSet you're going to have a bad time. Rust doesn't promise this is a good idea, it might cause infinite loops, memory leaks, all sorts of defined trouble, but it won't be Undefined Behaviour.
> Zig reference documentation badly needs examples. Can't figure out how to use std.fmt.parseInt.
While yes, Zig documentation badly needs examples, I'm not sure this particular criticism is justified. I would have thought that the usage of parseInt, was fairly obvious from the type-signature:
parseInt(comptime T: type, buf: []const u8, radix: u8) ParseIntError!T
Or, translated to C++ (and assuming the use of exceptions to return the error): template<typename T> T parseInt(const char buf[], uint8_t radix)
> Need to access myArray.items[idx] instead of myArray[idx]. I get it. But very unintuitive and requires knowledge of implementation details.I'm 50:50 on this one. While it might be nice to hide the implementation here and have a .item(n: usize) member function, ArrayList explicitly manages a contiguous region of memory, so I don't see the problem with exposing it as a slice as part of the interface.
> Should I pass the allocator to every function? Doesn't seem great. Maybe I'm supposed to create a global? Globals are evil and feel bad.
As a rule of thumb, libraries should take allocators as arguments to their functions, while applications can either do that or create a global allocator. There is absolutely nothing wrong with using a global allocator in an application; after all it's what almost every other language does. Zig just makes that global explicit.
Globals aren't always evil.
> Zig's inability to infer type is annoying. If I create var count and return it and the function return type is usize then var count is obviously a usize.
It's not. It could be any unsigned type of smaller size than usize.
I think it refers to Rust's ability to "indirectly" infer a type by it's later usage which can be somewhere entirely else in the function (for instance when it's used as return value), while Zig has a more direct type inference (but IMHO Rust's approach definitely isn't "obviously better", because sometimes one needs to read and understand the entire code of a function to get an idea what type a variable might be).
I understand what it is doing, but the fact that changing the return type of a method will change what previous code will compile to is something that I find extremely non obvious and surprising.
Inferring what the type is from the current code, sure. That saves me typing, but from what happens elsewhere looks like time travel.
I just always find it a PITA that Rust can't do global type inference and that I have to type most of the types for myself.
There's no coherent reason why it'd be OK for Rust to conclude that x is a u64 from
let x = some_u64.add(6); // the u64 type implement Add returning a u64
and yet not OK for Rust to conclude that x is a u64 from let x = someFunction(); // someFunction was defined to return u64
You had to explicitly change the return type of the method to cause yourself a surprise, Rust won't chase this rabbit through a warren, every individual function has to declare the types of its parameters and returns, so if you'd changed the function body not its declared return type, it could not have caused any change in type inference in other functions. let x = str.parse()?;
someFunc(x);
The function called here is dependent on what `someFunc` will accept.
That means that I can change what values will be accepted here by changing a completely different piece of code.With global type inference, the whole program is inferred for example.
let foo: T = bar.parse()?;
// or let foo = bar.parse::<u32>()?;
baz(foo);
In my experience I've never found this to be surprising behavior, or unclear.
It used to be that you'd get an allocator by taking a pointer to the allocator field (&gpa.allocator), in 0.9.0 that was changed to a function call (gpa.allocator()) and thus broke _a lot_ of things.
Exposing all these fields seems like it'll cause challenges when trying balance backwards compatibility vs breaking changes.
But I don't think that's the right level of abstraction to come from. ArrayList is just a wrapper around a slice that handles allocation and deallocation for you: the underlying data structure is part of the contract of the type, allowing it to be almost a drop in replacement.
That's less important than it sounds because the project is pretty explicit about not caring about backward compatibility till 1.0.
That being said, the state of the documentation is also the reason why I gave up on it for now. But I'm sure it will improve over time.
[1] https://github.com/ziglang/zig/blob/79628d48a4429818bddef2e8...
I think that's why ArrayList exposes it for you, so you can just use it as a regular slice everywhere. I, honestly, find that simple and liberating. I'm glad there's no other kind of accessor thing.
Continuous memory is just fast, and if you have some first class support for working with it, it makes no sense to duplicate the api. But again, if you want a hashmap you don't get the slice, for the slice doesn't make sense there...
Blog author here. In hindsight yes it's obvious.
I think my problem was a lack of understanding of comptime types. They're a little different from C++ / Rust since they're passed as args. Looking back I think I found the Zig docs less "sticky" than other languages. The concepts are familiar but just new enough I don't understand and don't fully remember them.
There's a point while working on AoC that I wasn't sure if Zig had first-class runtime types or not. It has comptime types and it also has a TypeInfo built-in with some degree of reflection? Is that runtime reflection maybe? I think partially? I honestly don't know.
Either way I was confused. :)
If you do that at comprime instead, the thing that runs at runtime is just memory offsets. You can for example generate code that parses json into 30 types you know at compile time, or you can use things like hashmaps and arraylists and whatnot to do it at runtime.
You get a lot of mileage by just doing the dynamic code away from runtime, and that's one of the benefits of comptime. And it's a slow thing to realize, also not necessarily something everyone cares about.
I do wish this was explained more prominently! The way this seems to work is that Zig has a lazy runtime for Zig code at compile time that does some interning (such that identical types obtained separately are equal, identical strings obtained separately are equal, etc.), and some types such as type can only exist in this runtime, not in the actual uh "run time" runtime. TypeInfo is for iterating over struct, union, or enum definitions at compile time to generate code. You'd use it if you wanted to write a generic json parser and serializer or if you wanted to write a type that's like an arraylist but internally uses a struct-of-arrays layout. Both of these are in the standard library and make informative reading.
TypeInfo itself is a union that can exist at runtime, but I don't think you can get into a situation where you have something at runtime and you don't know what type it is already. So actually using the TypeInfo at runtime may not be very useful.
> Zig appears to not report compiler errors for functions that get optimized out.
This is apparently caused by the laziness. The docs suggest using std.testing.refAllDecls(@This()) to have the compiler check unreachable code. I mean the actual usable docs at https://ziglang.org/documentation/master/#Nested-Container-T... rather than the automatically generated stdlib docs.
I also have to wonder about a fundamental misalignment of thinking when someone says downloading and replacing a single .exe is tedious, and isn't at all excited about 'Why Zig' list. If none of those things seem like compelling ideas to you then yeah, the language probably isn't for you.
If the tutorials aren't deliberately built up from the language's own defining philosophical fundamentals introduced singly and then in combination, then someone like me will just re-use the fundamentals I learned in whatever other languages I know that feel most similar to the new one. Most languages can be used in roughly parallel ways with just a change in syntax, even if that's not how they are supposed to be used, and if I (like most of us) have to get up and running ASAP in some new tool and don't have a quick on-ramp to establish the preferred patterns right from the start, I'll just have to re-use as much as I already know and just get on with it.
A lot of this seems to be a mismatch between people's expectations of zig's stability and the reality?
Similar to the process of getting the concepts of git when you've come from other version control systems.
(I have a horrible feeling that the solution to this is mostly going to be "lots of people writing tutorials that approach it in different ways until there's at least one tutorial out there that will work for (most values of) any given person", and having been there on projects of mine I sympathise with the unsatisfyingness of this conclusion)
const std = @import("std");
When I saw the second line and subsequent similar lines, I realized it was effing brilliant:
const os = std.os;
std and os here are names bound to types ... and you can have type variables and do compile-time manipulation and construction of types. This is confirmed in the manual when it talks about generic types being the return values of functions executed at compile time that take types (and/or other comptime values) as arguments -- while C++ templates are purportedly Turing complete, this is far more powerful in practice because it's vastly easier and more straightforward. Looking at code in the library like MultiArrayList--which implements AoS (array of structures) in the library rather than in the language--further confirms this.
So what's going on with those `std` and `os` bindings? The answer is given at the beginning of the Zig Language Reference (https://ziglang.org/documentation/master/) (which the OP apparently didn't read since it has a Hello World program that is an example of how to print): "The @import("std") function call creates a structure that represents the Zig Standard Library". So essentially, every source file represents an anonymous struct (all Zig structs are anonymous), the members of which are the top level declarations in the file--or rather, @import(filename) presents the file as such a struct, which can be assigned to a type variable like `std`. And `std.os` is in turn a type variable whose value is @import("os.zig") ... except that the actual value of `std.os` is computed based on the target machine. By turning files into comptime data structures containing all of the file's top level declarations, and having Zig code executable at comptime, immense power is achieved and one of the consequences of this is that zig running on any host is a complete cross compiler that can generate code for any target, using an appropriate target-specific version of the library. And it only took me a little bit of reading of docs and code to get my "aha" about how this works.
As a rank beginner, I had no problem learning these things ... they're mentioned repeatedly in material about Zig both from the project and from outside descriptions, the philosophy is the output of `zig zen` and in all the documentation, etc. The OP couldn't figure out how to print despite the Hello World program at the beginning of the Zig Reference Manual. I think most of the issues here are PEBKAC.
At least Rust, as blamed and loved as it is, delivered a stable compiler and people started working on the ecosystem (in the first years, most packages were working only on nightly, but at least there were crates available). The ecosystem for zig is insignificant now and a stable release would help the language.
[1] https://github.com/ziglang/zig/issues/234 [2] https://about.sourcegraph.com/podcast/andrew-kelley/
Zig was started in 2015. [2]
[1] https://en.wikipedia.org/wiki/Rust_(programming_language) [2] https://en.wikipedia.org/wiki/Zig_(programming_language)
http://smallcultfollowing.com/babysteps/blog/2012/11/18/imag...
The whole language got rebooted shortly after the blog post above, mostly because the borrow checker made so many other things suddenly unnecessary or trivial. What we call Rust today is at most 9 years old, and any similarities to pre-2013 Rust are strictly superficial syntax. They share a name and some syntax, sort of like Java and Javascript do.
Zig today at T+7 is not where Rust was in 2020 at T+7.
Compiling is faster than rust and C a lot of the times
We have packages, and a good few, thing is, this is no rust big, we don't have mozilla nor to backup and work into it.
I don't think is overpromise, Vlang is overpromise, zig atm, is slowly getting there, no promises on when
Have you, like, seen the release notes for 0.9.0?
https://ziglang.org/download/0.9.0/release-notes.html
> Zig still can't proper handle UTF-8 strings [1] in 2022
There's plenty of discussion on the subject in basically every HN thread about Zig: the stdlib has some utf8 and wtf validation code, ziglyph implements the full unicode spec.
https://github.com/jecolon/ziglyph
You might not like how it's done, but its factually incorrect to state that Zig can't handle unicode.
> In a `recent` interview[2], he claims that Zig is faster than C and Rust, but he refers to extremely short benchmarking that has almost no value in the real world.
From my reddit reply to this same topic:
This podcast interview might not be the best showcase of the practical implications of Zig's take on safety and performance. If you want something with more meat, I highly recommend Andrew's recent talk from Handmade Seattle, where he shows the work being done on the Zig self-hosted compiler.
https://media.handmade-seattle.com/practical-data-oriented-d...
Lots of bit fiddling that can't be fully proven safe statically, but then you get a compiler capable of compiling Zig code stupidly fast, and that's even without factoring in incremental compilation with in-place binary patching, with which we're aiming for sub-millisecond rebuilds of arbitrarily large projects.
> The ecosystem for zig is insignificant now and a stable release would help the language.
I hope you don't mind if we don't take this advice, given the overall tone of your post.
That sounds great! But at the same time people in other threads here are talking about 1-3 second compilation times for Advent of Code solutions (which I presume are smallish). Can you summarise where that really fast compiler comes from, to save me searching through that talk video? Is this something that everyday users will be able to use in typical workflows?
https://kristoff.it/blog/zig-new-relationship-llvm/
Long story short, we're currently working on a self-hosted implementation of the compiler and what people are using now is the old C++ implementation. As soon as the new compiler is feature-complete enough, we'll start shipping it and we expect much better compilation speeds, which will be even greater speed for debug builds once the native (i.e., non-llvm) backends catch up as well.
Latest progress update on this work: https://twitter.com/andy_kelley/status/1481862781380874240?s...
I'm sorry, what do you mean by this?
- latin1 is dead and should be in no stdlib in 2022 - uppercasing requires the current Unicode tables, so, a largish moving target that you probably don't want to embed in small programs.
In this environment you might very well not need actual uppercase/ lowercase but only the ASCII subset. Accordingly Rust provides that too, which is far less to carry around than the Unicode case rules. Since the ASCII case change can always be performed in situ (if you can modify the data) Rust provides that too if it's what you want.
My biggest problem with your comment is that it is completely and utterly false.
>At least Rust, as blamed and loved as it is, delivered a stable compiler
After MANY years and numerous complete redesigns.
I found zig language reference pretty good[0]. It is simple, lot of example. I would find 80% on there and had to google the rest. And as another commenter said, looking at zig source code is actually not a bad idea. The std lib is pretty clear and with comments.
My biggest beef with zig is a lack of a package manager. But apparently, it is high in the list of priority for the author so...
But if you hit a roadblock it's hard to track down more info. Some of that is due to the lack of adoption, though.
The comptime machinery is really cool, but it gives the compiler much less information to work with for producing good errors.
It'll be interesting to see how this plays out as the ecosystem grows.
Why do you think that's the case?
"I also think it's partially wrong. No one in the history of the world has ever been confused or upset by a + b calling a function."
It depends. If this is simple math on vectors I think it can be OK but it should probably be a built-in feature of the language as this is common, solved and we all implement it the same way (for short vectors at least)
But the + operator has been abused in the past, especially with strings concatenation and I think this is a huge liability. Such an innocent looking operator, the simplest of all operations, leading to a function call, a memory allocation and thus a very real potential memory leak, all of that hidden from the eyes...
In my opinion, clarity should have priority over anything else. If an operation is computationally complex (especially with side effects) it should be at least hinted to the reader by a function call.
However I think in general purpose programming we lost that battle when Java special-cased the operator. There's no overloading in Java, you can't have + do the Right Thing™ in Java for your 3D vector class, but it does concatenate strings because people had begun to expect that.
Also, while I'm in this topic, the author argues they don't want anybody overloading either of the member access operators (which is something Python kinda-sorta supports) but notice that Rust effectively does this all the time and nobody freaks out because it feels completely natural.
Your Box<Thing> isn't a Thing, so, why can you call Thing methods on it? Because it is transparently passing those into the underlying Thing by implementing the relevant overload features from core::ops and all of Rust's smart pointers work the same way.
misfortunate::Double shows that this gives very strange behaviour if somebody uses it inappropriately, a Double<Thing> is actually two Things inside, but when you change it, you're changing one of them, while any immutable references are to the other one... an eerie experience.
It was already special-cased by Pascal.
I think that it's possible if say Gosling hated + concatenating and had provided Java with a different String concatenate (it clearly wants a concatenate operator, but it needn't be named +) we'd see that popularised and while some languages with overloading might overload + it wouldn't be ubiquitous. I can't prove that of course, it's purely my opinion.
Even BASIC uses it, and it was everywhere during the 15 years that preceded Java.
Let alone all the other languages since Jovial that can't be bothered to dig out just to prove my point.
I stand corrected.
I hadn't considered the memory leak angle of this, and now that you mention it it's clearly a driving concern here. When freeing is explicit, it's a problem to heap-allocate temporaries. (And I guess the addition operator would also require a third operand to supply the allocator?)
I'm still a big fan of destructors, and how they make this problem mostly vanish. But I understand that it's hard to make them efficient without move semantics, and maybe also a global allocator so every object doesn't need to store an allocator pointer? What other interactions am I missing?
In my case (I use my own framework) I use a lot of makeStringWith* functions that all allocate in the same global scratch buffer (one per thread) and I don't care at all about releasing the memory.
There is simply one function to clear reset all scratch buffers, it has to be called explicitely by the user, most of the time once per frame (I create video games) but it can also be when the current job is done or never.
From my perspective it is very efficient and reliable, never had to solve a bug related to this mechanism.
It's not a given that floating-point addition is a simpler operation than string concatenation.
If you're serious about what you write, then all operators should be banned.
Integer division or multiplication of long (64bits) can also force the compiler to inline a bit more code or even to call a function on older architecture, hardware division was not a given on ARM before ARMv6 I think, so for example on the Gameboy Advance or even the Nintendo DSi you had to be careful with division etc.
But again, this will only slow down your code if you're not careful, not generating memory leaks silently in the background.
In the simplest case, you can just store all strings on the stack.
Things are less primitive these days but the memory still brings a smile to my face.
(I've yet to conclude if stealing ++ for this rather than preinc/postinc was a terrible mistake, so far in context it hasn't seemed to be but I'm still keeping an eye on the question)
Nothing says that you cannot use both infix "++" for concatenation and unary ++ for increment. "-" is both a unary prefix and binary infix operator, for example.
Nothing says that I can't do that but in context of deliberately -not- re-using operators for completely different things it would seem rather self-defeating.
Nonetheless I had gotten the impression from many posts on here before that Zig was a lot closer to finished than it really is, it seems like they’re not quite half way to something as polished as Rust 1.0. It is probably unfair to judge the language against Rust in its current state. Hopefully in a few years they will have figured out the memory safety problems and the stdlib documentation will stabilize more.
You're never deprived of control though, that's what unsafe is for, and it provides roughly the same level of safety as C or Zig would.
[0]: the compiler is allowed to emit extra reads/writes to references so a mut and shared one to the same memory location cant be alive at the same time, across threads, and across sometimes non-obvious scoping rules.
[1]: [pointer::set_ptr_value](https://doc.rust-lang.org/std/primitive.pointer.html#method....) is used by containers like Arc for `from_raw` in order to carry over aliasing, mutability, and other location meta data.
unsafe Rust is also less ergonomic when dealing with concepts that go against the borrow checker like Pinning for intrusive memory, ptr::read/write and ManuallyDrop for controlling unorthodox object ownership/lifetime, and MaybeUninit/NonNull for facilitating certain data layout/access optimizations. Such designs often can't be wrapped in safe-Rust without introducing runtime overhead the patterns were used to originally avoid. Languages like Zig and C however make these patterns natural or event pleasant enough to consider it over unsafe Rust.
Rust's value proposition is simple: no GC and no undefined behavior. Period. Nothing else has that.
Of course the swell is early, but waves are what technology is about, and the surfers are there and paddling out. It's a great time to be getting involved, especially for greenfield projects that have some time in themselves to reach stability and don't want to pay a language compiler/complexity tax for the rest of the project's lifetime.
You could also throw a dart blindfolded into the Zig community and be pretty sure to hit some seriously talented programmers to learn from. If you're investing in a deep understanding of the language now, I'm pretty sure it will pay off down the line.
https://github.com/bwickman97/ffmalloc
The point of C and Zig is that they are low level and you can do whatever, like not use an allocator, or write an allocator.
However a libc is a thing that zig probably wants and something I am also thinking on doing if I have the time
Much easier to just use C+.
I mean, as far as achieving it at compile time, I would say it remains to be seen whether Zig's approach is one of these ways. There are certainly other known approaches, like the way ATS models pointers as proof objects, but these are also fairly abstract.
> Zig is C, it's not meant to abstract away memory management.
This is just hiding the ball though. Why don't we want abstractions on memory management? There is no runtime cost to providing the abstraction that Rust does. As such, "machine-oriented" is an ambiguous description of the difference here. The whole lesson of Rust is that there's nothing fundamental about the abstractions the machine itself provides, and we can create better abstractions without losing our orientation to those concerns.
This is technically true, but if I want to write a program that does not call malloc in its steady state, I can't use the stdlib or any library that uses the stdlib.
https://kevinlynagh.com/rust-zig/
https://zig.news/aransentin/analysis-of-the-overhead-of-a-mi...
The point of languages like C and Zig is that they are only a bit higher level than assembly for portability, but otherwise they don't hinder you to do whatever. It's up to you to solve problems like memory safety. You might not even have a memory safety problem because you don't have memory in the first place, or your use case makes memory safety trivial.
> but otherwise they don't hinder you to do whatever
Neither does Rust! You can choose to ignore what the compiler wants you to do and just program with unsafe everywhere. The difference is that you don't have the option of writing safe code in Zig and C.
> It's up to you to solve problems like memory safety.
And so far, memory safety is not actually a problem that programmers have been able to solve by sheer force of will without borrow checkers or garbage collection. For a subset of trivial programs, sure. But not at any reasonable scale.
Not other ways. One other way. Just one. Garbage collection.
It just happens to have a better syntax for the C crowd.
I may be a bit biased, but many of the problems mentioned in the post are the sorts of things I haven't dealt with in Nim for probably 2 years. Nim's error messages are still hit-or-miss and {.gcsafe.} still haunts me in my dreams, but the stdlib and language documents are great, and most of the annoying gotchas have been fixed.
It's really that simple. There is no international conspiracy behind Rust's popularity. It's the only thing that can do what it does in that respect.
For example, to get the conversation started, how do both languages compare in terms of checked integer arithmetic? Do they enable checked arithmetic by default in safe builds with an opt-out for performance, or do they leave default builds unsafe with an opt-in for safety?
Another example, JavaScript is 100% memory safe, does it follow that we can expect to see less exploits against NPM than C? Both ecosystems are massive, but I'd wager that most product release security teams are more stressed out about NPM than C dependencies right now. Not to say that they shouldn't be running C dependencies in sandboxes or be evaluating the risk of C (and remember that Zig is an order of magnitude safer than C, much closer to Rust actually, and more so in some areas). But NPM is probably getting more attention. Same thing for bug bounty issues reported. Probably more for NPM supply chain vulnerabilities than anything else. All 100% memory safety, and yet security is still a thing.
It's the many small decisions like these, along with thinking of security not as a binary extreme but as a probabilistic spectrum, that are more interesting to me.
And nobody has ever claimed that memory safety is the only thing that matters with security, but it’s definitely high on the list. Can you imagine how much more of a nightmare NPM would be if JavaScript weren’t memory safe? This is the relevant counterfactual.
> Zig is an order of magnitude safer than C, much closer to Rust actually, and more so in some areas
Hard disagree. It’s definitely safer than C, (in terms of UB) but you can tell it’s not even close to Rust on a number of axes (yet?): https://scattered-thoughts.net/writing/how-safe-is-zig/
As for integer overflows, I don’t think they are nearly as big a security concern in a memory safe language with bounds checking. Feel free to correct me though.
See HeartBleed, CloudBleed, all buffer underflows resulting in buffer bleeds, letting someone read all your sensitive server memory, with no UAF.
They can also be caused just by integer overflow, which is what makes unchecked integer arithmetic so incredibly dangerous. It's easy for programmers to be oblivious to this, and overly rely on the borrow checker, thinking it can provide 100% memory safety, when by definition it can't.
And all kinds of software systems have these vulnerabilities. For example, I've worked on static analysis security software that could detect bleeds automatically in outgoing email attachments and it would find different bank systems leaking data in autogenerated statements.
From this experience, and from some bug bounty and security engagements I've done, I'm much more comfortable actually with Zig's approach to correctness and safety overall. I think the borrow checker is pretty awesome and has some serious muscle, but nevertheless Zig impresses me with its strict focus on explicitness, which I believe is the best approach still to eliminate these kinds of semantic gaps in general.
The borrow checker obviously can't protect you from bleeds (it can from some where static allocation is at play), but checked arithmetic would, or at least go pretty far — and again, it's about the spectrum, not the extreme.
It would be great for Rust to enable checked arithmetic by default for safety, with an opt-out for performance. Flipping this around would be a better default.
The implementation would probably be ugly, but I wonder if it could be implemented by using a comptime string to represent the operation, e.g. something like:
fn doMath(comptime op: []const u8, args: anytype) MathReturnType(op, args) {
// TODO: implement me
}
const result = doMath(
\\a + b
,
.{ .a = a, .b = b }
);
Where the implementation would call `.add` etc on the parameters when infix operators were used.The gist is that you can easily build them to preclude useful optimizations and efficient execution, when it's often more desirable to be fast than to have syntactic sugar, hence having explicit function calls like multiply_add(a, b, c) instead of a+b*c. If you really want syntactic sugar when it comes to math, operator overloading probably isn't the way to implement it, it'd be nicer to have something with the full context so there can be optimizing reductions. Lisp macros can do that, or you might have some other kind of parser (that might have to work on strings), or with sufficient cleverness you could build an overloaded operator nest full of context-accumulating operations-to-perform that either require some doMath wrapper at the end or a final overload of operations producing a fully computed return type.
I prefer languages that don't cripple expressive freedom and so overall I'm not anti-operator-overloading in general even if I think some overloads are pretty questionable (I dislike C++'s arrow overload for Optionals) but I no longer think that e.g. a math-focused library is an obvious win or exception to the downsides of the expressive power granted from operator overloading.
So even if zig allowed for operator overloading, those issues would have to be solved.
And I don’t believe I said they were just as clear, just that it’s no big deal to get used to and gave an example where overloading muddies the waters.
Do you have any examples of situations where people would be writing a lot of code with custom operators where each individual block is pretty small and having to deal with a DSL of some flavor would be especially burdensome?
As someone with only 18 days more experience than OP it’s a little silly for me to say, but I think OP is getting put off the the normal learning curve of a new language and standard library. The struggle to learn different patterns faded away after not too many more days. Zig is pretty simple, and while the standard library is barely documented, it is pretty well organized and straightforward to understand.
I’m pretty interested to see where zig goes, and I’m hopeful about it.
(As mentioned, zig is a work-in-progress. The final form might be significantly different that what is there now,)
That sounds rather like a different thing was wrong - not where the author tried to lean: Zig has crap documentation.
Zig is young. They're spending their time maturing the core ecosystem --- the compiler, the standard library, and so forth. It's perfectly reasonable to expect early adopters to "RTFS/read the fucking source". *When* they start writing documentation, I'd be perfectly happy to hear comments like "The zig documentation is crap".
At this stage of the project, I'll be honest, I find it upsetting to read comments like this.
Right now, Zig is young and the documentation is kinda ok (Language reference is clear enough, outside a few things that need better visibility, for example, explaining concepts that are only on the examples can be lost easy), the autogenerated std on the other hand used to work meh, until a commit broke it afaik, so rn is in a pityfull state, but there is no use in reworking it until stage 2, so that's that
Well, crap as an offering (what you get atm), not necessarily as in the quality of what little it is. It does though have a reference and documentation page:
These docs are experimental. Progress depends on the self-hosted compiler, consider reading the stdlib source in the meantime.
It also links to a page that explains how the standard library is structured:
https://github.com/ziglang/zig/wiki/How-to-read-the-standard...For auditing you are right, of course.
I mean, as a first note one of the ideas here is reading code already written in zig to understand how to write code in zig - so in that case you're effectively inferring the interface from what the authors of the code itself are treating said interface to be.
But also, interface documentation is pretty much never truly complete because there are almost always some implicit assumptions involved (and if you try and make all of those explicit you rapidly end up with documentation that's so verbose people's eyes glaze over when they try to read it so their model often ends up incomplete anyway so how explicit to be is itself a trade-off).
Then zig embeds its test cases in the source file, so you can look at what the authors have explicitly declared -must- work to help you know what the interface is intended to be.
Plus when I'm source diving, I've done enough of it over the years to be able to at least attempt to build up a model of not just the implementation, but of the author's mental model as they were implementing it, and if you can figure out their intent, it's much easier to guess what their code is -meant- to do and thereby what it will hopefully continue to do into the future.
If in doubt, though, leaving a comment in your own code as to what assumption you're making -and- writing a unit test in your own code that verifies the assumption continues to hold will mean that at least if it doesn't you'll see a test failure in your own suite that tells you it changed.
(or: "code to the interface, not the implementation" is absolutely the right thing to aim for, but in practice the line between the two is fuzzy and cases where you have to make a judgement call will always show up eventually)
I feel old, I know that any time is an opportunity to get distracted but 3 seconds doesn't strike me as a long compile time.
Or is that a typo for 30, which would make more sense, which is definitely long enough to be a frustration?
Zig comptime is the part that takes a bit more time compiling for being an interpreted version of the language, and the more comptime you use, the more time it takes (for a bit, even comptime heavy code can gain just a few seconds out of it, of course can gain 30 minutes, but you would be making like, weird jumps and generating strings on comptime that gets used on runtime, or an expensive algorithm that will take a long time to execute)
I use zig anyway.
$ time zig build-exe ./a.zig -O ReleaseSmall
real 0m1.728s
user 0m1.544s
sys 0m0.292s
Which is actually slower than running the entire solution in debug mode: $ time zig run a.zig -- input.txt
min 337 342641
min 470 93006301
real 0m1.138s
user 0m0.789s
sys 0m0.270s
Totally possible that it's much slower on the dev's machine, but if he's a VR developer it seems weird that he'd have a low-powered box.[1]: https://github.com/llimllib/personal_code/blob/master/misc/a...
Maybe this is a Windows issue? If I run "zig build" and then immediately run "zig build" again it still takes 3 seconds.
Speaking as somebody who mostly codes for *n?x but has some windows users of their libraries, I've run face first into "everything I know about build optimisation is inapplicable on windows" more than once.
('everything' is admittedly slightly hyperbolic but I'm sufficiently bad at windows tooling that it always feels that way)
Nuked zig-cache and zig-out, rebuild, no change in build times.
I think that's why he would like a `zig check` command that would give him these errors and warnings faster.
(or if you care only for one file) zig ast-check =)
For compile errors I think zig build is the only way
Personally, while I find the imperative paradigm more intuitive than the functional paradigm, I believe functional type systems are simpler and more "right" than imperative ones are (which have been taken over by OOP). There are multi-paradigm languages, but these have large numbers of features, and complicated type systems, and I want something simple. Is there interest in this? Thanks.
By SML type system, I'm thinking:
- Algebraic data types
- Anonymous functions
- Pattern matching
- Simple generics
- Type inference
- Simple modules and implementation hiding
combined with some approach to operator overloading (which SML doesn't consider at all). And to remind you, the paradigm should be imperative, not functional.I'm thinking a use case might be in numerical linear algebra, but with support for multiple different scalar types including floats, ints, complex floats, dual numbers (for autodiff), "codual numbers" (for better autodiff), bignums, quaternions, symbolic algebra, etc. The functional paradigm is not suitable here, but Julia is perhaps overcomplicated.
Zig is far from a finished project so it's premature to complain about most of the things you are complaining about.
That depends on `a`, `b` and the function - I have seen way too many `+` that weren't commutative and/or associative. Languages which do not allow to use _other_ symbols like `⊕` too almost always lead to confusing behavior of 'usual' operators.
And what does `a * b` do for vectors `a` and `b`? Is it the 'usual' 'dot product', is it component wise multiplication (`[.., ai * bi, ..]`) is it 'matrix multiplication' (`a * b^t`), is it the cross product (because the language doesn't allow to define `×` as operator), is it ...
Actually, in C and C++ the implicit conversion and promotion rules are footguns enough.
Of course it is, but you can use other functions names than `product`. If the language doesn't let you use any other symbol for the operator but `*`, you're out of luck.
No, as IEEE754 doesn't even guarantee that `a + b` == `a + b`. You always need to compare the absolute difference of two values against an ε.
So
a == b <=> |a - b| < ε
But for sensible comparison of floats, the addition is commutative.What? I'm pretty sure that is plain incorrect. Can you back that claim up with references?
See for example:
int main() {
double q;
q = 3.0/7.0;
if (q == 3.0/7.0) printf("Equal\n");
else printf("Not Equal\n");
return 0;
}
On an extended-based system, even though the expression 3.0/7.0 has
type double, the quotient will be computed in a register in extended
double format, and thus in the default mode, it will be rounded to
extended double precision. When the resulting value is assigned to the
variable q, however, it may then be stored in memory, and since q is
declared double, the value will be rounded to double precision. In the
next line, the expression 3.0/7.0 may again be evaluated in extended
precision yielding a result that differs from the double precision
value stored in q, causing the program to print "Not Equal". Of course,
other outcomes are possible, too: the compiler could decide to store
and thus round the value of the expression 3.0/7.0 in the second line
before comparing it with q, or it could keep q in a register in
extended precision without storing it. An optimizing compiler might
evaluate the expression 3.0/7.0 at compile time, perhaps in double
precision or perhaps in extended double precision. (With one x86
compiler, the program prints "Equal" when compiled with optimization
and "Not Equal" when compiled for debugging.) Finally, some compilers
for extended-based systems automatically change the rounding precision
mode to cause operations producing results in registers to round those
results to single or double precision, albeit possibly with a wider
range. Thus, on these systems, we can't predict the behavior of the
program simply by reading its source code and applying a basic
understanding of IEEE 754 arithmetic.
The relevant conclusion: Neither can we accuse the hardware or the compiler of failing to
provide an IEEE 754 compliant environment; the hardware has delivered a correctly rounded result to
each destination, as it is required to do, and the compiler has
assigned some intermediate results to destinations that are beyond the
user's control, as it is allowed to do.
https://grouper.ieee.org/groups/msc/ANSI_IEEE-Std-754-2019/b...> Implementations shall provide the following formatOf general-computational operations, for destinations of
> all supported arithmetic formats, and, for each destination format, for operands of all supported arithmetic
> formats with the same radix as the destination format
Which to my interpretation sounds like providing a `division(double, double) -> double` operation is required. I suppose the argument could be that in C the way of invoking that specific operation would be to add an explicit cast, i.e.
(double)(a op b)
But I do think that this is more a quirk on how operators are done specifically in C and not a general matter; the actual IEEE 758 operations are consistent and don't have such surprises.So I guess you were technically correct that IEEE 758 does not guarantee `a + b == a + b`, because IEEE 758 does not specify `+` (or `==` for that matter) operators at all. What IEEE 758 does guarantee is that you get the same result for the same operation.
In terms of C language, it is maybe interesting question if `FP_CONTRACT` pragma influences the result here at all:
> A floating expression may be contracted, that is, evaluated as though it were a single opera-
> tion, thereby omitting rounding errors implied by the source code and the expression evalua-
> tion method. The FP_CONTRACT pragma in <math.h> provides a way to disallow contracted
> expressions. Otherwise, whether and how expressions are contracted is implementation-defined.
Are the comparison and arithmetic operations considered to be contracted in these sort of situations?
That has nothing to do with C, but with the CPU and its registers. x86 has 80bit floating point registers (too), so if the compiler/CPU (doesn't matter which language) saves a value in a 80 bit floating point register and moves it from there to memory where it is stored in 64 bits, the number gets rounded (not truncated, it's actually converted).
See also the GCC page:
For instance, in the following code segment, depending on the compilation
flags and numbers and calculations used to find tmp, the following code may
print out that the values are different:
double tmp, X[2];
tmp = ....
tmp += ....
...
X[0] = tmp;
if (X[0] == tmp)
printf("Values are the same, as expected!\n");
else
printf("Values are different!\n");
This is because tmp will typically be moved to a register during register
assignment, which means tmp may hold a full 80 bits of accuracy, some of
which are lost in the store to X[0], and thus the numbers are no longer
equal. You may workaround this problem by always explicitly storing to
memory to force the round-down.
https://gcc.gnu.org/wiki/x87note> Are the comparison and arithmetic operations considered to be contracted in these sort of situations?
No. Or well, maybe. 'Contracted' means using e.g. FMA (fused multiply and add), so addition and multiplication like `ab + c` is done in a single step instead of two. Which means that the result is only rounded once (`ab + c` is rounded), whereas without this fusing/contracting `ab` would be rounded and `ab + c` would be rounded again (actually depending on the optimization/compiler flags). So the results (may) differ.
There are too many posts that focus on successes and embellish the state of affairs.
That said nothing was too surprising about a language that hasn't hit 1.0 yet.
> Lack of zig-analyzer makes learning hard.
> zig fmt src/main.zig is nice. Wish it automatically ran on all files.
I also did (well, "am doing", can only work a bit each day and am plugging through day 7 right now) AdventOfCode in Zig this year.
These points here didn't resonate with me at all. I wonder if the author knew about or tried ZLS[0]. I had it on and integrated with my VSCode and it would check a lot of things as I went and format on save. I think I followed something like this[1] to set it up.
[0] https://github.com/zigtools/zls [1] https://zig.news/jarredsumner/setting-up-visual-studio-code-...
> Zig is trying to be a better C
I recently started learning Rust by adding snippets to The Quick Snippet Reference [0]. Without any previous knowledge of the language I knew there had to exist a data type for dynamic arrays. After several Stack Overflow searches I found it: Vec. It was very interesting to compare to Python's Array. Getting to understand the drain() method was a refreshing experience. I'm still hesitating about the best way to handle errors (Rust has no exceptions, just panic!), but that will probably become clearer when working with threads. And Rust compiler's warnings and errors are really nice. I'll certainly give Zig a try.
[0]: https://github.com/snippetfinder/The-Quick-Snippet-Reference
Yeah in some way they are the same kind of reference. I already knew about Rosetta Code, explored their snippets and found them valuable as a learning tool by curiosity. But The Quick Snippet Reference presents examples to do the smallest tasks that we would need when programming. The tree of these minimal operations is shared between languages, so it's easier to compare, or find how to do something in a new language that you already use in another.
Should I cope with this? If so, do so. This is likely if you're a library and unlikely in some application software, especially AoC. Don't ask me to try calling fooA() and then call fooB() if it fails, just have that code inside fooA() already.
If not, might my caller want to know what went wrong? In a library your users might want to cope, unless you're sure the situation is fatal and all hope is lost, so always return Result and use appropriate Errors for error cases, the ? operator can help you pass on errors from other libraries.
Otherwise panic.
Because Rust's Result doesn't change control flow, you don't need to decide how callers should cope with Errors, only flag that it's their problem not yours. Because Result doesn't have a non-Error result in it if you returned an Error, your caller can't mistakenly carry on anyway, they'll just panic if they try to do that.
I think this is a fair post overall given what it is: a log of what it was like for one specific person to use Zig for AoC for the first time.
I'll spend some words on some of the things mentioned and then add some more useful advice for forrestthewoods at the end.
From my perspective the complaints mostly were about now knowing things which had varying degrees of discoverability.
The first point in Day 1 about mixing optionals and error unions is an example of something very easy to discover (by reading the language reference), so I personally would consider it an unlucky mistake by the author.
The second point (how to print an integer) is about something way less easily discoverable than in the previous example. You need to learn about how to print, then about `fmt` and from there the only authoritative place that contains information about format specifiers is the doc comment of `std.fmt.format`.
This second example is the kind of thing that we're the weakest at communicating to people because it's not a feature of the language itself (so it doesn't belong to the language reference) and, while you could in theory find this stuff in the stdlib autogenerated docs, you still have the problem of having to discover first what to search for. This is especially problematic when Zig breaks away from what's common in programming languages. One example of that is the fact that Zig has no print statement.
I've given some talks and written some articles about how to do common things and find stuff in the standard library:
https://www.youtube.com/c/ZigSHOWTIME
and, as the author metioned, Zig Learn is by far the closest thing we have to a "Zig book".
The problem with creating content for the standard library is that it keeps changing and it's not yet the focus of development. As the language matures more content is being created (just look at how many explainers are being posted to zig.news alone) but we're not yet in a position to invest into something like a book.
The author also mentioned issues with the autogenerated docs for the standard library. Those docs are currently incomplete and in fact greet you with this message as soon as you open them:
These docs are experimental. Progress depends on the self-hosted compiler, consider reading the stdlib source in the meantime.
I've seen some comments here about how recommending to read the source code is unhelpful. I vehemently disagree because of practicality first (if something is not documented elsewhere, then that's the best you can do) and second because reading the source code should not be considered something primitive that developers used to do before discovering fire.We all should read more and write less.
That said, Andrew is going to help me start this week the work on a new doc autogeneration system based on a different design that the current (incomplete) one. I'll do most of the work while streaming on Twitch so if anybody has opinions, complaints or questions, you know where to find me: https://twitch.tv/kristoff_it
Going back to the post, there were also some impressions that we all fully agree with (and are working on improving), a good example being
> Bit-shifting is a monumental pain in the ass.
Yes, it indeed is. See: https://github.com/ziglang/zig/issues/7605
Ok, so, here's my advice to forrestthewoods if you're ever going to try Zig again:
[1] Interact more with the community if you want to be pointed quickly in the right direction. You mentioned the discord server so you already interacted with people. Leverage them more: ask questions in #zig-help and don't be afraid to ask why things are a certain way. You'll probably be way less annoyed by, say, ArrayList requiring you to access its `items` field to iterate, if somebody explains to you that ArrayList is a normal userland struct defined in the stdlib and that in Zig there are no "magic methods" that the compiler picks up to do iteration etc. Or maybe you won't be less annoyed, but at least you will know if something is intentional or if it's just something that we haven't gotten around fixing yet.
[2] Try to learn Zig doing something else other than Advent of Code. AoC doesn't really let Zig show it's potential because it doesn't ask you to do good software engineering. You're just supposed to write a script that gets you to the correct answer. Try instead doing a small project where you have to validate input, produce useful error messages to the user, clean up resources properly (sockets, memory, ...), etc. You will find that Zig will be much better in this second case. I'm saying this based on first hand experience btw, I've streamed the first 16 days of AoC from this year, you can find the recordings in this playlist: https://www.youtube.com/watch?v=wo580tbLSR8&list=PL5AY2Vv6Es...
I stopped after day 16 because I felt it the exercises didn't really help me showcase any big strength of Zig, so I started working on other stuff instead.
I don't know if ZSF can guide this or if we can incept it into Andy's head, but if there were any coordinated plan to "solidify certain parts of the stdlib sooner than other parts" the std.fmt.format would be high on my list. And document it. Hell. document it, even if it changes around all the time. I'm often forgetting how that beautiful bastard works and having to jump to the comment in the stdlib is kind of annoying.
To be less dramatic: I do intend to wildly break formatted printing at some point.
Point taken though. Perhaps it deserves its own seat in the roadmap.
Does that mean that someone at his office had better luck and should be writing the success article? Or something else?
That's a bit overambitious maybe.
[0]: https://vlang.io/
To me, this seems like one of those projects that starts its marketing hype too early, gets a negative reputation and then never recovers even if the original promises have been fulfilled.
It was often around the difficulty of the exponential problems, but most were okay if you're very familiar with grid data, hierarchy traversal and caching.
Then when those design choices get hardened into a 1.0 it will become just-another-language and we will inevitably feel let down and move into the next attempt at a Perfect Programming Language. So it goes.
Incremental examples is where it's at. You should document a fonction/type/feature/whatever with a short example with basic operations and then a more involved example with good practices and potential caveats. I know this is a lot of work, I understand it can't be like that everywhere.
This is not a dig at anyone's project or language, but "read the source code" comments always strike me as unhelpful.
It helps you to see what are the idioms, what does optimized production code look like.
Compare C++ and Go std lib and it’s easy to constate the difference in language goals that they have.
1. It's useful when you want to check the exact under-the-hood behavior, to figure out side-effects or corner-case semantics.
2. After a good number of years, that code becomes less scary and more icky. Some of it remains kind of scary though...
3. Reading that code teaches you a lot about considerations you may be ignoring, especially your implicit platform-dependent assumptions.
4. Nitpick: There isn't any single C++ standard library source code, it's at least 3 popular ones (GCC's libstdc++, LLVM's libc++, and Microsoft's)
If you're a newbie, why are you learning an experimental language?
The zig standard library does the following to help me:
- Logical, straightforward code, no magic incantations
- Well named variables and functions
- Comments
- Folder structure is easy to navigate
- Easy to find on github - not hidden away
Obviously your mileage may vary.
Forcing people to read the code puts pressure on the code to be clean.
Not writing docs against an unstable API is a strategy to make it freer for the authors to change the API, possibly in sweeping breaking ways (for example allocgate).
Zig is 6 years old. Rust had much better documentation at this stage of development.
To be fair though: Rust also had corporate backing and a larger fan base contributing to the book and std docs.
Yes, and language competition is unfair. As a user, I don't care that Zig is more "cost/time effective" than Rust or Java, I care about the end result. The end result, right now, is a worse documentation than lots of other languages. That doesn't mean that the people behind Zig are bad or anything like that. That just means that when choosing between X and Zig, Zig's documentation (and lack of such documentation) will influence my choice.
I'm not sure why you would not compare by timeline, as this is probably the best predictor of where a language will be in the next years, which is important data as well.
1) I'm looking for the best, not the best by programmer hours or growth. Zig has accomplished a lot and I'm impressed by the work of people on it. But it's still a strictly worse option than many others. I think it's poised to be great in the future, but right now it's not good enough.
2) Organizations always become less effective as they grow, Zig being small would have an "unfair" advantage since they're still small.
3) Normalizing by programmer hours would lead you to a language that took a few hours to be developped, but is not good at all. This doesn't make sense.
I feel like you're evaluating Zig as if it was a company on which to invest. In that case, Zig is a good pick. It's small but already at a good point, and probably will have a good growth. But if you're looking for a supplier right now, I wouldn't bet on them. It's still small, young, and depend on a few people.
literally only you introduced this 'looking for a supplier right now' concept. I hope you enjoyed knocking down this strawman.
> It's important to consider that for zig at this stage, "not having valid documentation" is expected. At some point not having valid docs will become considered to be unacceptable.
> Seriously, comparing by timeline like that is an unfair comparison.
The article is about failing to learn right now. The first comment was about the absence of documentation right now. You're defending Zig being worse than other languages by saying that Zig is younger, which is a fair point. But you also said that comparing by timelines is unfair, which I think is wrong, which is why we're having this conversation. My point is that comparing by timeline is totally fair, because people are evaluating a language on what it is, and not what it can be.
To speak in more precise terms: the current position of Zig is worse than other languages, especially in terms of documentation. The position adjusted for the timeline of Zig is also worse than many other languages at that point. You're asking us to not compare by position or position adjusted for the timeline, but by position relative to the ressources invested, and/or position adjusted for the timeline relative to the ressources invested. I'm saying that this comparaison is useless, because 1) those relation doesn't scale linearly, and 2) because people are interested in the actual position and position adjusted for the timeline, not the efficiency.
Since you reject comparing the evolution of the documentation of Zig through time ("Seriously, comparing by timeline like that is an unfair comparison.") and right now ("It's important to consider that for zig at this stage, "not having valid documentation" is expected."), your only argument seems to be about the future: "At some point not having valid docs will become considered to be unacceptable.". The vast majority of people, when faced with a language that looks not great at the moment but could have a bright future, will react with "I'll take a look again in a year or two". You argue for the opposite, that people should contribute to Zig instead: "You should probably kick a few dollars into the zig foundation as penance for having made it publically. ;)". That is a weak point, and won't convince the people that have seen languages come and go. Especially these days when new languages have great documentation, error messages, and onboarding in general.
I'm trying EXACTLY to do that, and it's not logically inconsistent. I'm not sure why you are deliberately conflating "using a language" with "putting a (small) amount of money into a project". The exact point is "if you can't afford be present-focused in one way, be future-focused in a different way". I really don't have the energy to parse the rest of what you wrote.
That would work if Zig was the only language in development in that space. It isn't.
You seem to be ignoring that the thread you are replying stems from people challenging this statement:
“Zig is 6 years old. Rust had much better documentation at this stage of development.”
Patience.
I prefer to look up the source code either way, examples and the documentation can only get you so far.
Heh ... Good one.
I can understand why people would recommend reading the source code of projects like Redis or SQLite (and here Zig's std lib) to learn, simply because the code quality is so high.
It's also not like anyone is proposing that there will be no documentation when Zig turns 1.0 (as if high quality documentation and code quality were mutually exclusive!) simply that Zig is not YET at that stage [1].
There's much to love about Zig, and it's always a good time to be reading code, and all the more so while it's growing and shaping because you end up with a deeper knowledge and understanding of what makes a great project great.
[1] Personally, I do find the language documentation to be pretty much excellent and the IRC and Discord helpful—hey, you might even bump into Mitchell Hashimoto also asking questions there. It's the std lib that's a little less, but again, that's in flux and files like `std/mem.zig` are enough of a pleasure to read and learn from. Zig's definitely a swell worth paddling.
But I know others much prefer reading docs over code.
I agree, reading the source is a good way to learn more once a person already know the language.
It can be great to help track down a bug, or learn about an obscure corner case in a function, or see more advanced idioms in use, but it's a pretty terrible and unfriendly way to learn a language in the first place.
Almost every language has a dynamic array of some sort. They all do the same thing. They all have slightly different names. push / push_back / add / append. size / length / len / count. If you don't already know the name of a function then finding it is in a sea of text is very difficult and not fun.
Zig ArrayList: https://github.com/ziglang/zig/blob/master/lib/std/array_lis... C++ Vector: https://en.cppreference.com/w/cpp/container/vector Rust Vec: https://doc.rust-lang.org/stable/std/vec/struct.Vec.html
Sometimes I just want to see a list of functions. I can skim a list of 40 functions in seconds. Scrolling 800 lines of code and trying to pick out every single function takes much longer and is much more error prone.
grep '^ *pub fn' <file>
or equivalent for the language in question.A smarter approach would probably be a code folding plugin for my editor, but I tend to keep my editing environment deliberately stupid so I can spend my cognitive complexity budget elsewhere.
(this is not to downplay your annoyance, just it seems worth mentioning that stupid problems often have stupid workarounds and I have lots of practice at being stupid ;)
I'm not a professional programmer, this is all a hobby for me and i managed to pick and write an entire game with it
maybe you are just bad as a programmer (nothing wrong with that)
you have to learn how to learn, not everybody can do that (and it's fine)
that being said, zig is a language in the making, some areas are rough, but that's to be expected
This was actually my impression after reading the post. A bad programmer making a lot of ill-informed complaints. Zig is an unfinished low-level language. Not suitable for bad programmers.
Also it's a log of personal experience and not titled "an objective criticism of Zig". I've been interested in Zig a little for a while and this was awesome to read. It hasn't made my interest any smaller or bigger, it's simply a description of one beginner's journey, and if the Zig people are smart they will analyze it and then decide which points are valid and which ones can be ignored. Neither ignoring everything nor accepting everything would probably be useful.
> I try to keep my C++ simple and somewhat C-like.
... I knew this would not be very informative.
On one hand, the author's failure to take advantage of the power C++ offers means he will likely also fail to see how to use well what Zig does offer. (One doubts he makes any more effective use of Rust, another powerful language, than of C++.) At the same time, it should make him a better prospect to become a Zig user, not missing powerful features that could make him a more productive programmer.
But he is right that "no hidden control flow" is an anti-feature. What they call "hidden control flow" seems to be what we know of as destructors (or, in Rust, the Drop trait) that is about the only piece of programming automation invented in 40 years.
One might as well dispense with running water and cooked food.
Hmmm, "no hidden control flow" is at least what I want writing distributed systems, storage engines or device drivers.
Calling it an "anti-feature" is dismissing important domains for system languages where control flow in the control plane needs to be explicitly visible, and where the necessity for things like static allocation and NASA's "The Power of 10: Rules for Developing Safety-Critical Code" mean that destructors during the lifetime of a system are an anti-pattern anyway.
Static allocation is a valuable method that is wholly compatible with use of destructors. If you imagine that freeing dynamically-allocated memory is the main use for destructors, you have utterly failed to understand them.
I suppose we differ mostly not in our view of destructors, but in our approach to reducing dimensionality and how best to do that.
But I take issue with this:
// Obviously good and easy to read
return a*(1.0-t) + b*t;
// Obviously bad and hard to read
return add(mul(a, 1.0 - t), mul(b, t));
Sorry, but the first one is not obviously good, it's just what you're used to (the second one is indeed bad).Here's what I would consider obviously good :) for the objective reason that it does not require difficult to track implicit rules about operator precedence:
1, - t, * a, + (t, * b)
I invented this syntax myself :) but it's extremely obvious once you know how it works.It should be clear that this is similar to concatenative languages, where you push values onto a stack then apply operations on the stack. The `,` is used for pushing to the stack, basically. But then, it also allows you to "mix" more standard function notation into it... so when you write `a b` that means calling the `a` function with `b` as an argument (think LISP).
Now putting everything together
1, - t
Should read as `1 minus t` as `- t` is the function `-` being called with one value from the stack, `1`, and one from the "function call" (as if it was LISP `(- 1 t)`).Next:
1, - t, * a
Reads as "1 minus t times a" and has no ambiguity: the result of `1, - t` is the first argument to ``, the second argument is `a`, so you get LISP's `( (- 1 t) a)` but reads much more naturally.Finally, at the end, for readability, we use parens to group the final part of the equation:
1, - t, * a, + (t, * b)
Hopefully it's obvious what happens now?Does anyone like this form of expression or is just me? I tried to mix the best of LISP with the best of FORTH :) and I really like this.
That's also what everybody who as been in primary school is used to. Infix notation have its quirks but at least its ubiquitous.
That would (hopefully ;) have been
(1 - t)a + tbYour notation looks like a very bizare mix of infix arrangement postfix (RPN) behaviour plus added parens for good measure.
RPN has plenty of benefits, once you get used to its initial awkwardness, that I feel your notation fails to capitalise on.
The notation the author complains about is the first version using 'normal' (prefix) functions.
a + b = (+)(a, b) = add(a, b) = a `add` b (Haskell infix) = a .add. b (Fortran infix)
a * b = (*)(a, b) = mul(a, b) = a `mul` b = a .mul. b
so a*(1.0-t) + b*t
is (+)(a*(1.0-t), b*t) = (+)((*)(a, 1.0-t), (*)(b, t)) = add(mul(a, 1.0-t), mul(b, t)) (-> 1
(- t)
(* a)
(+ (* t b)))
Obviously use more or less whitespace to taste.1. It is a linear combination of a and b: there are two symmetric terms, a scaled copy of a and a scaled copy of b
2. The factors are functions of t alone which goes between 0 and 1 as t does
This isn't litigating about brackets, the traditional prefix notation also displays this structure (though less well than the infix notation IMO)
(+ (* a (- 1 t))
(* b t)) (+ (* a (- 1 t)) (* b t))
1 t - a * b t * +The first one is good because its what most people are used too - there's little that can be objectively said about notation once you reach a big enough audience. I could invent a word editor that produced documents from top to bottom rather then left to right and I could invent a million reasons what it's better, but that hardly matters.
Personally I don't think Polish notation is all that revolutionary for humans, it's just easier for computers; however just like anything else the human brain is quite malleable, and if you spend a year doing equations in polish notation you will come to prefer it.
(lerp t a b) ;; possibly, modulo argument order.