Zig, the Small Language
zserge.com
zserge.com
When I was in college we mostly used C++, which exposed (and thereby taught) a lot of core concepts around how computers and low-level languages work. But it was also a bit of a nightmare for... unrelated, obvious reasons.
Rust is great for building software, but as a learning language I think it introduces too many additional concepts, has too many additional constraints, and has too many magical abstractions. Eg, doing manual deallocation helps you later appreciate what the borrow-checker does and why it exists
Zig is supposed to be C++ without all the pain and misery (esp around tooling). That could make it perfect for teaching people about pointers and heap allocations and working with bytes, with minimal extra pain and friction
The stuff about how this actually works can come later, we're not teaching electronics students here (or are we? Zig as first language for electronics students makes some sense)
I agree that (safe) Rust has too much stuff for a first language for Computer Scientists. Lifetimes! Polymorphism! Two entire Macro systems! although it's simpler than C++, surely nobody should teach that as first language.
Zig has a compelling case as a second language if you learned something less practical like ML as first language. Apparently a colleague now teaches Haskell to the equivalent undergraduate classes as a second (or maybe third?) language, and I think Rust would also fit in there for the same purpose but given how practical the current first language is (now Java) he probably doesn't want more practicality.
One thing I'd say for Rust is that we need to start teaching Safety and Correctness early. There's a reason the Medics don't wait until year four to do Ethics and Consent. If we want to stop writing software which is crap we need to start teaching the next generation of CS graduates that's not OK up front, not just in some Friday afternoon class about the Therac-25 which has maybe 40% attendance.
From an educational viewpoint, I don't like that EE and SWE programs have very little overlap. A lot of EE tools could be improved if EE's learned more software (particularly parsing and compilers). A lot of SWE don't get into the dirty details of a CPU and what makes it tick and see it some sort of magical black box, and ergo don't understand digital logic.
Datapoint: There are now thousands of programming languages. There are primarily two digital logic languages: VHDL and Verilog.
If you ever cross over to the EE side, you'll understand how terrible some of the software is.
> I still think the first language when I was an undergraduate (so, last century) was the correct choice: an ML in our case SML/NJ
Again, I don't know. So much of Computer Science programs hinge on getting that first educational experience right. Not so much a problem for MIT, but the program I'm aware of that used Scheme for there CS100 course shrank over time. People just dropped and changed their degrees, when in reality if Scheme/ML were a junior level course, they probably would have more students, higher budget, etc -- and possibly most students would have an easier time after learning a little bit of imperative programming.
Perhaps it's a fair point that CS100 should be a weed-out course. But on the other hand, we don't teach the theory of calculus before we teach students trigonometry.
https://www.registrar.iastate.edu/resources/enrollment-stati...
In around 1995ish ISU moved to Scheme for their first level course, because they wanted to follow the MIT example.
At about 2000/2001 they had a healthy number of students in both the computer-science and pre-computer science programs. The problem you can see in the data is converting the pre-computer science majors to computer science. You can see that most just dropped.
Again you could say that maybe SICP should be a weed-out course, but the reality is, you can be a professional programmer without SICP. Yes SICP might be a nice to have for certain problem domains, but most people aren't functional programmers, and don't need to be.
Today, at ISU functional programming is a 300 level course, and the computer science program has 866 registered students for Fall of 2022.
But we were thrown into C and UNIX when I was at university. I think I’ll have Bus Error and Segmentation Fault core dump messages etched on my gravestone and may be request that the flower vases be labeled gdb (I forget the native debugger that came with the UNIX flavor I was using).
Either way we did data structures and algorithms with Pascal. But I think c should be essential learning from a systems perspective. Maybe I’m too old though.
1) Zig lacks features where they could teach about OOP.
There is no classes or inheritance. To include Zig doesn't lend itself so well to even the more general object-orientated concepts, which are in newer languages like Go or Vlang and others (structs that can be assigned methods and having interfaces). Like it or not (rightly or wrongly), class-based OOP still has many "hypnotized" and conditioned to its necessity.
2) There is the issue of having to address manual memory management.
While that would be the same problem for C, many teachers choose to go the easier route, and select languages like Python or Java. Others like Go, Vlang, D, etc... would make things easier as well.
And Pascal/Object Pascal (Delphi) is still around (a language made for being taught in CS), with a ton of books covering pointers, the heap, bytes, memory management, etc...
3) Zig is still a widely unknown and developing language.
Zig is years away from reaching 1.0. Many teachers are conservative. They often will go with what is already widely known and will likely create the least amount of controversy.
If more teachers can be convinced out of being so conservative, it would probably be because a language is being so strongly pushed by famous companies that are targeting schools or by demand in the job market. Google pushing Go, Oracle pushing Java, Microsoft pushing C#, etc...
First semester was Pascal, followed by C++ on the second one.
This makes a normal workflow with `watchexec zig test` basically impossible, since before I can even run the tests I have to spend time hunting down which variables are used/unused at the moment and (un)commenting them. And it seems like they're planning to double down on this by making you adjust even more trivial things like public/private and var/const before it will even compile your code. https://github.com/ziglang/zig/issues/335 https://github.com/ziglang/zig/issues/224
Maybe I'm a bad developer, but my code is never perfect the first time around. I always spend time experimenting and refining my designs. I want to find a design that works, and only then spend time polishing and prettifying before I git push.
I do understand the reasoning (they don't want people committing poor quality code), but this implementation just seems completely backwards to me. It breaks the natural order. It's like saying we won't let you ctrl+s until your tests pass to make sure you don't commit broken code. Stop nannying me and let me get my work done!
Not that this lint actually achieves that, or even prevents real errors. Go has the same, and it's so simplistic as to only be annoying. For instance not sure whether this fails in Zig but Go will allow this:
v1, err := Foo()
if err != nil {
return nil, err
}
v2, err := Bar(v1)
return v2, nil
Error of second call is never checked, but go has no issue with that, because it only tracks definitions per use. v1, err := Foo()
v2, err := Bar(v1)
if err != nil {
return nil, err
}
return v2, nil
also works fine, despite probably sending complete nonsense to Bar, for the same reason.But then it's an absolute pain in the ass every time you're fucking around and stop using a debug import or whatever.
export fn x() i32 {
var a: i32 = 1; // okay
a = 2;
return a;
}
This does not: export fn x() i32 {
var a: i32 = 1; // unused variable
var a2: i32 = 2;
return a2;
}
> Not that this lint actually achieves that, or even prevents real errors.I have to agree. If there's one thing I've learned over the years, it's that you can't prevent bad code by enforcing lint errors.
I completely disagree, I regularly catch mistakes/bugs in my own code due to enforced linting errors (both locally and in CI), and just in the past few months I can recall numerous instances where CI linting caught bugs in teammates' code too.
There are both false positives and false negatives with linting, but that doesn't make it useless at all.
And that's why this issue is kind of hard. It's why I end up block scoping a lot of error code in Go, which is ugly but generally safer.
It is.
> That `a` is reassigned doesn't make it unused. It is, after all, a variable.
But that's exactly why it's bad, "unused variable" is not a super useful signal in the first place, and if you assert that everything must be used then you need to protect against "unused stores" instead.
SSA tells us there is essentially no difference between assigning and reassigning.
A lot of simple cases can be handled with linting rules, BUT that, IMO requires a few additional things. Namely, being expression based, rather than statement based, and enforcing something like a hindley milner type system.
The missing error checks are annoying, but if you have appropriate editor config it is hard to miss them: https://cdn.billmill.org/static/newsyctmp/warning.png
Basically writing go without `staticcheck`[1] is not recommended. If you do have it set up, it's pretty easy to avoid simple errors like that. I do wish the compiler checked it for you.
Don’t get me wrong I like a good unused code warning.
What frustrates me is that Go’s is dumb / unreliable, and it will stop you from working entirely until you’ve complied with this whim, which has a fraction of a percent chance of identifying a real bug.
> Basically writing go without `staticcheck`[1] is not recommended.
So why have these things as mandatory compiler errors?
I guess they could have levels and allow them in debug mode or with special flag or something?
But most everyone else likes this, so it stays in.
There are no real right answers here. Adding a switch for it just brings its own annoyance (every compiler switch is a bug). You just wind up settling for a local optimum.
I totally get where this is coming from, but on the other hand it seems like Rust and Zig both get a lot of value from having the compiler understand the difference between debug and release modes, and it seems like modern C++ suffers somewhat from not having any built-in way to do something similar. The optimal number of compiler switches might not be zero.
It's not useless to have them but they can be utterly infuriating for exactly the reason you mention since most projects have warnings-are-errors turned on.
The way I shut it up in a class is usually by branching on the this pointer. Sneaky.
We simply have to have a very high bar for new features.
> hoping that you forget to remove the assert
I routinely add a lot of instrumentation when developing code. I use `git diff` to remove it before committing.
The compiler builds itself fast enough that this is not a problem. I have one window on the source, the other on the build.
What's damning is how much Stockholm Syndrome there is around this feature, with people saying it's no big deal and it helps catch bugs. It's more annoying than helpful, and it catches a very small amount of corner cases, while completely killing productivity.
And you know what's the reasoning behind this? "Zig doesn't have warnings." As if it's a massive undertaking to add warnings to a compiler. What a sorry excuse.
I have ranted about this before, so I already feel my blood pressure going up.
Setting aside the unused variables issue for a moment, you might want to take a moment to ponder the fact that not having warning messages is an explicit design choice, not a missing feature.
I see this comment a lot in response to feedback: "Don't like it? Don't use it!" Every single time, it reads to me as "stop complaining!" For as many times as I've seen the comment, I haven't been able to read it any other way.
Am I being uncharitable or is it really just a clear-cut attempt to silence debate?
On the other hand in case cloud services like those provided Google/FB/Twitter/Cloudflare etc it is much more difficult to just go with "Don't use if you don't like" because not only there are not much alternative but also because sometimes Cloud service providers act in unison. So if one is pushed out from a service they are likely to be pushed from all other services too.
Another situation where I think the compiler is too eager with errors: unreachable code. A bare `@panic("")` won't compile, but you can "turn off" the error by doing `if (true) @panic("")`.
Could it be that this problem is due to Zig's stance on "no warnings"? It's admirable to try and partition things into strict "correct" and "incorrect" categories. But with only static information it seems like there's always things that seem a little grey.
I love Zig. It's built on such a better foundation than any of its competitors, old or new. I've never been happier or more productive doing systems programming.
But the "no warnings" philosophy alienates developers and seriously hurts adoption. It's almost tragic that a language so impressive is going to die on this hill.
I originally looked forward to splitting my projects between Zig and Rust depending on the project parameters, but because of this one seemingly small issue I will never use Zig.
That's ok, I respect them having the right to do what they want with their language but it is a little sad because I liked pretty much everything else about Zig.
I had to work in python for a long time where code tends to be incredibly relaxed. When I switched to strict mypy typechecking at every build I became much much more productive and found and fixed dozens of bugs way earlier in development. I now am worried about all warnings and always strive to fix them.
I am also making an analogy to "not being allowed to smoke indoors" when I see these complaints. People hated those laws when they were first proposed but everyone's cool with having to do a tiny bit work for health and safety.
We'll see how communities play out, but I'm all for being rigorous.
I'd prefer `eslint-expect-error-next-line` instead to warm me when the comment is actually useless but it's better than what Zig seem to do.
The strict compiler check would have had exactly the opposite effect from the intended one.
Kelly is designing Zig to do nothing surprising or change things underneath you. Even C does things that are surprising.
If you tell the compiler "Hey, I know this is unused but I am going to write it anyway" then the compiler _can_ make decisions such as eliding the variable.
I think we do all agree that it is sensible that a variable in code that is not used should be a compiler warning (if it is an error or not is clearly debatable).
It really irks me when there isn't a way to have granular control over silencing warnings.
you get used to it after some time - much better than for example rust, where you have to figure out all the lifetimes, muts and generic traits before.
They aren't comparable. Everything you mentioned in Rust is necessary for typechecking. Unused variable lints aren't necessary for anything.
I am saying that prototyping (for me) is much easier in Zig. Some people like to define all the types (and traits) first, if it works for them then it's ok. I like to get feedback on something working as soon as possible because I'm likely a bit wrong and I will need to refine or rethink the whole thing again, rust is putting obstacles in my way, Zig is not.
[0]: https://www.scattered-thoughts.net/writing/how-safe-is-zig/
Put another way, the answer to your question is because even if program correctness were the only value that matters to anyone who does low level programming (and it isn't), there's no agreement sound memory safety in the language -- at least the way Rust does it -- is the best way to achieve correctness. Put yet another way, the answer is: for the same reason people would want to use a language with memory safety -- because they think it helps them achieve their goals better.
Zig does lack a general solution for temporal memory safety. That is a downside. But it can still have that safety in ReleaseSafe mode, at least - for example, by not reusing memory addresses. That adds overhead, which might rule Zig out in certain areas, but not others.
Also, while in many domains safety should be the #1 factor when when choosing a language (like building a web browser), that's not universally true. Other factors exist.
I don't believe this is true. Zig has pointers to unknown numbers of items, which don't seem bounds checked: https://ziglang.org/documentation/master/#Pointers
They prefer slices idiomatically, but that's not "full spatial memory safety". C also prefers you to pass the length of every array whenever you pass a pointer to it, in that correct code must do this. But the entire point of language memory safety is that we don't trust programmers to consistently do the right thing.
You might be able to say something like "Zig minus features X, Y, and Z has full spatial memory safety". I'd be interested to see what features those are: it looks like at a minimum you would have to get rid of multi-element pointers and extern unions.
> But it can still have that safety in ReleaseSafe mode, at least - for example, by not reusing memory addresses.
The overhead is extremely high, because it leaks an entire 4kB page if a single allocation from that page is still alive. In the worst case, it's equivalent to rounding every allocation up to 4kB. If you're OK with that overhead, you could just link in the Boehm GC and get the same safety with less memory usage, and as an added benefit you wouldn't have to call free anymore.
> Also, while in many domains safety should be the #1 factor when when choosing a language (like building a web browser), that's not universally true. Other factors exist.
In the vast majority of those domains, you could just use a GC'd language. I don't see much room for a new language that isn't memory-safe in 2022.
[*] is a syntactically delineated unsafe feature. Rust has the same thing. It's like saying Rust prefers safety, but doesn't enforce it because it has unsafe features; same goes for Haskell.
> I don't see much room for a new language that isn't memory-safe in 2022.
That statement [1] is about about as silly as "I don't see much room for a new language in 2022 that is obviously too complex to see wide adoption." If a language becomes popular then obviously there's room for it, and if it doesn't, then the question is moot. Maybe what you mean is that you don't think Zig would ever become popular -- and you may well be right -- but you work on a language that is subject to similar scepticism and doubt.
Both Rust and Zig have some huge challenges to overcome if either one of them is to have "room", and I wouldn't bet on their chances, but it's good to have some widely different approaches and see which, if any of them, turns out to have "much room".
[1]: Even disregarding the hard question over which of Zig or Rust make it easier to write correct programs, which could go either way; that sound memory safety is a better path to correctness than a balance of less soundness combined with simplicity is your opinion, and few would claim that Rust's flavour of safety doesn't come at a cost.
Well, `[*]` pointers and extern unions exist for C interoperability. I'm sure that Rust has to do something comparable when interfacing with C.
These pointers actually do help making C interop more safe. In C there is no distinction at the type level between a pointer to a single element and a pointer to the first element of an array, so if you at some point get it wrong, the compiler can't help you.
In Zig, if you annotate `extern` function signatures correctly, the compiler will be able to tell you when you're wrongly trying to iterate over a single-element pointer (or viceversa).
That's a good feature to have, and IMO probably the worst supporting argument for your usual complaint about Zig's existence.
I don't think people say a language is unsafe if it has opt-in unsafety in controlled, non-idiomatic ways. Java, C#, Rust, etc. all have unsafe things you can do, but they are still memory-safe languages.
Specifically for Zig, I don't see a reason the compiler couldn't have an optional warning on using an unsafe pointer, so codebases can be audited for this risk, making this a controllable form of unsafety (unlike most of the unsafety in C!).
> The overhead is extremely high, because it leaks an entire 4kB page if a single allocation from that page is still alive.
In general you are right, and in a long-running process that would be the case. But consider a short-running wasm event handler: such overhead is generally not going to be significant there. So this rules out some uses cases, but not all.
> I don't see much room for a new language that isn't memory-safe in 2022.
Yeah, I agree the space for a language like Zig is limited: GC languages are the right answer for most things anyhow, as you said, and when they are not, often memory safety is crucial (like in a web browser) and I'd strongly prefer Rust over Zig there.
Still, there are use cases where Zig seems nice, like wasm event handlers that I mentioned: they're sandboxed anyhow, binaries are small, and it's nice to not have GC overhead. Zig's simplicity and fast compile times are a bonus.
Another use case are low-level things that you'd need lots of unsafe in Rust anyhow. I'm not sure if I'd prefer Rust or Zig in such a case myself. I'd prefer either over C, though.
IMHO these are the only valid reasons for unsafe code in a systems language:
1. interfacing with unsafe external code (e.g. call the kernel directly)
2. implement the language runtime/GC itself
3. access hardware capabilities not exposed in the language (e.g. SIMD, GPU, memory-mapped IO, memory barriers)
These are semi-valid to bad reasons for unsafe code:
4. generate new machine code to, e.g. implement a guest language
5. explore the performance ceiling for a very specific code pattern that would otherwise be inline asm
For Virgil, I am at reasons 1 and 2. For Wizard, I am at 4 and 5.
I'm under the impression that crashes are considered perfectly safe because they don't lead to memory unsafety. They're inconvenient, but eliminating them is more or less impossible: if you don't want stack overflows, you have to ban recursion; if you don't want integer overflow, you have to ban addition or insert checks for overflow; etc.
What areas specifically?
It seemed back when Rust was an upstart that everyone would ride its behind about how this and that feature could add runtime overhead compared to C or C++. Especially compared to C. But now people are comparing Zig and C and are saying (indirectly) that they are fine with runtime overhead in exchange for more safety? Has a shift happened?
(it's of course not a "general solution", but I wonder if something like 'tagged pointers' could be that, even without hardware support, and at some runtime cost)
Also, safety is not the only consideration when writing software right? Otherwise, all software would have to be written using formal methods [2] ... clearly overkill for many (most?) applications.
Maybe it is a bit obvious, but I think someone would pick Zig because they like the set of features that the language provides, allowing this person to write reasonably bug free code in (hopefully) shorter time than using a different language or methodology.
I think most ppl choose the language because they like the tools/ecosystem/libs available, because it aligns with previous experience/s, because of how popular it is for the kind of app they are writing, and because they agree with the authors philosophy (as opposed to choosing it based on any particular language features).
--
When is Zig the best choice?
"Freestanding" targets, perhaps?
I often write software in C to run directly on microcontrollers; no OS, no runtime.
I can see the advantage freestanding Rust would provide for e.g. 802.11 firmware. Broadcom repeatedly suffers from their choice to write the firmware for their WiFi chips in C.
… but would Rust provide any clear advantage to the programmers of a Furby or a keyboard controller?
Add to that, Rust and Go have strong corporate backing, continually get lots of shine in the media, and often are given positive spin or ratings whenever compared to other languages.
As you mentioned Nim, look how far uphill it still has to climb, after the many years (around 14 years) it has been out. TIOBE Index doesn't even have Nim in their top 50, and they aren't even on IEEE's top programming languages list (of around 60 languages). This is just as likely the fate of Zig as well. Not saying some people won't find Zig useful for particular niches, but rather it will likely never be on the level of Rust, Go, D, Object Pascal, etc... or possibly even Nim.
- easy interfacing to C libs (glfw)
- very easy and readable metaprogramming (compile time inline for over all enum variants, duck-typed and strong-typed anytype at the same time, @compileError for advanced compile checks so you can teach compiler new rules about your code instead of trying to fit your code in the extisting compiler rules)
I don't know enough about Nim to pass judgment.
Two reasons I decided to give Zig a try: The official chat channel is on IRC, instead of Discord or Slack (so the people involved care about efficiency, open standards, and avoiding trends/bandwagoning), and it has an early but promising-looking Swift UI-like cross-platform UI framework in development: https://github.com/capy-ui/capy
Or it's enthusiasm for something that's genuinely worth being excited about. Such things do come along from time to time. It's probably not a good idea to be permanently jaded, to assume that everything is zero-sum, that we can never again make progress by bending tradeoffs.
1. Rust uses Zulip (ie, not Discord or Slack),
2. Rust doesn't lack for promising-looking, declarative/reactive UI frameworks: https://iced.rs/, https://crates.io/crates/egui, for example.
Computer security is a real problem that must be fully addressed at every step of the design of a system, language and tooling can help, a bit, not much.
Rust is not a "safe" language. It might be -safer- than C/C++, but only time will tell if the difference is significant.
In any serious system written in Rust, significant portions of it will have to be written in "Unsafe Rust".
You’d also have to define “serious system” and “significant portion” and if you really do mean “any,” that is, is a single counter example enough, and if not, what percentage of them would need to be so for this to be true. You’d also have to say what’s in bound or not, is it just the code you wrote for the system, is it any dependencies, etc etc etc. Does “significant” mean “large amount” or does it mean “most important”?
GhostCell is an interesting venture away from that idea. For the sorts of loopy data structures that upset the compiler, GhostCell provides a small, well-vetted, probably correct unsafe core and a safe API for accessing it that greatly reduces the amount of unsafe one should need. The core mechanism is just a higher order phantom (ghost) type parameter that gives the compiler enough information to vet that the thing you're handling was acquired properly.
The community has generally been pretty gung-ho about getting the most safe bang for your unsafe buck, and the compiler has enough bells and whistles to make many such ideas pretty seamless.
There is Go, Nim, and Rust. All pretty modern system programming languages and some of them even close to C performance.
Does Rust offer many advantages over Zig in this space?
I mostly program in C, but I've been eyeballing both for some time now.
If you are writing web apps then Rust is probably "safer". If you are writing low level code that requires obtuse memory management then I don't think you can make a blanket claim that Rust is just "safer". A lot of this nuance is not often discussed.
It sounds like you're saying rust because I don't know what other languages claim both.
Rust is an awful language. That's why. If you're using a systems programming language then chances are unsafeness isn't the problem you're most concerned with
Does that mean you've rethought that position?
Although the "Python Paradox" does not apply to Python anymore (it has become even more mainstream than Java), it does highlight an important point:
If you choose your language based on "I want the biggest pool of developers" you are also saying "I want the most average and mediocre developers that I can pay for cheap".
- my math type with + - * / overloads
- simpler way to fill an array, i can never remember the syntax, it doesn't feel natural `[_]u8{0} * 10;`
- smarter type system, i am tired of casting everything twice
A good language is not a language set in stone, a good language is a language that doesn't make me feel like i have to suffer because they made a stupid decision years ago and they refuse to make it better
But in all seriousness that's pretty much antithetical to zig's goals regarding explicitness.
But I wound up being convinced that on balance it was a good thing. But I was able to inculcate a culture that operator overloading should be restricted to the creation of user arithmetic types.
Not allowing the overloading of unary *, &, and dot also help discourage non-arithmetic overloading.
Because operators "should not be function calls"?
That would make zig unusable on soft float/soft div architectures, or would have to get rid of / for division and operators for floats.
But I would also assume add() is inlined and not be a function call in a sane language/compiler, so the explicitness even falls apart from the start.
Odin proved it that it can be made efficiently while keeping sanity
And vec2 + vec2 can probably be unambiguously translated to assembly code while mat4 * mat4 is an other story. In most cases it should be a function call, and not a trivial one, with SIMD it can be relatively fast but still several orders of magnitude slower than 1 + 1. (we're talking about more than 500 scalar-equivalent operations)
And unfortunately, from my experience, if you let people (especially new coders) use this kind of powerful syntactic sugar they tend to ignore the performance characteristics because it look so simple and basic, like if the CPU had a special instruction to multiply two 4x4 matrices.
I prefer when the function call is explicit, it is a bit more cumbersome to write, but there is less hidden complexity.
you hide and obfuscate basic operations with functions, that's worse, specially when you have to chain arithmetic operations, with functions it becomes ugly and unreadable
why? operator overloading doesn't help you solve any problems. you can have the readability with methods which are named appropriately.
operator overloading seems so powerful and useful until you realize one day that it only changes the appearance of things, and makes no difference whatsoever to anything you are actually doing.
https://dlang.org/phobos/std_checkedint.html
where operator overloading is used to create variations on integer types, like specifying the behavior when overflow happens.
Besides, `a + b / (c * d)` is far more readable than `add(a, div(b, mul(c, d)));
Just because people misuse operator overloading doesn't mean it should be removed from the language.
Also function overloads should never be allowed, either. After all, that's not explicit and the only thing that ever matters is being obsessively explicit for the sake of being explicit. How can I possibly tell what `min(a, b)` is going to do if function overloads or templates or macros exist?!? UNREADABLE!
I think at this point it's pretty well established that operator overloading is a net-good, and languages that don't have it end up with far more bugs than languages that do. Or they come up with arbitrary nonsense rules on why some types are allowed to have it but not your types you dirty filthy casual. Looking at you, Java, where the Integer class is allowed to overload operators but BigInteger isn't. Which, then again, is something Scala and Kotlin immediately completely reversed.
It depends. Low level coding the like of which zig is tailored to tends to involve a lot "write and forget about it" infrastructure stuff. At least that's my use case.
It sucks to painfully go over that cryptographic hash, that codec or that compressing algorithm you wrote 15 years ago and still works flawlessly because it wouldn't compile with the modern version of the language.
I think zig is very nice but I wouldn't use it even for my personal projects because it's still unstable (and from my POV it will remain that way for at least 5-10 years).
C is terrible but it's not sufficiently terrible that I would use a language that is orders of magnitudes less popular and still in its infancy. Too risky.
You have to specify the names more times but you can also do any syntax you want.
Current $work language is Go, which is so painful to use. I often look wistfully at all of Zig's features that improve on what Go does (particularly with regard to error handling). I hope in the future I can use Zig as a Go replacement and not just a C/C++ replacement.
I've only scratched the surface of Zig myself, but my impression is that replacing Go with Zig will probably be painful in most cases. I think of the stereotypical Go project as a backend API service, where memory is relatively plentiful, and "make a copy of this string" is something you do all the time without thinking twice about it. It seems like Zig wants you to be more thoughtful whenever you're allocating memory, which makes a ton of sense for low level libraries or kernel code, but which sounds painful for typical large applications.
Zig's scope is just to be a better C that's free to add modern features like optional types, compile time expressions instead of string-macros, source level modules, packages, a more expressive syntax for writing bit-packed structures, a standard testing framework, deferred function calls, and so on. You can also directly include c headers directly in zig code, so this gives you a pathway to modernizing C Code incrementally.
This isn't going to be appealing to Rust current users though, because “fighting the borrow checker” is a learning curve issue, you don't fight the borrow checker anymore once you've internalized its rules.
> or making gratuitous copies of data to satisfy the borrow checker.
You get it backward. In Rust there's less gratuitous copies, not more, because the ownership rules and the borrow checker gives you compile-time guarantees, while “defensive copies” are common in C++ for instance, “to be safe”.
I still fought the borrow checker after a year of using Rust.
And I found many situations while using Rust where I either needed to clone, or use unsafe where it's easier to make a mistake than in other languages because the syntax is extremely unergonomic and the memory semantics are much less clear.
(disclaimer: Zig has or had some pretty big holes around escape-checking function return values - that's something that even C compiler are warning about these days, at least for simple cases - but apparently improving this aspect is on the todo list).
The borrow checker doesn't stop most memory errors, just the easy ones at the cost of making it painful to write basic data structures. There are many memory related CVE on common rust libraries.
The thing about that though, is there are a number of better C languages (both old and new) that have various levels of C interop. Languages like: D, Odin, Nim, Vlang, etc...
First: there's only a certain amount of effort that one is willing to put into accomplishing one's goals, effort that will be spent on learning tools and building a project. It's not unlimited. If I just want to make a game where a character walks around the world, and I only have 20 or so hours of time before the money / motivation / market opportunity runs out, and it takes 300 hours to master Rust, that's a non-starter. If someone has 2000 hours of time available, then it'll be worth it to learn Rust before starting. If someone has even more than that, then it might be even better to use more advanced proof tools to be even safer than Rust, if safety is an absolute priority.
Second: Zig offers a lot more freedom than Rust. The borrow checker historically has difficulties with observers, backreferences, dependency references, several forms of RAII, and graphs. If one just wants a struct to contain a pointer, sometimes it's just easier to use Zig than to change one's entire architecture, which takes a lot of time and effort that one might not want to spend. Of course, this isn't a problem if one uses Rc and RefCell more, which brings its own tradeoffs.
Third: There are a lot of cases where the cost of memory unsafety just isn't that high, and Zig's mitigations are more than sufficient. For example, if I'm making a non-safety-critical app or webassembly program that's sandboxed on the client's machine and only talks to its own trusted server, that's more than enough security for a lot of use cases. In these cases, a couple memory safety bugs are just like any other bug, and it's not worth it to spend the time reducing this release's bugs from 23 to 21. This is how it was on Google Earth; the vast majority of our bugs were logic bugs, not memory unsafety (and this was C++).
Coming from Rust, the first point likely doesn't apply to you as much. The second and third points might apply, depending on the situation.
I think Rust's tradeoffs are stellar for cases like medical devices or HFT where the cost of memory safety and higher latency is much higher than your average webserver (where Go might be a better fit) or your average single player game (where memory unsafety bugs are just inconvenient, not critical).
If I was to sum it up, I'd say: Rust is pretty close to perfect on paper, when you don't consider the other human factors, such as limited time, training, or that sometimes it isn't worth it to prove memory safety. When you consider these factors, other languages can make more sense.
But then why not just use a GC'd language? It isn't 1995 anymore; most apps are written in GC'd languages. C++ makes sense when you already have a lot of C++ code, but if you're talking about new code, I don't see a lot of room for a language without memory safety.
I mostly agree with you, and I'm enthusiastically using Rust. But I'd love to be able to develop new software, with modern developer conveniences, that could run on a 1995 computer. Assuming that implies a 32-bit OS, that seems plausible with Rust, albeit not with std (but alloc should be fine). Maybe I'll play with that for a side project sometime.
Isn't your statement at odds with assessment from many big tech companies that memory safety is the prevailing reason for CVEs? I'm sure you're not suggesting that the folks who make these assessments are not experienced C/C++ programmers, so I'm curious about your thinking on this.
Think of what zig has to offer instead and try it out to see how it works.
If it's simplicity, it seems like complex programs would be supported by complex libraries instead of complex language features, leading to the same level of complexity but with less consistency.
"Smallness" seems to be a sought after feature but I'm not sure why.
Here's a simple but perhaps disappointing theory. Painting with a broad brush for a moment, there are primarily two sorts of thinkers: memorizers and logicians. The former have astounding recall, and consequently have no qualms about learning loads of vocabulary. They, for some reason I will never understand, actually delight in learning a foreign language. When they end up in the sciences, it is more often than not some form of biology where knowing all the assorted physiological systems and every cataloged failure mode and the symptoms thereof is of great benefit. The latter can't do any of that, but if you pause for half a second in the middle of a thought they will start trying to complete your sentence by suggesting words. When they find themselves in a foreign country, they figure out how to use the washing machine and vending machines and generally fend for themselves without any help despite not being able to read any of the labels. They might never learn the language, but they quickly figure out which characters mean "road", "park", "river" etc based on nothing more than the subway station names. When they go into the sciences, they tend to end up in math or physics. Newton only makes you memorize 3 things, and from there literally everything else follows if you think hard enough.
So, for the logician type thinker, small languages are just deeply pleasing for the same reason Newtons laws and Chess are deeply pleasing. Those are the antithesis of learning a foreign language. Small coding languages let you know everything while learning as little as possible. The memorizer types must see the appeal of small language the same way I see the appeal of memorizing all the characters in the Chinese writing system. I see no appeal in learning all the characters in written Chinese. Why would anyone ever consider all those characters a sought after feature?
Then the retort might be “but they wittingly or unwittingly pursue the goal of complexity for its own sake!”, in which case that would just prove that you have stacked the deck in favor of the “logician” (who makes simple languages for simplicity’s sake—simple).
Another problem is that some language designers are logicians by trade already.
> So, for the logician type thinker, small languages are just deeply pleasing for the same reason Newtons laws and Chess are deeply pleasing.
Chess has Opening Theory. I bet the Half-Delayed Swashbuckler Sicilian just comes naturally to the oh-so exalted logician.
(Then the retort might be, oh I’m just talking about the game in the abstract, not the metagame. Well it is very important how a game is actually played and how a language is actually put to use.)
A large language might be more consistent and well thought out than a smaller one. That, again, is a separate detail.
And we're talking about the main selling point of the language. Not just 'oh and its pretty small too.' Of all the things to focus on, it would not be my priority for sure.
Actually, there's a logic (a stroke order logic and also some common building blocks) to how to build the characters in Chinese (and other similar languages). So I would imagine the "logicians" might be OK with it :)
Very well put
We have to deal with complex libraries no matter of how big or small the language is. A small language that you can keep in your head makes dealing with them easier.
I would also argue that smaller languages lead to more consistency not less simply because there are fewer ways to do things. Look at early Python vs modern Python for example. In very early Python (pre-2000 or so) there was actually "one and preferably only one obvious way to do it". Modern Python with all the features it has grown over the decades is the exact opposite of that. C++ is another obvious example.
That said, I'm used to Java or C# as far as language size goes and I don't feel suffocated by it. I'm not sure if its a natural inclination or a learned skill.
The main problem is that every c++ programmer has their favorite little bit of complexity in the language, and that favorite bit is different for everyone. So you end up with a codebase sprinkled with the full spectrum of language complexity if you're working with a large team, meaning you need to understand all of it if you want to understand the codebase generally.
At least with complex libraries, the complexity can be safely hidden behind the API. Not so with complex languages, at least in my experience.
For example, if we have a bunch of language features, and we decide to add a new one, e.g. exceptions, we now have to consider:
- How exceptions work
- How exceptions interact with threads; how exceptions interact with dynamic binding; how exceptions interact with dynamic binding in threads; how exceptions interact with mutable variables; how exceptions interact with mutable variables in threads; how exceptions interact with dynamically-bound mutable variables in threads; and so on.
Also, all of the language's tooling needs updating: compilers, linters, documentation generators, static analysers, formatters, syntax highlighters, auto-completers, etc.
- Is small/simple/principled (i.e. understandable)
- Enables many desirable programming styles or capabilities
- Complements other features: ideally working even better in combination; or at least, working in an obvious way
Algebraic effects tick all these boxes IMHO. In particular, there are many languages which feature try/catch AND for/yield AND async/await AND call/return; whereas having algebraic effects subsumes all that, turning it into library code.
Note that we can get pretty much the same thing with, say, delimited continuations; so they're also cool. However, having delimited continuations AND algebraic effects would seem redundant/bloated.
In my experience, languages with complex features still have complex libraries. Library complexity is a constant. If you simplify the language, at least you free up a little cognitive load there.
In addition, simplifying the language often implies (or is even equivalent to) minimizing footguns.
Now you can go see all the arguments about RISC vs CISC, which is the same argument. Should I have a small instruction set that bloats the final program's instruction count? Or a complex set with slower instructions?
Note that small doesn't imply lacking features. For example, something like a std::transform is often "smaller" than iterating through an array with a for-loop, but it's also clearer and less error prone. It's a result of the language itself providing more powerful data structures, and interfaces that make manipulation of those structures easier, which is what I'd argue "small" is actually shorthand for. I don't want to be using a language that only has if-statements and goto, which would be "small" in one sense of the word; I want to use a language that has the minimum necessary number of keywords, decorators, or meta-languages, and instead provides clear, ergonomic ways to express my intent to the machine.
-- Warford J.S. (2002) The BlackBox FrameworkFor me, it isn't just the 'smallness' of the language that matters, it is the understandability of code bases written in the language that matters.
Having fewer language constructs (especially ones that offer overlapping alternatives) means that programs that solve the same problem will tend to look the same. This makes reading a new code base easier, which makes on-boarding engineers faster. This is the main reason I like smaller languages.
Languages like C++ (that seem to adopt every language feature that can be implemented) enable developers to solve similar problems in vastly different ways. In small programs that isn't an issue, but in large, old code bases it can become a very costly mistake. It can be hard to understand a large code base as it is, but when 100 different developers have touched the system and each has a pet style, you are in for it.
Of course, "no code at all" isn't useful, so there's a medium to be found.
Libraries _can_ hide some of this, but then that might also be yet another function or operator that does the same thing. Internalizing when to use one method vs. another adds to the total cognitive load of using the language, and it takes away from the mental capacity available for the problem you are actually trying to solve.
Sometimes these differences are necessary and important, other times they are less so.
complex features (in a language that does the simple things correctly) are always easily composed from the simple features available.
the antithesis of small languages is C++, which is probably the most popular language on the planet in which 0% of its users know 100% of.
with a small language, i can’t use language features which you do not understand when it comes to read my code or take over maintenance, or even understand it.
The problem with C++ is not the size of the language, it is how inconsistent it is.
I can do puzzles in any of them, which many will fail to guess, specially if runtimes and standard libraries are part of the question pool.
Zig is still pretty feature rich (for instance, full generics support) but remains small and simple
For example, one might be making an embedded program where all memory is pre-allocated up-front and there's no heap usage. There's not as many opportunities to mess that up than your average program. This is where Zig would shine compared to e.g. Rust whose borrow checker would be helping with less, yet is imposing its usual complexity burden on the programmer.
Libraries then framework are the other layers, built on top of the language.
A language designer is unlikely to understand all the needs of the language users, trying to solve complex problems at the language level can seem neat and powerful for the exact use case it was intended, and also unnecessary/cumbersome/too general or not general enough for many more use cases.
_ = my_unsed_var;
grep '$ *_ ?=' *Even C has accreted a lot of features.
I think Python was small at some point.
Pascal was small at some point, but by Delphi had become quite big.
Go has just added generics.
I think looking at a "smallness" for a relatively young language is likely to be misleading.
Check the short list of changes to the language specification in the Go 1.18 release notes: https://go.dev/doc/go1.18#generics
Go is still a small language, even with generics.
fn count_nonzero(a: []const i32) i32 {
var count: i32 = 0;
for (items) |value| { // "for" works only on arrays and slices, use >"while" for generic loops.
if (value == 0) {
continue;
}
count += 1; // there is no increment operator, but there are shortcuts for +=, \*=, >>= etc.
}
}When the return value is unspecified, Zig will default to an implicit return of the last-calculated rvalue. The same behavior is also found in Ruby, Lisp, and some other languages as well.
For example, according to xbps on Void Linux, GCC is a 172MB install. Zig is only 166MB. A Clang or Golang install is twice the size of Zig. A Rust install is three times the size of Zig. And so on.
As a software engineering lead I am becoming more replaceable every year as others learn to maintain various components of the project. Recently, I am proud to have dropped to having fewer than 50% of total commits in the Git repository for the compiler.
I also find it incredibly frustrating when an application I rarely use fails to launch because it wants to force me to download, process and install massive updates for features that I don't even want.
As an engineer, I care about binary size because gigantic binaries tend to require far more time to compile, move around and work with. Gigantic binaries tend to be a reliable signal of slow, toxic, internal processes that chew up positive energy and produce cynicism and frustration in it's place.
It's not about the disk space per se, it's about treating waste as if it's a virtue.
I'm more likely to believe Zig's "small binaries" are more from lack of optimizations than some obsessive focus on L1 icache density. Which, given it's not a 1.0 language, isn't something that can be held against Zig. But it'd hardly be a strength, either.
By the way, I wouldn't be surprised if Rust 2.0 will feature a typechecker that guarantees that everything fits in the L1 cache.
It is a simple proxy/heuristic for the level of bloat / crust / overengineering of the runtime and to the quality of the compiler.
If a language / framework can't do an extremely simple task in an almost optimal fashion, then it is unlikely that it will improve with more difficult tasks.
It is like looking at the CPU/GPU load of an app in idle mode.
In embedded development you can forget the overhead due to binary formats such as ELF (while the toolchain might output these as an intermediate, what is flashed is just the relevant sections), but if you are doing things like loop unrolling all over the place, you're going to get less functionality for static program space.
This is why ARM and PowerPC have "embedded" variants of their ISAs. What this means in practice is a compressed form of their instruction representation. Why? If you have compressed instructions, you can fit more of them in program flash. So both Thumb and VLE are variable length encodings. Thumb, for example, can use 16 bits for many instructions, i.e. two bytes, whereas ARM by default would simply take four bytes for all instructions. (For the avoidance of confusion, they still "address" 32-bits of memory and are thus still 32-bit microcontrollers). PowerPC's VLE is similar.
The driver here is cost. You could put in more program memory (the address space has plenty of room) but that costs more money, and when you don't need it, why do it?
Also IMHO, statically linked programs simply shouldn't contain any code that will never be called, just out of principle.
I use golang with all the size optimization steps that I can use... it makes a difference to make the binary smaller... it doesn't have to be minimal, but 4megs is better than 10megs.
A lot.
- How do they communicate? Maybe Json if it’s just data. For more control you need FFI
- And FFI is such a hassle that some languages seem to focus mostly on getting a nice C FFI
- ... That lingua franca that proves that you can have all the simple languages that you want as long as its name is one letter and starts with “c”
- All kinds of minute differences that are real tradeoffs when it comes to each individual language become major pains when switching between them
- Sometimes you seem to have to rule out languages just because they don’t have good libraries for X. Or can’t do concurrent tasks. How are multiple small languages supposed to flourish when 20+ year old languages can’t get good library coverage or get rid of their global interpreter lock?
I would like small languages if they all played nicely together. But they almost never do.
I'm sure there are plenty of other languages such as Nim etc that never get similar exposure.
// Arrays may contain a sentinel value at the end, here array.len == 4 and array[4] == 0.
const array = [_:0]u8 {1, 2, 3, 4};
Small typo. Len should be 5, I believe.If you're coming from Python and Go, then producing the fastest possible program that uses the least amount of memory probably isn't your top objective. Zig is fast, but you're going to have to think about memory management and other things you take for granted in garbage collected language. Zig also has nowhere near the number of libraries that exist in Python or Go, so you'll be building a lot more stuff yourself.
Also, the simplicity of a language like Zig may be a hindrance if you're working with really big code bases. If you need something that will scale to really large projects, like say a web browser, you're probably better off in Rust.
If performance is not an issue and you want something that will let you focus on the program's logic rather than the implementation, Python is probably what you want.
If you're writing a microservice to run in a datacenter, you'll be up and running in Go much more quickly and the performance is probably good enough.
But if you're writing a program for an embedded or limited environment, or you need something that's simple and fast, then Zig is worth a look. And of course, Zig is worth a look because it's very simple and easy to learn so you can be up and running very quickly.
I wouldn't put those two languages in the same class. Go has about the performance of Java/C# with about the memory usage of C++. It's much closer in performance and resource efficiency to those languages than it is from a scripting language like Python.
The syntax of Zig appears to be quite user friendly but not sure if there are hidden pitfalls.
I am excited for Zig but careful in adopting new languages but if this takes off, what sort of changes might we see? Cheaper C/C++ programmers?
It depends on what you are doing.
If your python code is mainly calling out to highly optimized math routines, well, not much.
If you are doing lots of computationally intensive work, quite a bit.
If your Python code does a lot of allocating of memory, throws around lots of large objects, or has to parse a lot of inputs, then moving to a native language can be very beneficial.
> I am excited for Zig but careful in adopting new languages but if this takes off, what sort of changes might we see? Cheaper C/C++ programmers?
Newer more powerful languages tend to result in more complex software being written. Ignoring CPU and memory limits, no one would have been capable of writing a modern AAA game with the tools available back in 1992.
Zig is very likely to appeal to people who loves C, or newcomers who are looking for performance.
The quality-of-life gains might lead to marginal changes in productivity/cost, but I would not bet on it.