[1] https://github.com/commonmark/cmark
[2] https://github.com/tiehuis/tiehuis.github.io/blob/38b0fd58a2...
105 karma · joined November 17, 2016
marc@tiehu.is
https://tiehu.is
https://github.com/tiehuis
[1] https://github.com/commonmark/cmark
[2] https://github.com/tiehuis/tiehuis.github.io/blob/38b0fd58a2...
The sub-items of this task are still valuable to complete even if this overarching proposal were declined. Most of them are not being completed as a prerequisite for this proposal. There are benefits gained even without the full removal of LLVM.
I can really respect the really wide-reaching views and goals of Andrew with zig and in proposals like this, even if I don't agree.
[0] https://github.com/ziglang/zig/commit/c51288f1f6be20be9f162c...
[1] https://github.com/ziglang/zig/pull/13821#issuecomment-13448...
LLVM is being placed as a an optional component in the future [2]. Using arocc to fill this gap would mean that seamless C compilation could still be done, without LLVM being involved.
I have previously used kcov [1] to perform coverage testing which due to zig having good DWARF debugging information worked with no fanfare. However, I'm not too familiar with the shortcomings of this method vs. the more fully instrumented binary approach provided.
The code itself should be up to date for most, but I'm aware of a few examples than need touch-ups. Unfortunately there is a bit more boilerplate compared to other languages for some simple tasks, since Zig requires you to be quite explicit about things (e.g. i/o, memory allocation).
Regarding the rejection of '\r', there is the intention for `zig fmt` to handle these minor issues and reformat code as needed which hopefully reduces this barrier. There was a long issue regarding hard tabs with discussion here: https://github.com/ziglang/zig/issues/544
const std = @import("std");
const wl_list = struct {
prev: *wl_list,
next: *wl_list,
fn init(list: *wl_list) void {
list.prev = list;
list.next = list;
}
fn insert(list: *wl_list, elem: *wl_list) void {
elem.prev = list;
elem.next = list.next;
list.next = elem;
elem.next.prev = elem;
}
};
const Data = struct {
data: i32,
link: wl_list,
fn init(data: i32) Data {
return Data{
.data = data,
.link = undefined,
};
}
};
pub fn main() void {
var foo_list: wl_list = undefined;
foo_list.init();
var e1 = Data.init(1);
var e2 = Data.init(2);
var e3 = Data.init(3);
foo_list.insert(&e1.link);
foo_list.insert(&e2.link);
e2.link.insert(&e3.link);
var entry = @fieldParentPtr(Data, "link", foo_list.next);
while (&entry.link != &foo_list) : (entry = @fieldParentPtr(Data, "link", entry.link.next)) {
std.debug.warn("{}\n", entry.data);
}
}
[1] https://ziglang.org/documentation/master/#fieldParentPtrFor example, the following code
fn errorOrAdd(a: u8, b: u8) !u8 {
if (a == 2) return error.FirstArgumentIsTwo;
if (b == 2) return error.SecondArgumentIsTwo;
return a + b;
}
pub fn main() void {
if (errorOrAdd(0, 0)) |ok| {
} else |err| switch (err) {
// will force a compile-error to see what errors we haven't handled
}
}
emits this error at compile-time: /tmp/t.zig:13:18: error: error.SecondArgumentIsTwo not handled in switch
} else |err| switch (err) {
^
/tmp/t.zig:13:18: error: error.FirstArgumentIsTwo not handled in switch
} else |err| switch (err) {
^
You can use the global `error` type which encapsulates all other error types if needed,
replacing `errorOrAdd` in the previous `main` with the following function fn anyError() error!u8 {}
now requires an else case, since it can be any possible error. /tmp/t.zig:13:18: error: else prong required when switching on type 'error'
} else |err| switch (err) {
^
This works pretty well and is very informative in most cases. The tradeoffs
are it can make some instantiation of sub-types a bit clunky [2] and you need
to avoid the global error type everywhere. The global error infests all calling
functions and makes their error returns global as a result. You can however
catch this type and create a new specific variant so there is a way around this
for the caller at least.Do remember that Rust allows passing information in the Err variant of a Result while Zig's error codes are just that, codes with no accompanying state.
[1] https://github.com/tiehuis/zig-bn/blob/3d374cffb2536bce80453...
[2] https://github.com/tiehuis/zig-deflate/blob/bb10ee1baacae83d...
Sure. I think that's more an effect of the choice of keyword defaults here. A straight union is very uncommon and is typically solely for C interoperability.
> And the "more powerful" is about the fact that a Rust enum can actually carry data, while it doesn't seem to be the case with Zig.
A tagged union can store data as in Rust. See the examples in the documentation [1]. Admittedly Rust's pattern matching is nicer to work with here.
To summarise the concepts:
- `enum` is a straight enumeration with no payload. The backing tag type can be specified (e.g. enum(u2)).
- `union` is an unchecked sum type, similar to a c union without a tag field.
- `union(TagType)` is a sum type with a tag field, analagous to a Rust enum . A `union(enum)` is simply shorthand to infer the underlying TagType.
> So, it sounds like deallocations are not checked by default, right?
If referring to if objects are guaranteed to be deallocated when out of scope then no, this isn't checked. There are a few active issues regarding some improvements to resource management but it probably won't result in any automatic RAII-like functionality. This is a manual step using defer right now.
> At first glance, Rust's `enum` looks safer and more powerful than Zig's `union` + `enum`, while Zig's `union` + `enum` appears more interoperable with C.
A `union(TagType)` in Zig is a tagged union and has safety checks on all accesses in debug mode. It is directly comparable to a Rust enum. Any differences are probably more down to the ways you are expected to access them and Rust's stronger pattern matching probably helps some here.
> I don't see traits in Zig.
Nothing of the sort just yet although it is an open question [1]. Currently std uses function pointers a lot for interface-like code, and relies on some minimal boilerplate to be done by the implementor.
See the interface for a memory allocator here [2]. An implementation of an allocator is given here [3] and needs to get the parent pointer (field) in order to provide the implementation.
It isn't too bad once you are familiar with the pattern, but it's also not ideal, and I would like to see this improved.
> I don't see smart pointers in Zig, and more generally, I have no idea how to deallocate memory in Zig.
You would use a memory allocator as mentioned above and use the `create` and `destroy` functions for a single item, or `alloc` and `free` for an array of items. Memory allocation/deallocation doesn't exist at the language level.
> I don't see anything on concurrency in Zig's documentation.
There are coroutines built in to the language [4]. This is fairly recent and there isn't much documentation just yet unfortunately. Preliminary thread support is in the stdlib. I know Andrew wants to write an async web-server example set up multiplexing coroutines onto a thread-pool, as an example.
> Zig supports varargs, Rust doesn't (yet).
It's likely that varargs are instead replaced with tuples as a comptime tuple (length-variable) conveys the same information. I believe this fixes a few other quirks around varargs (such as using not being able to use varargs functions at comptime).
[1] https://github.com/ziglang/zig/issues/130
[2] https://github.com/ziglang/zig/blob/15302e84a45a04cfe94a8842...
[3] https://github.com/ziglang/zig/blob/15302e84a45a04cfe94a8842...
[4] https://github.com/ziglang/zig/blob/15302e84a45a04cfe94a8842...
That being said, some other potentially interesting safety aspects that are present (or are being explored) to give some idea of the target audience:
- compile-time alignment checks [1]
- maximum compile-time stack usage/bounding [2]
I would expect safety at the end of the day in Zig will be more similar to
modern C++ and its smart pointers, alongside (optional) runtime checks,
than a full lifetime system. Will have to see what the future holds.It is also easy to overlook how well optimized GMP is across a wide range of less common architectures and chips and I wouldn't be surprised if my particular implementation lost a bit of ground on other architectures like ARM (would be a good thing to test).
Zig doesn't have a default memory allocator. Allocators instead are expected to be passed as an argument to functions as they need them. This makes it trivial to replace an allocator with something custom or use multiple different allocators within a small code block.
A contrived example:
const std = @import("std");
pub fn GiveMeAnInt(alloc: &std.mem.Allocator) -> %&u32 {
return alloc.create(u32);
}
test "using two allocators" {
const int1 = try GiveMeAnInt(std.heap.c_allocator);
*int1 = 2;
// Would usually store the allocator with the type on construction.
defer std.heap.c_allocator.destroy(int1);
const int2 = try GiveMeAnInt(std.debug.global_allocator);
*int2 = 2;
} #include <stdint.h>
#include <string.h>
typedef struct {
int32_t a;
int32_t b;
} Foo;
int main(void)
{
uint8_t array[1024];
memset(array, 1, sizeof(array));
Foo *foo = (Foo*)(&array[0]);
foo->a += 1;
}
Using clang 3.8.0-2. Compiling examples with `clang -S llvm-ir`.It appears that the array is aligned with the minimum ABI requirement 16 by default? May be a note of this in the standard, can't recall of the top of my head.
%array = alloca [1024 x i8], align 16
...
%6 = load i32, i32* %5, align 4
...
store i32 %7, i32* %5, align 4
We can also explicitly specify the alignment required in C11. #include <stdalign.h>
#include <stdint.h>
#include <string.h>
typedef struct {
int32_t a;
int32_t b;
} Foo;
int main(void)
{
uint8_t alignas(alignof(Foo)) array[1024];
memset(array, 1, sizeof(array));
Foo *foo = (Foo*)(&array[0]);
foo->a += 1;
}
Results in the following IR. %array = alloca [1024 x i8], align 4
...
%6 = load i32, i32* %5, align 4
...
store i32 %7, i32* %5, align 4It seems as there have been some attempts to revitalize it [2] but these are targeting older versions and would probably require a lot more work to ever get back in-tree.
If you wanted to send data you would need some other means like you suggest. I'd be interested in finding an ergonomic solution to this but it probably wouldn't be at the language level.
Error values under the hood are just unsigned integers and are returned on the stack. In fact, the granularity at which allocators are exposed in the stdlib makes any possible dynamic allocation very explicit in the language.
This link [1] provides an overview of errors and some of the surrounding control flow.
For zig it is not all or nothing of course. It is pretty easy to link and use c if wanted. The following for example asks the compiler to link against libc.
zig build_exe main.zig --library c
One other benefit that zig gets is more compile-time execution opportunity. The math library for example being written in zig means that we can determine `math.log(7)` and use it for compile calculations when using libm we couldn't.Hassle-free error management
Consider you write a function
fn div(a: u8, b: u8) -> u8 {
a / b
}
If later you find an error case that needs to be handled updating the code
usually will not require much extra addition. error DivideByZero;
fn div(a: u8, b: u8) -> %u8 {
if (b == 0) {
error.divideByZero
} else {
a / b
}
}
The caller can choose to ignore errors using `%%div(5, 1)` or can propagate
errors to the caller, similar to Rust's `try!`, `?` using `%return div(5, 2)`.I find this so easy that I'm much more inclined to think about edge cases and handle errors up front. I find when writing Rust the extra setup and management of errors adds a fair bit of tedium (although to be fair, with error-chain and proper setup at the beginning of a project this isn't too bad).
Compile-time programming
Zig has pretty strong compile-time programming support. For example, its printf formatting capability is all written in userland code [1]. It doesn't at this moment support code-generation like D's mixins but I personally have not found this too problematic.
Generic functions can be written in a duck-typing fashion. With compile-time assertions the inputs can be limited to what they need pretty clearly and the errors during usage are pretty self-explanatory.
error Overflow;
pub fn absInt(x: var) -> %@typeOf(x) {
const T = @typeOf(x);
comptime assert(@typeId(T) == builtin.TypeId.Int); // must pass an integer to absInt
comptime assert(T.is_signed); // must pass a signed integer to absInt
if (x == @minValue(@typeOf(x))) {
return error.Overflow;
} else {
@setDebugSafety(this, false);
return if (x < 0) -x else x;
}
}
Zig doesn't have any form of macros. Everything is done in the language itself.