Fuzzing Zig Code Using AFL++
ryanliptak.com
ryanliptak.com
I know it's a small thing but strings are common enough that the little ergonomic costs add up.
const fuzz_exe_path = try std.fs.path.join(b.allocator, &[_][]const u8{ b.cache_root, fuzz_executable_name });
// We want `afl-clang-lto -o path/to/output path/to/object.o`
const fuzz_compile = b.addSystemCommand(&[_][]const u8{ "afl-clang-lto", "-o" });
can be updated to newer syntax: const fuzz_exe_path = try std.fs.path.join(b.allocator, &.{ b.cache_root, fuzz_executable_name });
// We want `afl-clang-lto -o path/to/output path/to/object.o`
const fuzz_compile = b.addSystemCommand(&.{ "afl-clang-lto", "-o" });
docs: https://ziglang.org/documentation/master/#Anonymous-List-Lit...added in zig 0.6.0: https://ziglang.org/download/0.6.0/release-notes.html#Tuples...
IMHO for a systems programming language it's the right decision to not have a builtin string type (and instead treat strings as bags of bytes), proper string support is better provided by a library (or even better several specialized string processing libraries).
Obviously Zig isn't going to offer what Rust does here because it involves safety guarantees, but the assumption could be exactly the same, this is a UTF-8 string, and even in systems programming that's a valuable built-in.
Twenty years ago it's unclear what you should pick here, and so "bytes" ends up being a reasonable compromise. But today it isn't unclear, the answer will be UTF-8. Languages like C++ are stuck offering people a 16-bit character type if they want one because they were around in the 1990s when that seemed potentially viable and they're stuck in a compatibility quagmire, not because it is any actual use today.
const str = []const u8;
...in Zig isn't it? Both are slices (aka pointer/size pairs). Rust's str seems to have some additional string-related functions, but in Zig those could be provided in the stdlib (along with the 'str' type wrapper), and then on the next level a String-like type which includes memory management.PS: Ah ok, one important difference is that Rust's str is guaranteed to be valid UTF-8.
I like Rust's approach of having purpose-specific string types that may differ (e.g. if your OS is using UTF-16 encoded strings but the C strings are just arrays of bytes). It guarantees that the programmer checks/converts the string is in/to the correct form, or just let the program panic with .unwrap(), when interfacing with the external world. It makes the program safe and makes the potentially expensive string conversion calls explicit. Also, I think valid UTF-8 strings is a good default.
I haven't used Zig yet (I haven't found a good use case for it), it seems closer to C both in spirit and in implementation while removing some foot guns.
str could have a part of "core" library.
I thought it was just a `struct str([u8]);` (a dynamically-sized type).
You're correct that Rust has a bunch of actual type rules here to say that a str really is UTF-8 and not some random bytes -- and Zig presumably wouldn't do that, but by explicitly building in a named str type, rather than just shrugging and saying []const u8, there'd at least be a social convention that str is actually, you know, a UTF-8 string, not just some bytes.
And yes, obviously the memory managing high level String with features like concatenation and mutability is way more heavyweight and shouldn't be built-in to the language fundamentals of something like Zig, it's just the cheap "size + pointer" immutable string slice that I believe is worth building in explicitly now that I've seen Rust do it and after decades of experience noticing that software almost invariably has string literals representing human text (thus suitable for UTF-8) that we shouldn't muddle with slices of arbitrary bytes.
Btw one thing that Zig has over Rust is that string literals are zero-terminated for C-compatibility (so the proper type is actually "[:0]const u8" - which is a "sentinel-terminated array"). This is in addition to strings being slices (pointers/size pairs), so despite being zero-terminated it doesn't inherit C's problems.
And then also the same misfire can hurt in the other direction. Since Rust's str is really a UTF-8 slice, the zeroes are just zeroes, valid UTF-8 characters, you can ask if a Rust str contains(0 as char) because that's a completely reasonable question - but if your code thinks zero might be a "terminator" it can't actually handle arbitrary UTF-8 this way.
Actually I think this makes the call for an actual str built-in even more appropriate because Zig's owners would need to wrestle with this and work out what they intended whereas just saying []const u8 allows you to hide inconsistencies in the libraries and people will get bitten.
https://ziglang.org/documentation/0.8.1/#Sentinel-Terminated...
pub const string = []const u8;
This lets me do: &[_]string{“hi”}
A little better, but not great.