Zig 0.9.0
ziglang.org
ziglang.org
I can't wait for 1.0! Or at least until the docs are a bit more accessible... :)
That's all that's needed (and should be implemented) on the language level, everything else should go into the standard library, and additional specialized string processing libraries (because string handling can never be a "one-size fits all" solution, it's too complex for that).
The downside is of course the fact that you have to account for encodings in your language now, but picking the one and only sensible encoding really shouldn't be a problem in 2021.
Precisely. That's why I think representing strings as character slices and letting third party libraries handle them is not a good solution.
Deciding what features to put in the language vs. the stdlib vs. third party libraries is one of the hardest parts of language design. Personally I believe strings are important and frequent enough they deserve special treatment at the language level.
Edit (late addition):
I think not treating strings specially is mostly fine in C, but Zig seems to aim at being a little less lowlevel.
Zig does not do this. It represents strings as byte slices. A UTF-8 character could be multiple codepoints with each codepoint being multiple bytes.
A big thing about Zig is not hiding complexity. If UTF-8 were implemented at a language level (whatever that means), then "language level" string operations would be non-linear, which would be very non-Ziggy. I could see value in a standard library UTF-8 implementation, but a LOT of forethought would need to be put into it. I think keeping UTF-8 string manipulation at the third-party library level is a good choice for now. Maybe once the language is finalized, the ecosystem is more developed, and lessons have been learned from the third-party libraries, then the standard library can implement this.
Sorry, byte slices is what I had in mind.
I'm not talking about language level string operations as in concatenation with + or something like that at all, because that certainly wouldn't make sense in language like Zig.
I'm not advocating for string functionality in the language, I'm advocating for a way to not allow byte slice functionality on a thing that is clearly not a byte slice.
This already exists in the form of structs or opaque types. Both of these approaches would end up being implemented in "userspace" anyways, whether that's standard library or third-party.
However, (UTF-8) strings are byte slices. You can do simple manipulation with them as byte slices safely and validly. Split on spaces? Sure. Tokenize? Sure. Find substring? Sure. You can't do things that depend on say UTF-8 graphemes, but you can safely do most things that depend on bytes. For most purposes, treating strings as byte slices is the safest and correct approach.
That's a nice analogy.
> Doing find substring by find byte subsequence won't behave correctly in many cases, where semantically equivalent strings have multiple different bytesequence representation.
Unfortunately that's nearly impossible to do sanely in the general case, no matter how the string is represented.
Not necessarily the shortest (NFC means not using composed characters from later revisions of the standard), and you only get a normalised representation if you've actually normalised it - if you've just accepted and maybe validated some UTF-8 from outside then it probably won't be in normalized form. IMO it's worth having separate types for unicode strings and normalized unicode strings, and maybe the latter should expose more of the codepoint sequence representation, but I don't know if any language implements that.
Like subslicing. And accessing individual bytes in it.
That's what the parent meant. Char(aracter) is a byte in C.
And regardless of where they put the code, this is something that needs to be done by the core team. Otherwise you will end with too many string libraries, all of them trying to solve a particular problem and doing bad at everything else, with bad documentation and different API styles and making difficult to mix code that uses two different libraries.
I would even argue that having string handling in a standard library (or language) has the potential to cause a net increase in bugs, because of people thinking they are handling strings when actually they are just screwing around with codepoints. Go's string handling is completely broken, for example. As a result of strings in the language, Go programs tend to be more broken than C programs in terms of string handling.
That sounds nice and all, but wr have 50+ years of protocols and formats and APIs built up around strings. Unless you're just writing code to run on a small microcontroller, you need to be able to parse and generate strings. So its going to be pretty frustrating not to have good support for them, or to have every codebase use its own libraries and idioms for working with strings.
Edit: so far these examples have been given:
* HTTP: wrong. the spec does not tell you to decode any strings
* CSV: wrong. the spec does not tell you to decode any strings, nor is it necessary to have any unicode awareness in order to properly read and parse the data or deal with the delimiters.
Additional request: if you attempt to provide a counter-example, please also point to the place in the spec where it tells you to decode a string.
This is a validation pass only and it doesn't make any meaning of the code points, except to validate that none of them are surrogate code points (disallowed in UTF-8).
1. "JSON syntax describes a sequence of Unicode code points. JSON also depends on Unicode in the hex numbers used in the \u escapement notation."
2. "A JSON text is a sequence of tokens formed from Unicode code points that conforms to the JSON value grammar."
3. "A string is a sequence of Unicode code points wrapped with quotation marks (U+0022). All code points may be placed within the quotation marks except for the code points that must be escaped: quotation mark (U+0022), reverse solidus (U+005C), and the control characters U+0000 to U+001F. There are two-character escape sequence representations of some characters."
4. "Any code point may be represented as a hexadecimal escape sequence. The meaning of such a hexadecimal number is determined by ISO/IEC 10646. If the code point is in the Basic Multilingual Plane (U+0000 through U+FFFF), then it may be represented as a six-character sequence: a reverse solidus, followed by the lowercase letter u, followed by four hexadecimal digits that encode the code point."
5. "Note that the JSON grammar permits code points for which Unicode does not currently provide character assignments."
JSON does require Unicode awareness, both for general parsing, and for correctly interpreting strings. Backslashes are allowed for special escape characters, which means that you must be aware of the format (and the encoding of the text) in order to be able to decode.
Note that JSON also doesn't specify a required encoding, only Unicode correctness, so a parser may need to be able to handle multiple Unicode encodings and differentiate between them.
The spec doesn't specify what to do with code points which are not understood as Unicode (given especially the allowance for unassigned characters), but explicitly-invalid Unicode should be rejected.
[1] https://www.ecma-international.org/wp-content/uploads/ECMA-4...
If you have valid UTF-8 already, then yes, the task is a lot easier. But depending on the level at which you're parsing, this might not be the case — i.e., if you're writing a JSON parser from the ground up, you do need to know what UTF-8 and Unicode are, and will need to validate the input data.
> Converting the unicode character escape codes to utf-8 would require knowledge of utf-8 encoding
Agreed. Even if you're not working at the "array-of-bytes" level, you will need to be able to parse and translate "\u..."-style strings into the appropriate output character encoding.
> but this unescaping is not a feature that would be provided by the language regardless.
I'm not sure we're talking about this being handled at the language level. This translation is something that would likely be offered at the parser level (working with the features offered by the standard library), but the parser does need to know about it — and does need to be able to work with strings at a granular level to be able to parse it out. By definition, it cannot leave the input data as an undecoded bag of bytes.
Note, too, that the JSON spec does not specifically require UTF-8. UTF-16 is a completely valid encoding for JSON (though much less common than UTF-8), in which case none of these characters are an ASCII subset, and greater awareness is needed to be able to handle this.
But all it's doing here is taking a hex string (which is entirely ASCII) and converting it into the respective hex representation. Since ASCII translates unambiguously to bytes, it doesn't really matter if `str[0]` is operating on a byte stream, codepoint stream or grapheme stream, because in utf8, they're all the same thing as long as we're within the ASCII range.
Where things get hairy is stuff like `str.reverse()` over arbitrary strings that may or may not be in ASCII. This repo[0] talks about some of the challenges associated with conflating characters with either bytes or codepoints. The problem is that programming languages often approach strings from the wrong angle: you can't just tack on handling of multi-byte codepoints on top of ascii handling; you lose O(1) random access and you don't actually model the linguistic domain properly by doing so, because in the first place, humans think of characters not in terms of bytes or codepoints, but in terms of grapheme clusters. Clustering correctness falls deep in the realm of linguistics, and is therefore arguably more suitable to be handled by a library than a programming language.
> hex string (which is entirely ASCII)
My point is that JSON doesn't need to be UTF-8 or a superset of ASCII to be valid. It can be any representation of Unicode, including UTF-16, UTF-32, GB 18030, etc.; so long as the text is is comprised of Unicode code points in some Unicode transformation format, the JSON is valid.
As I said in the parent comment: if you are working within UTF-8 exclusively, and can assume valid UTF-8, then great! But this isn't necessarily true, and in some cases, you will still need to care about the encoding.
(Either way, this starts straying slightly from the more general discussion at hand: regardless of the encoding of the string, you will still need an ergonomic way of interacting with the contents of the data in order to meaningfully parse the contents — even past the hurdle of decoding from arbitrary bytes, you still need to manipulate the data reasonably. In some cases, this means working with a buffer of bytes; in others, it makes sense to manipulate the data as a string... In which case, you may run into some of the string manipulation ergonomic considerations being discussed around these comments.)
Sure, it can also be gzipped, encrypted, etc but that goes back to the point that there's nothing inherently special about JSON as it relates to encoding to a byte stream. All there is to it is that somewhere in a program there's an encode/decode contract to extract meaning out of the byte stream, and in a protocol one most likely only looks at byte streams as sequences of bytes (because performance-wise, it doesn't make sense to look at payload size in terms of number of codepoints/graphemes at a protocol level)
[1]: https://eclecticlight.co/2021/05/08/explainer-unicode-normal...
I could create a standard library called "arrayofbytes" that lets you, for example:
- Search an array of bytes for a smaller array of bytes
- Split up an array of bytes based on an array of delimiter bytes
- Selectively convert an array of bytes in range 41-5A to bytes in the range 61-7A
But wouldn't it be more appropriate to describe these as "string handling" functions using the generally accepted terminology of our profession?
Case sensitivity, ordering, and case changing rules, are reasons not only to have strings, but extensive culture support built into them. String types also let you compare utf16 bytes to utf8 bytes, for example.
To think that strings are opaque bytes is massively naive. It presumes one encoding exists, the input is always valid, in addition to what is said above.
An example of what you're setting yourself up for, is an exploit based around differences in handling invalid utf8.
> the spec does not tell you to decode any strings
Which "spec" is that? My user ask for csv, see csv, upload csv, do SEARCH on CSV, transform them, etc. Even do some scripting on them.
Nowhere, EVER "opaque encoded bytes" is on the spec of the users!
Which should be ascii though no?
CSV is defined in terms of characters, not bytes, so CSV parsing does require you to be encoding-aware: if your CSV file is encoded as UTF-16, bytewise parsing will destroy the data.
That doesn't make sense. If you're working with utf16, why would you slice bytewise? That's like slicing a zip file bytewise and wondering why it got corrupted.
The whole point of the argument for string support at a library level (rather than assuming some sort of equivalence between stringness and its underlying byte buffer at the language level) is that fixed width bytes fundamentally cannot model human language characters unambiguously because what a "character" is depends on the encoding/decoding contract of the program manipulating the byte buffer.
Assuming equivalence between 0x2C and `,` stems from the ancient history of ASCII, english and usage of C `char` as a mechanism to squeeze performance out of string operations by not properly supporting the full gamut of valid human language characters.
For a low level language that might be used to implement protocols, it totally makes sense that foo.len is length in bytes, because you're pretty much never going to want to know number of grapheme clusters at a protocol level. It doesn't make sense for a language level .len to be length in terms of codepoint count because that assumes encoding, which is fundamentally a business logic level concern.
When have you last browsed the web using telnet? It's all plain text, so you should be able to, right? Press Ctrl-U right now and get a taste of how readable it is, and even that is with the benefit of syntax highlighting.
On the machine level, processing binary data is effortless. At most, you have to swap the bytes around to conform to wrong-endian network byte order. Check that the length of the data fits into your buffer. Not a problem with Zig, or any language that doesn't silently ignore integer overflow as C does!
Scripting languages make handling strings seem easy, because that's what they were built for. And of course in most languages there are "mature" libraries for parsing JSON or XML. But all the fundamental complexity is still there. Each layer of abstraction may introduce some bug that can be exploited. Compared to classical buffer overflows, it may take several more steps to gain arbitrary code execution, but every bit of extra code you depend on still increases the attack surface.
Ideally, only user interface code should be dealing with strings at all, but legacy protocols may unfortunately require it too. This should be isolated as much as possible from the rest of any application, and not dictate what features are "first class" in a programming language.
Swift is pretty pedantically strict about Unicode correctness, and avoids some pitfalls which lead to incorrect handling (e.g., integer-based indexing). This means that the traditional way many developers might be used to interacting with strings is cumbersome and annoying — but once you get past the initial hurdle*, most operations are actually (1) really easy, especially when expressed in terms of generic Sequence/Collection operations, and (2) much more difficult to get wrong.
*I think the largest part of that hurdle is overcoming what you may have gotten used to from other languages, i.e., treating strings as an array of "characters", for some language/library definition of "character" (whether bytes, code points, etc.). It's relatively rare that you actually care about indexing into an arbitrary spot in a string: instead, combinations of slicing operations (including `prefix(_:)`, `dropFirst(_:)`, `take(while:)`, etc.) and generic Collection operations will get you what you want. Things like `.reversed()`, `.sorted(), and `.shuffled()` all work trivially correctly too (since you're not operating on a bag of bytes), and it's exceedingly rare that user input will confound the operations you might need to perform. (Exception: operations like case folding and collection, which are locale-specific, need special handling through a framework like Foundation.)
To be clear: not everything is sunshine and roses, but an amazing amount of functionality "falls out" of basic protocol conformances on String, and its exposure as a Collection of grapheme clusters.
Given a specific string manipulation task, I'd be happy to provide an example of what it might look like in Swift!
How would you safely get the nth index of a string, clamped to the valid indexes? So if the nth index is out-of-bounds you get the first/last index instead?
extension String {
func character(atClampedIndex index: Int) -> Character? {
guard !self.isEmpty else {
return nil
}
let clamped = max(0, min(index, count - 1))
return self[self.index(startIndex, offsetBy: clamped)]
}
}One thing to note: aside from programming interviews, these operations are fairly rare. And that's a good thing, because none of these produce results that are very intuitive, because they are not very well defined on strings in general (I don't fault Swift for this, but it's just a general problem with text). Using any of these to create a new String may cause entirely new characters to show up, or the length of the text to change. So Swift actually doesn't expose these as "string" operations, but operations on the characters themselves; in each case returning a new collection of characters that is not a String. Now, you can reconstitute them into a String pretty easily, but you should keep the this in mind when doing so.
IMO they are a great example of how Golang is a pragmatic rather than "clever" or "pure" programming language.
s := "naïve"
// bad
fmt.Println(len(s)) // 6
fmt.Println(string(s[2])) // Ã
// good
r := []rune(s)
fmt.Println(len(r)) // 5
fmt.Println(string(r[2])) // ï
https://go.dev/play/p/YbMo49wU7vuI suggested no such thing.
> individual codepoints inside grapheme clusters
That's less severe than invalid codepoints.
Perhaps the whole thing whichever way it is represented should not be mutable given that there's no way to make it mutable in a sensible way?
I have lost count of how many times I have wanted to find substrings, transform cases, catenate strings, find patterns, substitute patterns. I'd be happy to do that in a language that didn't permit me to index the underlyinc characters or the bytes. Keep 'em opaque, sure. But I think it would be a mistake for a language not to have an idiom with a favorite library to perform these operations on encoded text. If resolving library dependencies is easy enough, then it doesn't need to be "standard" but it should be "the defacto standard." And if it turns out the defacto standard stagnates and doesn't keep up with the needs of developers, a new one can come take its place.
>>> "ñ"[0]
'n'
>>> "ñ"[1]
'̃' >>> "ñ"[0]
'ñ'
>>> "ñ"[1]
IndexError: string index out of range
Which is what I would expect.Ah, interesting, I suppose then there is a difference in the way Linux handles strings? Didn't realize that, very unfortunate if true. I am running MacOS.
>>> b'n\xcc\x83'.decode()
'ñ'
>>> b'n\xcc\x83'.decode()[0]
'n'
>>> b'n\xcc\x83'.decode()[1]
'̃'
But I agree, it's rare case when you need to deal with non-normalized data.import unicodedata list(unicodedata.normalize('NFD', "ñ")) >> ['n', '̃']
list(unicodedata.normalize('NFC', "ñ")) >> ['ñ']
both are correct, the issue is that unicode allows accented letters to be written as _accented_letter_ or _letter_, _accent_. The idea of "character" in uncicode is not very useful, most of the time you will want graphemes, not codepoints. User-friendliness wise, this is what Python should use (another rant - strings should not have length method, they should have byte_length, codepoint_length and grapheme_length).
Classic Andy, shitting on other language with no references or examples. Go has some of the best string handling I've used. Seamless byte, rune, string conversion. Simple iterating and slicing. Plus helpful tools like strings.Builder and strconv.AppendInt. while Zig has nothing.
Python has the GIL.
Go itself has a few such cases (usually revolving around "NIH" and misguided simplicity).
Zig has the prejudice about proper string handling.
You are welcome to show fast single-core python interpreter without GIL. Coreteam will gladly accept your patches.
I always like to ask for what kind of task one needs GILless python?
CPU bound? If you are using native python code for CPU bound tasks you are already in a bad place. C extensions can release GIL. For example numpy.
What else?
> CPU bound? If you are using native python code for CPU bound tasks you are already in a bad place.
While I think you may have misread my comment for a value judgement on Python's GIL, I don't see as particularly useful dismissing major potential (multithreaded) performance gains for a slow language just because there's faster languages. Languages that are better in some way - like speed - should stand as a benchmark and a goal, not a reason to give up improvements.
https://godocs.io/strings#Builder
ArrayList cant do that. And strconv.AppendInt can convert a number to byte slice, then it appends to an existing byte slice. formatInt cant do that.
"appends to an existing byte slice" is not a completely coherent thing to ask for. If I give you a byte slice, the byte after the end it might belong to something else. If you want to deallocate my byte slice and make a new, replacement byte slice, you'll need an allocator to do that. An ArrayList(u8) has a byte slice, and knows how much of it is unused, and has an allocator so it can make a new larger byte slice if needed, and exposes a writer so that you can call std.format.formatInt to write into the byte slice (or allocate a new byte slice, copy the entire contents, and write into the new one, if appropriate).
> strconv.AppendInt can convert a number to byte slice, then it appends to an existing byte slice. formatInt cant do that.
As we have learned just now, we actually can use formatInt to append a formatted int to an existing byte slice (to the extent that that's a meaningful thing to ask for), by passing it the writer exposed by an ArrayList(u8)!
Is your complaint that you're not aware of any function in Zig that takes a u32 representing a Unicode codepoint and encodes it to UTF-8 and writes it to a writer (which is what "appending a rune" seems to mean)?
Is your complaint that people do not, by convention, allocate new, larger byte slices without an explicit reference to an allocator?
Is your complaint is that Zig does not have a garbage collector?
Can you share a Go program that uses the Go stdlib functions and cannot be trivially ported to use the Zig stdlib functions instead, for some reason other than that Zig programs must decide where the bytes will live, and Go programs need not do that?
Heres an easy one:
My original comment, was that Andrew has a habit of shitting on other languages without proper references or examples. Nothing you can really say is going to change Andrews behavior, so maybe you should stop, unless you can justify Andrews comments.
Most of the difference comes from the demand that we represent unicode codepoints as integers at some point in the program, which is a nonsensical thing to do, because unicode codepoints don't correspond to anything useful in the actual text being represented.
You seem to have a habit of making false claims about the standard libraries of languages you dislike, and when pressed on the matter ask other people to do your homework. I certainly should stop doing other people's homework.
So for Andrew to shit on Go string handling with not a single example is rude, and frankly just wrong as I have demonstrated.
https://play.golang.com/p/Dla3sXciYXC
I also think string handling in Go is pretty sane. But range vs indexing on strings is something you need to be aware of.
String literally are not really UTF-8. Rather, zig source files are UTF-8 (by definition), and string literals are u8 literals; putting some UTF-8 between quotes just puts the literal UTF-8 encoded text into the literal because those are the bytes that are in the source file. In particular, string literals can contain arbitrary binary data and null bytes by way of \x00 escapes.
For Unicode handling there's some basic stuff in std.unicode (conversion between different UTF encodings, checking validity, decoding to codepoints etc.). This is used e.g. on Windows for checking filesystem paths. I don't general libraries of encodings is really that important today, iso-8859 and shift JIS might be useful sometimes, but everything else probably doesn't need to bloat up a standard library (iirc Python's codecs package, which contains dozens upon dozens of encodings, is like a third of the standard library by size).
it doesn't look like there will be language support for things like codepoints or grapheme indexing or treatment of strings as anything but byte arrays, so ddevault is sad.
there is intention from andrewrk and jecolon to provide such features in the standard library before 1.0 release.
downside to library vs lang support that is you can expect a good chunk of programmers to ignore the less-ergonomic library features, use the language features at hand, and handle strings incorrectly as byte arrays.
And this is met with answers like "just avoid string handling" from the language designers...
It's probably because they don't work in any related domain, and even less so have to do with international strings (except as byte buckets they don't care about and don't have to do anything do).
When you send your unicode string to an external system (for example a storage server with a database) and latter retrieve the string, only to find out that it has been normalized differently so it no longer match byte-for-byte what is stored in your program, and all of a sudden strcmp no longer works.
Or all kind of weirdness like that because every system outside of your program will handle unicode differently, and you will need to adapt to them, and having a string library to do most of the heavy-lifting will avoid the need for every user to rewrite a library from scratch.
Not to mention that for every developer which rewrite unicode handling functions, you will probably end up with a function with subtly different behaviors from others, which aggravate the problem for others when they will try to communicate with your system.
I hadn't, and as long as I control the data I'm displaying, I won't have to.
> Or all kind of weirdness like that because every system outside of your program will handle unicode differently
Blame those systems, not me.
What you suggest is surrendering to the state of affairs, which we collectively self-inflicted.
When I have to deal with normalization issues and have to interface with external systems, I can still go looking for a library if I don't feel like implementing it on my own (which is likely).
But unless I need to do normalization, I'm way worse off with a complicated library than with just doing memcpy() or using a simple decode_codepoint() routine.
Yes I completely agree with you, and if you don't need it, any unicode handling library is overkill and add more headaches than simply handling utf-8 string as byte arrays.
I just wanted to insist on the fact that some people will have to deals with theses kind of issues. And these issues are self inflicted, but it gets worse every time someone try to reinvent the wheel or rely on byte array when they shouldn't.
Having a standard library in the language make the issue less worse: the core of the language still handle only byte arrays, and for the cases where it's not enough you still have only one library so you don't add your own subtly different mishandling of the standard by implementing your own.
So memcpy is fine, but that's about it: for example, please don't use strcmp when you need to sort data alphabetically and please don't try to reimplement the standard algorithms designed for that, otherwise you will be part of the problem.
"Which third-party strings lib of several half-complete incompatible libs" will be a much realer concern...
I've worked on a full-duplex file synchronization system that had to support cross-platform operating systems and file systems, across a variety of Unicode normalization schemes (and versions [1], which is why I introduced the question), and I'm personally satisfied that baking this into the language specification would be a mistake.
[1] For example, depending on the file system, there's simply no way to get the normalization right unless you reverse engineer the actual table they're using, or probe the file system to do the normalization for you.
[1]: https://github.com/bellard/quickjs/blob/master/libunicode.h
Assuming that you really want to use UTF-8 internally, which is probably a sensible choice, the reusable part of a string library is basically the UTF-8 encoded/decoder. A useful implementation of UTF-8 is about 100-200 lines, I could probably rewrite what I use in an hour or two without an internet connection. The rest of the work is integration stuff that doesn't make sense to put in a library IMO. The idea of a string library fits much better with garbage collected and scripting languages (which includes C++ with RAII mechanism, but consider that std::string and similar often cause bad performance).
Many programs, in particular non-graphical programs don't need any UTF-8 code at all - UTF-8 handling is basically memcpy().
argv to main is utf8 on my system.
On Windows, I believe it is Unicode converted to current codepage.
In any case I don't need to care about it since I can simply treat arguments as ASCII-extended opaque strings as described.
That sounds totally compatible with programs that don't know anything about utf8. Do programs need to normalize the utf8 you pass in before using it as an argument to open(2) or something?
Also, a lot of the time the problem isn't that people are using fundamentally incompatible string libraries, but that there isn't one correct answer to the question they're asking and they chose different ways to convert the question into code. A reasonable question to ask is "How many extended grapheme clusters are in this string?" The answer is "It depends on what font you plan to use to render it." Not great!
Some programmers would still like to e.g. write a reverse() function that returns ":regional_indicator_f::regional_indicator_r:" unmodified (because it is the French flag emoji) and returns ":regional_indicator_i::regional_indicator_h:" when given ":regional_indicator_h::regional_indicator_i:". If such people want to avoid having nonsensical behavior in their programs, the only solution available is to decide what domain the program should work on and actually deal with the complexity of the domain chosen.
For example, consider JS' introduction of String.normalize(). This is a slippery slope. It had a huge impact on Node's build process and binary sizes because now all the tables had to be shipped. But it's still broken in JS, because no matter the Unicode normalization support provided, it will never match the exact tables used e.g. in Apple's HFS.
I feel that by the time it gets to String.normalize(), it's too far gone.
There are plenty of loadable modules in C++ for Linux and BSDs, and plugins for Postgres, in places where there is no expectation of upstreaming them.
Zig is in a similar boat.
But of course it would be less of an upgrade, and the Zig parts would have to stay clearly segregated.
It’s a different tradeoff which may be worth it for some use-cases. But it’s definitely not obvious whether it is worse, as the alternative comes with a huge complexity on the language side.
People on HN find it interesting.
At least for TigerBeetle [1], a distributed database, the story was strong enough even last year that we were prepared to paddle out and wait for the swell to break, rather than be saddled for years with undefined behavior and broken/slow tooling, or else a steep learning curve for the rest of the project's lifetime. We realized that as a new project, our stability roadmaps would probably coincide, and that Zig makes a huge amount of sense for greenfield projects starting out.
The simplicity and readability of Zig is remarkable, which comes down to the emphasis on orthogonality, and this is important when it comes to writing distributed systems.
Appropriating Conway's Law a little loosely, I think it's more difficult (though certainly possible) to arrive at a super simple design for a distributed consensus protocol like Viewstamped Replication, Paxos or Raft, if the language's design is not itself also encouraging simplicity, readability and explicitness in the first place, not to mention a minimum of necessary and excellent abstractions. Because every abstraction carries with it some probability of leaking and undermining the foundations of the system, I feel that whether we make them zero-cost or not is almost besides the point compared to getting the number of abstractions and composition of the system just right.
For example, Zig's comptime encouraged a distributed consensus design where we didn't leak networking/locking/threading throughout the consensus [2] as is commonly the case in many implementations I've read, even in high-level languages like Go. It made things like deterministic fuzzing [3] really the natural solution. People who've worked on some major distributed systems in C++ have commented how refreshing it is to read consensus written in Zig!
Zig also has a different/balanced/all-encompassing approach to safety that resonates more with how I feel about writing safe systems overall: all axes of safety as a spectrum rather than as an extreme (this helps to prevent pursuing one axis of safety at the expense of others), safety also including things like NASA's "The Power of 10: Rules for Developing Safety-Critical Code" [4], assertions, checked arithmetic (this should be enabled by default in safe builds, which it is in Zig), static memory allocation, and compiler checked syscall error handling, the latter of which is really the number one thing by far that makes distributed databases unsafe according to the findings in "An Analysis of Production Failures in Distributed Data-Intensive Systems" [5].
While we could certainly benefit from the muscle of Rust's borrow checker in places, it makes less sense since TigerBeetle's design actively avoids the cost of multi-threading, with a single-threaded control plane for more efficient use of io_uring (zero-copy when moving memory in the hot path), plus static memory allocation and never freeing anything in the lifetime of the system. The new IO APIs like io_uring also encourage a future of single-threaded control planes (outsourcing to the kernel thread pool where threads are cheaper) since context switches are rapidly becoming relatively more expensive. Multi-threading for the sake of I/O is less of a necessary evil these days than it was say 5 years ago.
At some point, the benefits didn't outweigh the costs, and we had to weigh this up. In the end, it came down to simplicity, readability and state-of-the-art tooling.
[1] https://www.tigerbeetle.com
[2] https://github.com/coilhq/tigerbeetle/blob/main/src/vsr/repl...
[3] https://github.com/coilhq/tigerbeetle#simulation-tests
[4] https://web.cecs.pdx.edu/~kimchris/cs201/handouts/The%20Powe...
[5] https://www.usenix.org/system/files/conference/osdi14/osdi14...
> The Zig standard library is still unstable and mainly serves as a testbed for the language. After the Self-Hosted Compiler is completed, the language stabilized, and Package Manager completed, then it will be time to start working on stabilizing the standard library. Until then, experimentation and breakage without warning is allowed.
IIRC Forwards compatibility is not an issue because those deployments are one-and-done.
Sure. But people do all kinds of not advisable things, even if revenue is at risk.
In fact, it might even be fine, for their use cases. The build it once, it has the libs they want covered, it works like the want, that's it.
That can work with any early release language/lib.
The problem is with someone casually putting it into production, and then not expecting breakage, missing libs they will need, finding that there's not much tooling, and so on.
In other words "can it be put into production?" is another question compared to "is it ergonomic, stable, full featured enough to be a good and easy production choice?"...
Almost everything can be put into production (and even work reasonably well), even a 1000-liner Perl script written by someone who first tried Perl that same week.
Zig's C interop means it can leverage some of the most battle-tested libraries out there, with no shortage of libraries given C's massive ecosystem. And with Zig, it's not like you have to rewrite your whole system in a new language, you can incrementally rewrite the parts that make sense, test, and repeat.
Zig's tooling around compilation is also arguably ahead of most languages, and Zig's progress here is flowing back and adding value to many communities. For example, this 0.9.0 release can now build native Node addons without requiring node-gyp.
hehe
I've been looking around their web site and the only thing I could find was pretty generic:
Zig is a general-purpose programming language and toolchain for maintaining robust, optimal, and reusable software.
This page makes it sound like the self-hosted compiler will be the compiler, but various parts of the compiler infrastructure will be either interchangeable with not-quite-self-hosted modules via flags, or these LLVM-leveraging modules (e.g. for C/C++ compilation) will always be hard dependencies even in the self-hosted compiler.
The self-hosted LLVM backend is not done. Zig 0.9.0 is the self-hosted compiler, you are using the self-hosted compiler when you use Zig today. But it relies on the bootstrap compiler for the LLVM backend.
I know it is weird to say "self-hosted LLVM backend" but I don't know what else to call it. It's the .zig code that lowers Zig IR to LLVM IR.
It's possible to build Zig without an LLVM dependency by not passing `-Denable-llvm` to `zig build`. In this case the LLVM backend is not available (neither the self-hosted one or the bootstrap one).
My guess from the preliminary Plan 9 target work mentioned in the release notes is that something like a SuperH backend (for example) would indeed be in scope (provided someone's willing to contribute one, of course), but a confirmation would be neat. I suppose the C backend would do the job in these cases, too, but I'm sure it'd be nice to not have to include GCC (or some other C compiler besides zig cc) in the mix (especially for projects like OpenBSD that try to minimize copyleft code).
Speaking of: how is zig cc anticipated to work with a self-hosted Zig? Will there be a dependency on clang (as suggested by the punt_to_clang(...) call in the current main.zig)? Will it be possible to swap that out with something else that could turn C into ZIR (or something else the self-hosted compiler could then punt to whatever backend)?
Relatedly, would zig cc support the planned C backend? If so, would the resulting C output be equivalent to the input (notwithstanding the current limitations re: macros, struct bitfields, etc.)?
Yes! We won't block 1.0 on the quality of the less mainstream targets, but that's what the tier system is for - to ship a compiler that has varying levels of quality for various targets, while communicating clearly to users what kind of experience they can expect for each one.
SuperH patches are absolutely welcome.
> how is zig cc anticipated to work with a self-hosted Zig? Will there be a dependency on clang [...]?
The main distribution of Zig will be LLVM/Clang-enabled. However it is already possible to build a version of Zig that does not have these features enabled. In such case, compiling C, C++, and Objective-C code will result in an error.
However, the arocc project[1] is emerging, which, depending on a combination of how much funding ZSF gets and how much enthusiasm the unpaid contributors working in their spare time have, is looking like a promising C frontend that would be available even without LLVM/Clang. It is C only, however, with no intention of compiling C++ or Objective-C.
> would zig cc support the planned C backend?
As it is currently implemented: no. Zig invokes clang to turn C source code into object files.
However, with the arocc frontend mentioned above, this would be converting the C source code into ZIR (or perhaps AIR), which could then be lowered with any of the backends, including the (partially complete) C backend. In such case, the C output would look drastically different than the input. It would look more like a machine-generated IR than natural C code that a human would write.
https://ziglang.org/download/0.9.0/release-notes.html#Self-H...
Notice the bubble that says "LLVM Codegen" (44% done). This is the LLVM backend. All the other bubbles do not depend on LLVM at all.
I suggest to check back in with the next release of Zig and see where we are at. I suspect we will have at least the x86 backend fully operational by then.
[0] https://ziglang.org/download/0.9.0/release-notes.html#Compil...
Edit: To be clear, love enforcing the idea for production code, but wish they had embraced a '-dev' mode or equivalent flag that made it easier to experiment.
test "example" {
var x: i32 = 1234;
_ = x;
}Do discards not fulfill your use case?
Not a deal-breaker (honestly, I write Python most days), but a huge annoyance.
(From my use in Go, I'm definitely overall a fan of the feature. Tho of course how much personal value vs annoyance you get out of it is going to depend on your own code style and behaviors.)
_ = bla;
_ = blub;
...which have been forgotten during development.So the next thing that's needed is an error if 'bla' or 'blub' are actually used elsewhere ;)
Would catch this:
var a, b
c = a + a // whoops, should have been a + b, compiler complains about unused b
Doesn't catch this: var a, b
foo(a, b)
c = a + a // whoops, should have been a + b, but still compiles
All in all, unused variables being errors is an awful feature that isn't very helpful in practice, at the cost of making experimentation a pain in the arse.There is nothing that kills my state of flow more than having to comment a piece of code that is unreferenced, because the compiler complains, while I'm trying to hack and explore some idea.
Luckily though my Rust setup doesn't fail to compile with unused stuff, it just warns - and then on CI i have it reject all warnings.
I agree it's very frustrating not having a -dev flag or something less restrictive.
Unused variable warnings are a good idea, but making them hard errors is a language design mistake. Warnings are good because they allow programmers to quickly make changes and test things, while providing a reminder to clean things up in the end (which is why the "just use tooling that removes unused variables" response misses the point--when making quick temporary changes for debugging, you want the warnings as reminders to go back and undo the change). Additionally, warnings allow for adding helpful static analyses to the compiler over time without breaking existing code like introducing new errors does. As I recall, there were some cases in which Rust 1.0 accidentally didn't flag unused variables, which was fixable post 1.0 without breaking existing code precisely because it was a warning, not an error.
Does Rust emit warnings for cached compilation units?
There is an interesting thread with community consensus against the use of #[deny(warnings)] at [2]. The most important takeaway for me is that the right place to deny warnings is in CI. You don't want end users who compile your crate to be have their builds fail due to warnings, because they might be using a newer version of the compiler than you were. You don't want to fail developers' builds due to warnings while hacking on code, because of the overhead warnings-as-errors adds to the edit/compile/debug cycle. CI is the right place to deny warnings, because it prevents warnings from getting to the repository while avoiding the other downsides of warnings-as-errors.
[1]: https://rust-lang.github.io/rfcs/1193-cap-lints.html
[2]: https://www.reddit.com/r/rust/comments/f5xpib/psa_denywarnin...
I would agree that forcing uninitialised variables as errors is a design mistake.
foo := someDefaultValue
if someCondition {
foo := someMoreSpecificValue
}
Whoops, accidentally created a fresh, unused variable in the nested scope instead of changing the value of the original variable.I do that frighteningly often.
We found a few cases in TigerBeetle where some nasty bugs would have been detected and prevented by this. Some of these were at the entry point to a function, where the function failed to check all the pre-conditions and was ignoring some of the API interface, and others were where we were calculating some of the derived buffers we would need, but then never use them, using the original buffer incorrectly instead.
Here's to hoping Zig gets explicit SIMD intrinsics in 1.0. It is one of the few things holding me back. That and the pedantic compiler errors around unused variables.
1: https://github.com/ziglang/zig/issues/903#issuecomment-45950...
Anyway, please file feature requests for missing SIMD intrinsics!
https://www.reddit.com/r/javascript/comments/c8drjo/nobody_t...
zig fmt accepts and converts tabs to spaces, \r\n to \n, as well as many other transformations of non-canonical to canonical style.
Which is... still saying "spaces instead of tabs"In a nutshell, what you're currently using is the stage 1 compiler, aka the bootstrapping compiler to compile the official zig compiler going forward. zig fmt is opinionated because it was only meant to enforce formatting for the zig project. At that stage of the project's life, they felt it was more important to get shit done in a consistent manner w/ the people that are actually contributing than to cater to a hypothetical accessibility-impaired developer that isn't.
Zig is a very ambitious project. Prioritizing pragmatism over ideology is a fairly common theme with it currently. Another example: they repeatedly break stdlib APIs because catering to a larger audience is currently less important than getting other things nailed down first.
Right now you might be more worried about how the current compiler can't even compile Zig correctly before worrying about what stylistic choices it can handle. I'm a tabs guy myself but it's really quite irrelevant at the moment - the focus is still on figuring out how Zig should work and making that happen not day to day usability of the current toolsets for end users.
So yes, all of the quotes and conclusions you found are either old or wrong according to the very pages you linked. No, it doesn't need to be updated (or prioritized?) as they all conclude with supporting tabs you just hadn't read all of what you're citing.
.
If there is still any doubt for some reason rather than trolling through GitHub issues from 2017 I'd recommend sticking to the current FAQ from 2021 which has a section dedicated to this question: https://github.com/ziglang/zig/wiki/FAQ#why-does-zig-force-m... as this is the fully current and authoritative answer on the status of tab support, regardless of where bits of past discussion lay on the matter.