Using Zig to Unit Test a C Application
mtlynch.io
mtlynch.io
This is for example the case in libaegis: https://github.com/jedisct1/libaegis
Calling C functions from Zig is easy (C headers can be imported directly) and doesn't have any overhead. So I can take advantage of the convenience of Zig, even if the tested code only requires a C compiler.
I also now always add Zig build files as an alternative to make/libtool/automake/cmake/meson/etc. The main advantage is that cross-compilation to many targets is supported out of the box, including to WebAssembly. So I can quickly test if the C code compiles fine before actually trying to run it on an emulator.
Thanks for sharing!
For the readers who aren't familiar:
The code for running test and then cross-compiling (on one machine and OS) for different target platforms is:
https://github.com/jedisct1/libaegis/blob/main/.github/workf...
and the only zig file in the repo which drives the build process is:
I see there's "translate-c" to help migrate C to Zig, and maybe that's enough.
Generally I fall on the side of if a compiler can toss out a type warning, it could certainly give a reasonable solution also. Best-guess with a "// Warning type was guessed" is better than an opaque "RTFM" in most every case.
Do you mean building a .h file from zig code?
The other direction is just @cImport.
I wouldn’t so much pin this on zig as much as the complete clusterfuck that is building c libraries.
The D compiler(s), as Zig, can be used to compile C code... and D code can import C files as if they were D modules (as in Zig, there are some special types to represent C strings and other types that are not exactly the same).
The approach I am more familiar with is using Google test and C++ to test C code. It's pretty easy if you already have a cmake project set up, and most C developers can wrap their heads around GTest.
However, for the complete newcomer to Zig, this post seems a bit convoluted in this presentation.
What about the Zig authors do a step by step tutorial on how to leverage Zig to implement tests in a C codebase?
Would be a great entry point for the broader community!
many system languages are making the headlines lately its very hard to pick one to learn
not sure how to deal with this, learn them all, bet on one, what should we do
- it is used to create an OS https://github.com/mirage
- it is used to create a transpiler https://melange.re/v2.2.0/
- it was used to create Rust first compiler
ocaml is surely a systems languageWriting a OS is more "system", and certainly those using the Mirage operating system use OCaml as a system language.
A colloquial definition of systems language seems close to: "exposes low level details and doesn't have garbage collection." By this definition c, c++, zig and rust are system languages, java and ocaml are not. Go is debatable (and people do debate this).
Personally though, I prefer to think of what types of _systems_ I can build with a language. I can design a framework in a high level language which then transpiles to, say, c. I consider this high level language to be a systems language because it facilitates a system (the framework). Others will disagree but it comes down essentially to a semantic question over what counts as a system language. That debate doesn't seem especially fruitful to me.
Rust is created in the same spirit that created and evolved C++: create a complex and featureful language that enables compiling your solution from a high level representation in an expressive/safe/performant way.
Zig is created in the same spirit that created and evolved C: create a simple language that allows you to directly and transparently represent and reason about what you want to have happen.
You'll probably think that one of these statements is more biased than the other, and that probably reflects your own preferences :)
Unlike C++, ISO C is nothing more than culmination of features that more than 1 compiler has implemented (and doesn't interrupt the compilation process of a micro-controller firmware that was released literally 40+ years ago). Anything else, is GNU C. And it is so incredibly complex and obtuse at times that clang still can't compile glibc after years of work.
Zig was not created with the same spirit that created and evolved C. Zig was created with the idea of a simple C, one that does not match reality, and frankly leans more on Go rather than C. Zig, Odin, V, nearly all these better-C languages are more inspired by Go itself, than what C actually is. What they want from C is just the performance; that's why they're so focused on manual memory management one way or another.
If you squint zig's error return fusion looks a bit like go's tuple error return but it actually is more "first-classing certain c conventions" than "adopting a go pattern". Same goes for slices.
C never had the philosophy of keeping things simple through the years. If it did, we would not have time traveling UBs to begin with. The lauded simplicity and explicitness comes directly from Go, where the philosophy was crystalized and preserved very early on.
You might say it is semantics, to call improving upon C being a derivative of Go (with manual memory management). You would be partially correct, it is semantics, but one that holds up very well if you look at how languages developed over the decades.
I have two words for you: json marshalling
1. It only allows function calls instead of any expression.
2. It allocates memory dynamically and attaches the function call expression to the function, rather than the current scope exit. This has surprising and harmful consequences if you use it inside a loop.
So, I wouldn't say that zig's defer is borrowed from Go.
Edit: ok rewatched it and I didn't see that come up, i was just misremembering
Zig - attempts to stay simple, like C, but with warts fixed and with cool compile-time programming. Its biggest strength seems to be not the language itself, but the compiler and build system which can cross-compile seamlessly, including C code.
Nim - a systems-language that looks like Python and tries to be fun to write. Has macros that may remind you of Lisp macros. Compiles to C or JS.
D - older but very cool as well... I was surprised to find out its metaprogramming capabilities are as good as Zig's or Nim, and that is has a lot of cool features not seen in mainstream languages, like contract programming and executable documentation. Much more mature than the previous ones. Also seamlessly compiles and imports C.
Odin - really reminds me of Go. It's used in production to create fluid simulation for Holywood movies apparently. Very minimalistic language but I couldn't see what it brings to the table that the ones above do not. It's kind of similar also to Jai which is also upcoming but focusing on game programming from what I understand... that's still not even publicly available yet.
Which one to choose really depends on your taste, hope my descriptions above help, even if they're pretty rough simplifications.
If you want the most popular language in this area, that's undoubtedly Rust though.
If you don't know it already, you learn C, as well as you can. It is not a hard language to learn, but it has a lot of footguns. That's why you learn it along with tools like Valgrind & sanitizers.
Then you look at Rust. That's somewhat harder to learn, but nails some important details you need to think about while writing C. It will make some things obvious that you'd need to learn by shooting yourself in the foot repeatedly in C. You don't need to use Rust once you got what you need out of it education-wise, but a lot of people like and use it.
At this point it is kind of unimportant what other "systems" language you decide to learn, but here's my opinion of some of them:
- Personally I like Odin's ergonomics. It is incredibly convenient. You can just jump in and start writing OpenGL code without dealing with wrappers and all that. Included vendor libraries take care of a lot.
- I also like the explicitness of Zig. It seems like it'll be the most popular one in the future, most likely not because of the language itself but because of the tooling. By the way the reason I say "not because of the language" is that the maintainers seem uninterested in having some way to constrain generics. The language sorely needs some sort of comptime interface / traits / concepts, anytype-everything is not nice. In 10 years someone will come up with a Boost-like library that implements just that in userspace and it'll be horrible.
- Ocaml is garbage collected Rust, or rather, Rust is non-garbage-collected Ocaml. It is underrated. Jane Street people are adding borrow checker to it. Could be more popular in the future. Also, all languages are "systems" languages depending on how you wield them. No need to bikeshed about Ocaml's "system"ness status.
- D is pretty cool. Very fragmented library ecosystem if you want to do betterC or no-gc though. Be prepared to just use C or C++ libraries, which it can talk to pretty easily.
- Nim is a nice language. Does reference counting so it has a low memory footprint compared to other GC'd languages. It compiles to C so technically it is the most portable language in this list. You can easily run it on microprocessors that others won't run on. Try it, you'll very quickly land on "like it" / "don't like it" territory depending on your programming style.
- Jai is non-existent right now. Doesn't warrant a discussion until Jon Blow feels it is ready for prime time. But since that's his strategy, expect something practical and polished. If it sucks, two possibilities: 1) he didn't deliver and it won't get drastically better or 2) your use case was not in consideration.
- C3, it exists, it is usable, it is like a halfway between C and D. I didn't spend much time on it yet.
- Free Pascal: I didn't use it but just putting it here because this list is getting long & it kind of deserves a shout. Lazarus looks nice.
- Go: Use Java or C# instead, they can compile to native now.
- C++: It exists, it is used everywhere, it sucks. As opposed to most other languages on this list, it wasn't designed. It kind of picked up random features along the way because they looked good. You kind of design it by picking up a subset and putting up with its weirdnesses. Don't use it if you can help it. If you have to use it you most likely didn't have a choice in the first place.
I'm not sure I agree. Anything which attempts to constrain "anytype" makes the language much, much more complicated.
For example, if it were constrained, Zig "allocgate" would likely have necessitated compiler changes instead of just library changes.
And I often think that we conflate two different things--"Generic" and "Libraries like the big boys build".
I generally don't need fully generic programming.
What I do need is the ability to build a library that works exactly like the standard library. All libraries need to be precisely equal in expressive power and composability to those that have been officially "blessed".
Oddly, Rust fails at this due to things like the Orphan Rule even though its "genericity" is quite expansive. The invasiveness of Serde is a prime example. If an external crate doesn't support Serde, you can't add it. You have to literally copy the entire library over to your code in order to add Serde support. This blocks an alternative to Serde from ever arising because it will never be as convenient as Serde which has been "blessed" by the community.
Technically speaking, Java and C# would serve somewhat different purposes and I would rather recommend Kotlin and C# with the former serving higher-level code goals better with existential types, SRTPs and overall strong type system and the latter for lower-level and/or performance-sensitive code with SIMD, pointers/byrefs, struct generics (zero-cost abstractions) and free/cheap interop.
The only concern is how well Kotlin Native works today regarding compatibility. NativeAOT has seen a lot of work to improve this and scenarios that will never be supported (runtime reflection emit which needs JIT or arbitrary unbound reflection) are now well-documented.
You may also be interested in Bflat[0] which has 'UEFI' as a target or even Zerosharp[1] as a demonstration how far you can push this (which is, of course, impractical, just use Rust :D)
[0] https://github.com/bflattened/bflat
[1] https://github.com/MichalStrehovsky/zerosharp/tree/master/no...
- Typescript for anything related to webdev, not a huge fan of the language but can't deny the ecosystem for productivity
- Rust for systems-level projects
- Crystal for quick scripts, although I'm still on the fence here
$ pwd
/home/kaz/ustreamer
$ gcc -D_GNU_SOURCE -shared src/libs/base64.c -o base64.o
$ valgrind txr --free-all
==8785== Memcheck, a memory error detector
==8785== Copyright (C) 2002-2017, and GNU GPL'd, by Julian Seward et al.
==8785== Using Valgrind-3.13.0 and LibVEX; rerun with -h for copyright info
==8785== Command: txr --free-all
==8785==
This is the TXR Lisp interactive listener of TXR 292.
Quit with :quit or Ctrl-D on an empty line. Ctrl-X ? for cheatsheet.
Psst! The complimentary Allen key that comes with TXR is inpired by IKEA.
1> (with-dyn-lib "./base64.o"
(deffi us-base64-encode "us_base64_encode"
void (buf size-t (ptr (array 1 str-d)) (ptr (array 1 size-t)))))
us-base64-encode
2> (let ((out (vec nil))
(sz (vec 0)))
(us-base64-encode #b'00112233445566778899AABBCCDDEEFF' 16 out sz)
(list out sz))
(#("ABEiM0RVZneImaq7zN3u/w==") #(25))
3> (let ((out (vec "abcde"))
(sz (vec 5)))
(us-base64-encode #b'00112233445566778899AABBCCDDEEFF' 16 out sz)
(list out sz))
(#("ABEiM0RVZneImaq7zN3u/w==") #(25))
4> (let ((out (vec "xxxxxxxxxxxxxxxxxxxxxxxxxx"))
(sz (vec 25)))
(us-base64-encode #b'00112233445566778899AABBCCDDEEFF' 16 out sz)
(list out sz))
(#("ABEiM0RVZneImaq7zN3u/w==") #(25))
5> (let ((out (vec "xxxxxxxxxxxxxxxxxxxxxxxxxxxx"))
(sz (vec 27)))
(us-base64-encode #b'00112233445566778899AABBCCDDEEFF' 16 out sz)
(list out sz))
(#("ABEiM0RVZneImaq7zN3u/w==") #(27))
6>
==8785==
==8785== HEAP SUMMARY:
==8785== in use at exit: 0 bytes in 0 blocks
==8785== total heap usage: 24,782 allocs, 24,782 frees, 4,806,317 bytes allocated
==8785==
==8785== All heap blocks were freed -- no leaks are possible
==8785==
==8785== For counts of detected and suppressed errors, rerun with: -v
==8785== ERROR SUMMARY: 0 errors from 0 contexts (suppressed: 0 from 0)
We covered the cases when the the destination buffer is already allocated, and smaller, equal to or in excess of the space being required. Thus testing the cases when realloc is or isn't necessary. No memory errors or leaks.We can see in the last case that when the buffer is larger, the function leaves the size alone, returning our original 27.