ZML - High performance AI inference stack
github.com
github.com
And ZML the framework also resolve issues with the complex dependency graph of stablehlo/pjrt.
That being said, I have a question regarding the ease of use. How difficult it is for someone with python/c++ background to get used to zig and (re)write a model to use with zml?
Coming from Python, the hardest part is learning memory management. What helps with ZML is that the model code is mostly meta programming, so we can be a bit flexible there.
We have a high level API, that should feel familiar to Pytorch user (as myself), but improves in a few ways
Note that our focus is being platform agnostic, easy to deploy/integrate, good performance all-around, and ease of tweaking. We are using the same compiler than Jax, so our performances are on par. But generally we believe we can gain on overall "tok/s/$" by having shorter startup time, choosing the most efficient hardware available, and easily implementing new tricks like multi-token prediction.
my only question is: is zig stable enough to base such a project on?
Stable as in reliable enough, I’d say so.
We are also looking ahead at Zig roadmap, and trying to anticipate upcoming breaking changes, and isolate our users from that.