MLC's Apache TVM implementation is also excellent. The autotuning in particular is like black magic.
Speed can be improved. Quick and dirty/hype solutions, not sure.
I really hope ONNX gets traction it deserves.
Speed (and ergo costs) trumps elegancy in industrial settings imho, specially considering how expensive to run LLMs are. Such an OSS dependency can be improved or at least "wrapped" to isolate it from the rest of your codebase.
Even if llama.cpp was ~30% slower and the same size as ONNX, something like a custom grammar implementation or an extended context would be a huge deal.
Curious what you mean by this
There's only going to be more of these models and ONNX starts from the right place, cross-platform and from base principles rather than coupling tightly to one model structure.
Most importantly, it is freakin' awesome, the comments thus far, 30 in, don't reflect what its like to use or it's technical realities.*
* the main threads of discussion are "not even wrong": float16 is big compared to float4 (its trivial to quantize to your liking) and looking for an alternative that supports CoreML (ONNX is the magic that makes your model take advantage of CoreML / WebGPU / WebGL / whatever Android's marketing name for its API is etc. etc. etc.)
Funny how it's the opposite.
ONNX in this case, outside of the HN headline and saying "we did it" is almost useless.
LLMs are so heavy that you can't afford running a suboptimized version. This FP16 ONNX takes 4x as much memory and is probably 5-10x slower than something hand optimized such as llama.cpp or exllama with 4 bit quants.
I think ONNX would need to natively support this "packed" format or otherwise quantize fp16 models from disk on the fly... Which is a problem, as the FP16 models are huge.
gpt-q is also more than just group wise quantization and takes a decent while. Quantizing models without a major performance hit is not possible on the fly.
It can be cheaper to deploy Llama.cpp and then foobar.cpp (6 months from now) than it is to have inference that is 2x slower.
Interestingly in the LLM space all these model servers seem to be converging to using the same API as OpenAI, making it easy to swap containers to get a different model+inference server with 0 code change.
> its trivial to quantize to your liking
...Except its not implemented in the demo. Also, quantization is far from simple.
All this sounds good, but I have seen cool sounding ONNX demos for years, and (outside of some fairly quick one off TensorRT demos) I havent really seen the pavement hit the road.
What do you mean by this? The demo UI? Code quality?
There are currently also more quantization options available as mentioned. Though those incur a performance loss (they make the model faster but worse) so it depends on what you're optimizing for.
> specific instruction sets like for apple silicon
You are thinking of the Accelerate framework support, which is basically Apple's ARM CPU SIMD library.
But Llama.cpp also has a Metal GPU backend, which is the defacto backend for Apple devices now.