ONNX runtime: Cross-platform accelerated machine learning
onnxruntime.ai
onnxruntime.ai
(disclaimer: I work at GH/MSFT, not connected to the Llama 2 project)
[0] https://onnx.ai
It's reliance on google's protobuf with it's 2gb single file limit is an extreme limitation. Yes you can keep weights outside your model file, but still many operations (model slicing) fail.
Second, inability to offload parts of the model to disc or cpu(like huggingface accelerate) while the rest executes on the gpu.
Thirdly, inability to partition existing large models easily. You can delete nodes, but then fixing the input/output formats means manually editing text files. The work flow is ridiculous (convert onnx to txt with pdoc, edit in text editor, convert back to binary).
I really wish they fix all this stuff and more.
is your complaint that the whole context/weights needs to be sent through the whole file?
I'm asking from a position of ignorance, I'm surpised to see a serialized transport of that size, and wondering why it's specifically limited at 2gb. like to be able to mmap on 32-bit hardware?
I'm not sure if this is a file size limit too or just an object memory representation size limit. For me using a library designed for message passing to save/read your AI models is a bad design decision.
>is your complaint that the whole context/weights needs to be sent through the whole file?
I store large onnx models with "external" weights, but even so many operations fail with the dreaded "ModelProto exceeds maximum protobuf size of 2GB: 3385275542". So the complaint is that you simply can't do a lot of stuff with models over 2gb.
You can create a session to execute the model, you can run the vanilla optimisation over it. But trying to run transformer specific optimisation errors out as well as making any attempts at slicing the model. Additionaly some conversion processes.
It appears the code that does that looks up the model size and just errors out if over 2GB.it doesn't even try loading it
Message. Each message cannot be bigger than 2GB(and usually should be not this big), but a file could contain multiple messages. This limitation helps them prevent integer overflows, since every length is 32-bit but every application process is 64-bit, you can convert the length numbers to 64-bit before doing any arithmetic operation. Therefore, fundamentally, there is no way to make it secure on 32-bit platforms, or no way to support more than 2GB on 64-bit platforms without totally rewriting the code.
With LoRA / QLoRA, my bet is that edge training capabilities are as important in the next decade. I don't have any citations though.
Is it? From what I understand, to use an analogy, ONNX is the bytecode specification and JVM whereas Pytorch, TF and other frameworks combined with converting tools are the Java compilers.
Your training framework and a suitable export is the compiler.
Onnx Runtime (which really has various backends), tensorrt, .. (whatever inference engine you are using) is your JVM.
I just did an install of the runtime on Python ( pip install onnxruntime ) . Here are the additional packages it installs.
Package Version
------------- -------
coloredlogs 15.0.1
flatbuffers 23.5.26
humanfriendly 10.0
mpmath 1.3.0
numpy 1.25.2
onnxruntime 1.15.1
packaging 23.1
protobuf 4.23.4
sympy 1.12
https://onnxruntime.ai/docs/install/If an ONNX implementation wants to do codegen, like what XLA does, then usually it is based on LLVM and it needs to be shipped with a copy of LLVM.
Eventually we went with pytorch only support for the time being, with still exploring OpenXLA in place of ONNX, as a universal adapter: https://github.com/ipcamit/colabfit-model-driver
Many users didnt want to install random binaries (security), and the devs didnt document or link directly to the corp websites.
Now its as easy as pip install, going to make things easier.
The community is moving faster that the corps making the tools.
I have written about it in my blog: https://www.zaynetro.com/post/run-ml-on-devices