ONNX Runtime merges WebGPU backend
github.com
github.com
ONNX support in the browser was lacking and limited to CPU, but with a WebGPU backend it may now finally be feasible to run models in the browser on a GPU, which opens up interesting oppertunities. Although from this PR it looks like only a few operations are implemented, no browser-based GPT yet.
>Some context for those who aren't in the loop: ONNX Runtime (https://onnxruntime.ai/) is a standardization format for AI models.
It's just an IR, one of many - every framework has its own.
>Nowadays, it's extremely easy to export models in the ONNX format, especially language models with tools like Hugging Face transformers which have special workflows for it.
Meh it's poorly supported by both PyTorch and TF. Why support Microsoft's IR when you have your own.
>probably the most performant ML runtime at this point.
Not even by a long-shot - first party compilers are generally faster because of smoother interop but even amongst third-party you have TRT and TVM. TBH I have no idea what anyone uses ONNX for these days (legacy?).
[1] https://github.com/apache/tvm/tree/unity
[2] https://github.com/openxla/iree
[3] https://pytorch.org/get-started/pytorch-2.0/#inference-and-e...
It's surprising to me how many people just shoot from the hip/guess on this stuff. IKYK and if you don't maybe don't guess?
This does not match my experience. Most new model architectures port fine to ONNX from pytorch, only occasionally having to fill in rare functions.
> Not even by a long-shot - first party compilers are generally faster because of smoother interop but even amongst third-party you have TRT and TVM.
Again, in my experience, there is little by way of documentation for any of these platforms nor as broad support for as many OS and Hardware combinations.
> I have no idea what anyone uses ONNX for these days
Huggingface has strong support for ONNX and leverages it for improved performance in places.
https://huggingface.co/docs/optimum/onnxruntime/usage_guides...
https://huggingface.co/docs/optimum/onnxruntime/usage_guides...
It's curious how impressions can differ so.
What I'd love to see is a lean, inference-only version of pytorch that just works with existing code, and is smaller without the overhead of autograd and GPU support.
Tensorflow has breaking changes in model behaviour on patch versions!
Pytorch and tensorflow require a huge python environment to run, and how on earth do we get that onto a client system without them having virtualisation?
Worst of all, they both have very significant changes in prediction depending on the CPU (or god forbid GPU) on the system.
Once we export to Onnx we find we get reliable output and performance across runtimes, which seems mandatory for running any kind of product.
Every time I've convinced data scientists to just hand off an ONNX file to me for Production everyone comes out pleased: building an ONNX file is easier than productionizing a notebook any other way; the speed and performance of ONNX runtime are great and it is more easily integrated in C#-based pipelines (potentially avoiding expensive data transfers to/from Python VMs).
The biggest, ugliest hurdle I've seen is how many pre/post-processing steps data scientists tend to accidental convince themselves "can only be done in Python" either because they don't have time to research alternatives, believe Python to be inherently magical, or found some massive, obscure Huggingface-like corpus with gigabytes of data that would result in the most bloated ONNX files and "obscured in a Python or VM install step" hides how big the corpus is.
ONNX Runtime enable also provider specific inference code like TensorRT, CUDA, CoreML etc... without having you to change your code
what else needs to be implemented? will it be possible with the WebGPU API/spec?
I might do something like this if I was in a rush and just needed a save point, especially if I was working on a branch I knew I'd squash and merge with a more meaningful commit message. I certainly wouldn't be proud of it, but I might do it.
It has the energy of someone who has been forced into version control and deeply resents it.
When you do large PRs you want to checkpoint your work and commits have less meaningful descriptions, it's just next step of steps that you couldn't plan ahead, they're popping up as you go.
My approach is to combine commit messages when squashing. With that approach they are useful.
Completely valid commit strategy if you ask me (even though I bet some might beg to differ).
[1] https://xenova.github.io/transformers.js/
[2] https://twitter.com/xenovacom/status/1650634015060156420
It still took a lot of effort but the final version is very performant and reliable.