OpenXLA Is Available Now
opensource.googleblog.com
opensource.googleblog.com
I want to see a DSL that can be used to describe models elegantly and then export them either to a shared object or to something that can be run with a runtime (in this case IREE). Things like ONNX and TorchScript promised this but I've had little luck getting these to work well enough to trust them in large scale production deployments.
I understand that PyTorch is an awesome tool for researchers, but it doesn't necessarily fit into a prod environment.
You need to write some infrastructure around PyTorch to make it work. Something like a key/mapping in each checkpoint that says which architecture to choose with which parameters.
It sure could be easier, but is saving the model's code into the checkpoint enough? Things like the data pre-processing expected by the model would also have to be included for it to really be self-contained.
Why should I look at Aesara’s representation of multi-dimensional array programs when I might already use JAX’s?
Does Aesara support a staging transformation that allows me to construct programs in your representation from a subset of Python?
I’m personally interested in the answers to these questions, given what I know about IREE, JAX, and XLA — as a user in the space, I haven’t been able to determine how Aesara would actually benefit me over JAX.
Note that I know that Aesara can use JAX as a backend — but I’m trying to ascertain what one extra layer buys me.
Admittedly we're on a reasonably easy situation: we just have to deploy models (some from scikit-learn, some from Keras, some from PyTorch) to various users who mainly run a specific version of python under Windows and Linux, with CPU and GPU support.
Some of the largest deployments of ML are using PyTorch models, e.g. OpenAI, Meta, Microsoft.
But for someone who wants to jump the bandwagon, does anyone have a "guide/map"? To put it simply, "How do I start AI/ML in 2023? And then what?"
The "2023" part is important. If you're bringing someone new to JS world, you probably show them Vue/React and not jQuery/Prototype.
When you see a new compiler/runtime, it usually impacts deployment and optimization post training (TensorRT, FasterTransformers, TVM, OpenXLA) and/or target new specialized silicon (AWS Inferentia and Trainium, Google's TPUs, and others).
So to answer succinctly:
> How do I start AI/ML in 2023? And then what?
You implement your model in PyTorch, you train on Nvidia A100 and you deploy with the framework that gives you the best speedup for your architecture.
if your just trying to learn the "basics" I would argue use google colab or something similar to get your feet wet. it's also a good way to learn python
Also you might look at huggingface or model zoo for existing models. It feels alot like early does of public apis. Where you can do alot of cool things with apis and mashups.
Edit: I would personally suggest reviewing the basics: linear algebra, optimization, then the classical ML technique. Then move on to deep learning. Look at important research papers and GitHub code, and implement models to really understand how they work.
So why does there seem to be no published metrics showing performance of various common ML models on common hardware with OpenXLA vs other frameworks/compilers?
But people like Google offering ML compute as a service want to keep the actual optimization modules secret and closed source. That way they can make their hardware perform superspeed while competitors hardware looks slow.
For simple stuff, we can compare JAX to PyTorch on a 4090, and JAX seems faster by 10-50%. It's way way way faster on TPU.
That said, JAX is apparently a pain to work with in comparison to PyTorch (disclaimer I haven't used JAX enough to opine yet).
I'd also want a comparison between JAX and Taichi-lang for parallel stuff, even though the problem domains aren't exactly aligned.
What benchmarks are you looking at here?
Not super scientific, Im sorry to say, but methodologically benchmarking this is really hard.
The point made above is that WebGPU can only be used for GPU's and not really for other types of 'neural accelerators' (like e.g. the ANE on Apple devices).
Accelerators inside GPU (like Tensorcores) seems like a lot better deal as you can easy utilize it without 4 abstraction layers with only some unknown to us mortals operations support inside. (And my god i hope apple will allow to programmable run ANE or at least put this api inside Metal framework cause right now working with Coreml for anything new is a nightmare and even some old models are broken on new versions of coremltools)
Also worth pointing out IEEE intends to target wasm, vulkan/spir-v (webgpu's wgsl isn't entirely unrelated but would take work). So if you don't have ml on your system you still have good targetting options. If the web platform gets support, the browser could internally target vulkan as a baseline.
I'm curious how different the training vs inference needs are, and whether these tools can adequately serve both.
In IREE, we have prototypes targeting Wasm and WebGPU with ahead-of-time compilation, and we'd like to see more hardware exposed as compute devices via Vulkan/WebGPU (possibly leveraging extensions for computations like matrix multiplication).
Also, OpenXLA is one of the external organizations in the Coordination section of the working group charter: https://w3c.github.io/machine-learning-charter/charter.html. We're looking forward to collaborating with WebML folks!
TensorRT is proprietary to Nvidia and Nvidia hardware. You'd take a {PyTorch, Tensorflow, <insert some other ML framework>} model and "export / convert" it into essentially a binary. Assuming all goes well (and in practice rarely does, at least on first try - more on this later), you now automatically leverage other Nvidia card features such as Tensor cores and can serve a model that runs significantly faster.
The problem is TensorRT being exclusive to Nvidia. The APIs for doing more advanced ML techniques like deep learning optimization requires significant lock-in to their APIs, if they are even available in the first place. And all these assuming they work as documented.
OpenXLA (and other players in the ecosystem like TVM) aim to "democratize" this so there are more support both upstream (# of supported ML frameworks) and downstream (# of hardware accelerators other than Nvidia). It's yet another layer or two that ML compiler engineers will need to stitch together, but once implemented, they in theory can do a lot of optimization techniques largely independent of the hardware targets underneath.
Note that further down in the article they mention other compiler frameworks like MLIR. You can then hypothetically lower (compiler terminology) it to a TensorRT MLIR dialect that then in turn runs on the Nvidia GPU.
>You can then hypothetically lower (compiler terminology) it to a TensorRT MLIR dialect that then in turn runs on the Nvidia GPU.
there's no tensorrt dialect (there are nvgpu and nvvm dialects) nor would there be as tensorrt is primarily a runtime (although arguably dialects like omp and spirv basically model runtime calls).
Lots of innovation here. Time for a proper DL compiler.
Much of this code is not "new" in the sense that much of the OpenXLA effort has been extracting the existing XLA representations and compiler from the TensorFlow codebase so it can be more modularly used by the ecosystem (including PyTorch).
A better frame is TensorFlow exporting its stable representation that many vendors have already built around, more than a "new" standard.
I'd also like to comment on our (StableHLO's) relationship with related work. StableHLO was a natural choice for the OpenXLA project, because a very similar operation set called HLO powers many of its key components. However, I would also like to give a shout out to related opsets in the ML community, including MIL, ONNX, TFLite, TOSA and WebNN.
Bootstrapping from HLO made a lot of sense to get things going, but that's just a starting point. There are many great ideas out there, and we're looking to evolve StableHLO beyond its roots. For example, we want to provide functionality to represent dynamism, quantization and sparsity, and there's so much to learn from related work.
We'd love to collaborate, and from the StableHLO side we can offer production-grade lowerings from TensorFlow, JAX and PyTorch, as well as compatibility with OpenXLA. Some of these connections in the ML ecosystem have already started growing organically, and we're super excited about that.
Definitely feels like critical mass
While ONNX Runtime seems to split efforts away from .NET ML.
It is an area where they seem to suffer from the same management issues like in desktop GUI frameworks of lately.
(I work for Google and I work on client-side StableHLO, but I don't speak for Google).
Also, as someone interested in MLIR - I’m excited that (perhaps sometime in the future), I’ll be able to read the op semantics outside of the TensorFlow docs :)
Is this going to be addressed?
Can a large tensor be split into several small ones?
In the StableHLO spec, we are talking about this in more abstract terms - "StableHLO opset" - to be able to unambiguously reason about the semantics of StableHLO programs. However, in practice the StableHLO dialect is the primary implementation of the opset at the moment.
I wrote "primary implementation" because e.g. there is also ongoing work on adding StableHLO support to the TFLite flatbuffer schema: https://github.com/tensorflow/tensorflow/blob/master/tensorf.... Having an abstract notion of the StableHLO opset enables us to have a source of truth that all the implementations correspond to.
> Extension mechanisms such as Custom-call enable users to write deep learning primitives with CUDA, HIP, SYCL, Triton and other kernel languages so they can take full advantage of hardware features.