35 karma · joined March 21, 2023
Since we don't want to rewrite everything multiple times, it also has to be multi-platform and optimal, so the feature set must be per-device, not per-language. I'm not aware of a tool that does that, especially in Rust (which Burn is written in).
The goal was to improve performance and flexibility by using Tensor cores when available, performing bounds checks when necessary, supporting any tensor layout without any new allocation to transpose the matrices beforehand, and implementing many improvements.
The performance is greatly improved, and now it works better with many different matrix shapes. However, I think we created an atrocity in terms of compilation speed. Simply compiling a few matmul kernels, using incremental compilation, took close to 2 minutes.
So we fixed it! I took the time to write a blog post with our solutions, since I believe this can be useful to Rust developers in general, even if the techniques might not be applicable to your projects.
Feel free to ask any questions here, about the techniques, the process, the algorithms, CubeCL, whatever you want!
We've introduced a new tensor data format that offers faster serialization/deserialization and supports Quantization (currently in Beta). Loading and saving can be up to 4X as fast.
As always, we've added numerous bug fixes, new tensor operations, and improved documentation. Thanks to all contributors, over 50 for this release.
Why it Matters
CubeCL tackles three major challenges in GPU computing
- Portability: The same codebase can be used to program any GPU without a loss in performance.
- Usability: No need for a new shader language — simply add an attribute on top of your Rust code and voilà, it can now run on any GPU.
- Performance: We generate fine-grained kernel specialization via an innovative compile-time system to use the most efficient instructions available.
How it works
CubeCL leverages Rust's proc macro system in a unique two-step process:
1. Parsing: The proc macro parses the GPU kernel code using the syn crate.
2. Expansion: Instead of immediately generating an Intermediate Representation (IR), the macro generates a new Rust function.
The generated function, semantically similar to the original, is responsible for creating the IR when called. This approach differs from traditional compilers, which typically generate IR directly after parsing. Our method enables several key features:
- Comptime: CubeCL functions can contain sections marked as Comptime. These sections are executed during compilation rather than at runtime. This allows for the creation of highly specialized kernels by incorporating compile-time information directly into the generated code.
- Automatic Vectorization: By simply vectorizing the inputs of a CubeCL function, we can determine the vectorization factor of each intermediate variable during the expansion.
- Rust Integration: The generated code remains valid Rust code, allowing it to be bundled without any dependency on the specific runtime.
Our goal extends beyond providing an optimized compute language; we aim to develop an ecosystem of high-performance and scientific computing in Rust.For now we have highly optimized matrix multiplication kernels, leveraging Tensor Cores on NVIDIA's hardware when available. We are going to focus on adding more algorithms, but community contributions are more than welcome. There is still a lot of work to be done!
Don't hesitate to check the repo and ask any questions that come to mind.
We updated the API to remove instances where the device chosen was the default one, potentially causing bugs due to device mismatch. Now, our API is more explicit about where the device should be specified, aligning well with the Rust philosophy.
The book has seen various updates, including a new section on dataset manipulation requested by the community. We also plan to create a contributor guide, to help new contributors get familiar with the internals of the project.
A lot of work has been done to improve our JIT compiler, where we can fuse WebGPU tensor operations into a single kernel for impressive performance improvement. We added automatic vectorization of element-wise operations, as well as integration with autotune. Additionally, kernels created on-the-fly can now be executed in-place for reduced memory usage. We now support running multiple optimization streams independantly, which helps when metric updates and training run on the same device, but different threads. This feature isn't enabled by default yet, but you can enable it with a backend decorator. Future releases will add more optimizations to the compiler, and we will probably ship it by default. We also have plans to add other compilation targets in addition to WebGPU, namely Vulkan and CUDA.
One of the major quality of life improvements is the addition of the new PyTorch recorder that allows loading PyTorch weights into Burn modules. We also support specifying regex to dynamically map the weights to your Burn model if the structure isn't the same as the PyTorch implementation.
With this new release, we spent a lot of time solidifying our infrastructure, testing our framework on additional OS (Windows and MacOS). Overall, our CI is more mature and allows us to more easily ensure the quality and correctness of every code change across backends and operating systems.
Release Notes: https://github.com/tracel-ai/burn/releases/tag/v0.12.0 Burn Book: https://burn.dev/book/
In our most recent blog post, we explore two different operations commonly found in almost all deep learning models: Reduction and Matrix Multiplication (Matmul). The post highlights the difficulties of manually choosing the right kernel, given the uncertainties arising from the hardware and various input shapes and strides.
For instance, in some scenarios, our first reduction algorithm can be 3X as fast as our second one, but in other scenarios, it can be 19X slower, highlighting the importance of selecting the right kernel for the job. For Matmul, often we have the best kernel performing 3X as fast as the worse one, but the fastest one changes constantly across a spectrum of scenarios.
We hope the post highlights our flexible solution to this problem and how it can support our mission of creating the fastest framework on all hardware.
Let us know what you think of it and how we may improve it further.
Happy reading