OneAPI – A cross-industry, open, standards-based unified programming model
oneapi.com
oneapi.com
I'm struggling to see what's enticing in this for python people that already have largely optimised torch and tf cpu and gpu backends, especially for batch work. And for latency-sensitive inference, I thought the 'industry' was going for TVM or other 'target all the things at code level'.
I'm thinking: gimme your network in onnx format, everyone gives a C inference/compilation API, and let everyone optimize /behind/ that... Xilinx, AMD, whoever-Google-bought-last-week...
Some incarnation of oneAPI is bound to exist as long as Intel has a foothold in the HPC market. For example, MKL and MKL-DNN have been rebranded with a oneAPI prefix. So no, the long term stability argument doesn't hold water.
I'd also note that they appear to have added support for cuBLAS and cuDNN as backends in their respective oneAPI libraries. It would be hilarious if that lead to more people running oneAPI on non-Intel hardware than first party stuff.
How many real alternatives are there though aside from OpenCL? Not to mention if you're interested in targeting something that isn't an Nvidia GPU.
Better in what sense?
> they had async compute way longer
What is that?
May be some other model can be designed that's nicer. CUDA isn't really using shader paradigm, right?
AMD’s compute cards have generally been worse than Nvidia’s as far as I know.
See: http://developer.amd.com/wordpress/media/2012/10/Asynchronou...
It's awful. And it is not even supported on RDNA and RDNA2. They even dropped Polaris support last month so that it's only supported on Vega.
And that's very much not a high quality GPU compute stack, sorry...
CUDA isn't a dead end and isn't going away any time soon.
CUDA isn't dead, but there should be a bigger effort to get rid of it because it's not a proper GPU programming tool but rather a market manipulation tool by Nvidia.
RDNA2 with its cache and low main memory bandwidth didn't help either, because while the cache is there, you're going to exceed what it can absorb by a lot in compute workloads...
As of no difference in compute workloads, I can put Blender as an example: https://techgage.com/article/blender-2-91-best-cpus-gpus-for...
(6900 XT is 20.6 TFLOPS FP32, RTX 3090 is 35.6 TFLOPS FP32 _and_ also has the tensor cores, AMD can't win there because of sheer brute force)
And AMD's software stack removes them from consideration for a huge extent of compute workloads. If they aren't ready to have a proper one, it doesn't matter how good how much their hardware is good or not, it might as well be a brick.
Either way, CUDA should be gone for good and replaced with something that's not tied to specific GPU. I welcome competition in hardware. I have no respect for lock-in shenanigans.
I wish CUDA wasn’t vendor locked, too, but AMD and Intel need to start taking GPGPU software seriously. They’ve dug themselves a massive hole at this point.
Also, I the case of "asynchronous shaders" it's a command list dispatching feature, which is equally applicable to CUDA/ROCm and compute shaders.
Edit: Maybe I should be more precise with my language. Using CUDA streams, you can queue up and execute multiple kernels (basically commands?) in parallel. Kernels invoked on the same stream are serialized. I have no idea what the exact analog is in the compute shader world.
Nvidia hardware won't dispatch multiple kernels simultaneously, all of the streams stuff you're seeing is hints to the user space driver in how it orders the work in it's single hardware command queue.
I thought this allowed kernels to run in parallel if both kernels could fit in the open resources. It doesn't have a priority system, and I'm guessing isn't smart enough to reorder kernels to fit together better. Further Googling to be sure that they mean parallel AND concurrent didn't turn up much, but Nvidia does appear to mean both.
The GPU is able to run up to 128 kernels in parallel on modern NVIDIA GPUs, concurrently.
> Also, I the case of "asynchronous shaders" it's a command list dispatching feature, which is equally applicable to CUDA/ROCm and compute shaders.
Async compute affects how you have to synchronize your code, and what your hazards are. If you have a queue of items, you know when one starts that the other is done. With async compute, you can give up that guarantee in exchange for overlapping work. NVIDIA still doesn't have this in modern GPUs.
See this slide deck about it: https://developer.download.nvidia.com/CUDA/training/StreamsA...
(for the implementation in Fermi, it's even more flexible since then)
I found ISPC (https://ispc.github.io) to be a little easier to work with than CUDA, for SIMD on CPUs, as it sits a bit in the middle of the CUDA and regular C models. Interop can work via C ABI i.e. trivial, it generates regular object files, and so on.
Univ. of Tennessee, Knoxville
Univ. of California, Berkeley
Univ. of Colorado, Denver
Date
October 2020
The goal of the MAGMA project is to create a new generation of linear algebra libraries that achieves the fastest possible time to an accurate solution on heterogeneous architectures, starting with current multicore + multi-GPU systems. To address the complex challenges stemming from these systems' heterogeneity, massive parallelism, and the gap between compute speed and CPU-GPU communication speed, MAGMA's research is based on the idea that optimal software solutions will themselves have to hybridize, combining the strengths of different algorithms within a single framework. Building on this idea, the goal is to design linear algebra algorithms and frameworks for hybrid multicore and multi-GPU systems that can enable applications to fully exploit the power that each of the hybrid components offers.Designed to be similar to LAPACK in functionality, data storage, and interface, the MAGMA library allows scientists to easily port their existing software components from LAPACK to MAGMA, to take advantage of the new hybrid architectures. MAGMA users do not have to know CUDA in order to use the library.
There are two types of LAPACK-style interfaces. The first one, referred to as the CPU interface, takes the input and produces the result in the CPU's memory. The second, referred to as the GPU interface, takes the input and produces the result in the GPU's memory. In both cases, a hybrid CPU/GPU algorithm is used. Also included is MAGMA BLAS, a complementary to CUBLAS routines.
oneAPI might be more general but at the end of the day heterogeneous HPC converges on linear algebra. So it might have different ergonomics but it's very similar.
The right way to go would be to first upstream, finish and optimize AMD and Apple M1 support to PyTorch and Tensorflow and XLA, while trying to abstract as much to common libraries as possible.
Trying to create an API first just slows development down even more.
Intel's effort seem to be banking on the Aroura system at ANL. Beyond that I don't know who's lining up to by Xe GPUs? Though I guess oneAPI will fit into whatever they do further down the line.
And it seems to have decent Python bindings.
What is the value proposal here?
What oneAPI (the runtime), and also AMD's ROCm (specifically the ROCR runtime), do that is new is that they enable packages like oneAPI.jl [1] and AMDGPU.jl [2] to exist (both Julia packages), without having to go through OpenCL or C++ transpilation (which we've tried out before, and it's quite painful). This is a great thing, because now users of an entirely different language can still utilize their GPUs effectively and with near-optimal performance (optimal w.r.t what the device can reasonably attain).
[1] https://github.com/JuliaGPU/oneAPI.jl [2] https://github.com/JuliaGPU/AMDGPU.jl
So for years, before they started having a beating OpenCL was all about a C99 dialect with printf debugging.
SPIR and support for C++ came later as they were already taking a beating, trying to get up.
Apparently that is also the reason why Apple gave up on OpenCL, disagreements on how it should be going, after they gave 1.0 to Khronos.
Just compare Metal, a OOP API for GPUs, with Objective-C/Swift bindings, using a C++ dialect as shading language, a framework for data management, with Vulkan/OpenGL/OpenCL.
opencl 2 went into the weeds going full c++, but that all got rolled back with version 3.