I wonder why Microsoft helps nVidia, instead of using their own technology?
Here’s an example: https://github.com/Const-me/Cgml
I wonder why Microsoft helps nVidia, instead of using their own technology?
Here’s an example: https://github.com/Const-me/Cgml
https://blogs.windows.com/windowsdeveloper/2023/11/15/elevat...
https://blogs.windows.com/windowsdeveloper/2023/12/14/direct...
So probably, the group wasn't even aware of that technology because it's far away by either professional connection, or by personal/relationship connections, or they knew about it but had another goal than "maximize use of own stuff" and made the call that the tradeoffs wasn't worth it.
I think the only reason for that monopoly is the mental inertia of everyone involved. The complexity of these AI models is contained within the data in the models (gigabytes of numbers in these tensors), the GPU-running code is rather simple, most of that code is basic BLAS stuff. Unlike traditional GPGPU applications (FEM, numerical simulations, fluid dynamics), the compute kernels used in AI are easily portable across GPU APIs.
Maybe when I have some free time, I should port my library to Linux + Vulkan, just to prove the point.
In my experience, using any non-native GPU API is asking for troubles. On Windows, the native ones are D3D 11 and 12, on Linux and Android it’s often Vulkan, and on MacOS it’s either Metal or that newer thing they have built specifically for AI. It seems ILGPU only has backends for CUDA, OpenCL, and CPU SIMD.
I have general suspicion towards custom compilers. Writing a good compiler is hard, for GPUs even harder due to the weird execution model and insufficient documentation from GPU vendors. HLSL compiler is supported by Microsoft, and compute shaders are used by many millions of gamers every day. Similarly, CUDA is supported by nVidia, usually pretty stable, my only issue with CUDA is vendor lock-in. I have an impression people are often unhappy with the quality and hardware compatibility of less popular GPU APIs like ROCm and OpenCL.
BTW, I asked a friend to test my program on their low-and laptop. The laptop has some Intel core i3 with integrated GPU. The performance wasn’t great at about 1 token/second (single-channel memory), but at least my code worked. I’m not sure it would have worked on that computer if the backend was based on OpenCL.