HNHacker News
TopNewBestAskShowJobs

e-kayrakli

12 karma · joined January 12, 2024

submissionscomments
e-kayrakli··on Introduction to GPU Programming in Chapel
Locales do control where the data is stored. For example:

  var HostArr: [1..10] int;  // allocated on the host memory
  
  on here.gpus[0] {
    // now we are on a GPU sublocale...
    var DevArr:[1..10] int;  // allocated on the device memory
    ...
  }
In the near term, we are planning to publish our 2nd GPU blog post where we will discuss how to move data between device and host.
e-kayrakli··on Introduction to GPU Programming in Chapel
I'll also add to @danilafe's reply that we have a GpuDiagnostics module which would count kernel launches or report them as they occur in a section of code. Something like that can be used to debug parts of your code where you do or don't expect kernel launches to occur. See https://chapel-lang.org/docs/main/modules/standard/GpuDiagno...
e-kayrakli··on Introduction to GPU Programming in Chapel
Thanks for bringing this up. I posted an answer to the same question here: https://news.ycombinator.com/item?id=39009566

But seeing that we already have two questions about it makes me think whether this should be something we should think more about when we are prioritizing work.

e-kayrakli··on Introduction to GPU Programming in Chapel
This is a good question, and I completely agree that programming them can feel... strange.

We approach this as two separate problems:

1. Can Chapel utilize tensor cores under the hood? We have LinearAlgebra and BLAS (and LAPACK) modules in Chapel. They have not been integrated with the GPU support so far. But we want them to be able to use libraries like cuBLAS for GPU-based array under the hood for GPU-allocated arrays. That model can enable Chapel to exploit tensor cores efficiently and seamlessly. Of course, you can extrapolate this to ML/AI APIs and potential Chapel modules correspond to them for example.

2. Can Chapel enable general-purpose tensor programming? This is definitely more challenging, where the main challenge is whether we can make tensor programming portable in such a way that the same code can be used on a non-GPU (and non-vector) processor. Probably a relatively shorter-term solution is to provide a low-level interface that's not much different than CUDA's Warp Matrix Functions or ROCm's rocWMMA interfaces and live with that for a while even though it is not portable. One of the thoughts that I can ignore is whether general-purpose tensor programming is _that_ common vs using tensor cores under the hood through some library as I described in (1).

e-kayrakli··on Introduction to GPU Programming in Chapel
Thanks, @subharmonicon!

While Chapel can run on many different systems, the main goal is making HPC programming much easier. Therefore, we are currently focusing on hardware that you can find in HPC systems (NVIDIA, AMD and Intel). Metal doesn't fall into that category, unfortunately. So far, the name came up infrequently in our discussions IIRC (especially targetting SPIRV), but we haven't heard from any [potential] user who may be interested in it. I would encourage you or anybody else interested in it to create an issue asking for the feature: https://github.com/chapel-lang/chapel/issues/new. Seeing public interest in that direction can change our prioritization.

One thing that I wanted to add that's not in the blogpost is the "cpu-as-device" mode. With that mode, you can use any machine, even one without a GPU, to write applications using Chapel's GPU features. That mode is for those who want to do initial development/debugging on their personal laptops before putting their application on an HPC system. In other words, while you can't use Metal directly, you can still write GPU-enabled applications in your Mac using Chapel, if the end goal is to run it on an HPC system. More details on cpu-as-device: https://chapel-lang.org/docs/main/technotes/gpu.html#cpu-as-...

e-kayrakli··on Introduction to GPU Programming in Chapel
Re Intel support: That's definitely in our plans. However, there are also many other areas that we are actively working on to add more features, fix bugs and improve performance. When prioritizing we typically make decisions based on what our current and potential users might need in the language. Frankly, we are not seeing a big push for Intel GPU support so far. So, currently it is not near the top of our priorities. If you (or other readers) have any input on that matter where lack of Intel support might be a blocker for testing Chapel and/or its GPU support out, definitely let us know.

Re implicit serialization: To clarify; the serialization based on order-dependence is not implicit. The users should use `for` loop if their loop is order-dependent and `foreach` (and `forall`) if their loop is order-independent. In other words, the Chapel compiler doesn't make decisions about order-dependence. In particular for GPU execution: a `for` loop will never turn into a GPU kernel.

There are however some cases where a `foreach` does not turn into a kernel. You may be referring to those cases, but that's not related to order-dependence. Some Chapel features cannot execute on GPU. If your `foreach` loop's body uses any of those features then it will not be launched as a kernel even though `foreach` signals order-independence. Now, a subset of such features that makes an order-independent loop GPU-ineligible are there because we haven't gotten a chance to properly address them, yet. Another subset of such features will remain thwarters for a longer time and maybe forever. For example, your `foreach` loop could be calling an extern host function.