1) Let me just say that language integration is a very serious issue. It is not just between completely different languages, but also between modules written in the same language. Once you start using algebraic datatypes to emulate language features the main language lacks, you essentially step into a dynamic sublanguage and have to deal with the friction caused by crossing module boundaries. This friction is both on the programmer's side – he has to deal with writing boilerplate for crossing the boundaries, and on the computer's side which has to do marshaling which is terrible for performance and just nasty.
2) This friction in the case of Futhark will be magnified manifold as it is a completely different language. Since Futhark is not a general purpose language, a realistic use case for it to be called indirectly directly by other languages who will generate code for it. This is generally how high-speed anything is used today. There is a collection of highly optimized, assembly written routines (such as BLAS) bundled into a library and they are called from very slow high-level languages such as Python.
Futhark today is not fit for such a purpose.
* It would be difficult to partition the program written for it into separate pieces. For very simple programs, at a minimum it will generate 2k lines of code (in the C backend).
* This is compounded by the fact that it does not link to the aforementioned optimized libraries, but generates all the code internally. You could then imagine using Futhark intermittently – calling those fast libraries in the main language and using Futhark for the rest since writing code in Futhark is much more convenient compared to C, but then you would need to partition the program and will immediately run into the code bloat issues in the first bullet point.
* Futhark is a very high level language and takes on all the responsibility for managing memory on itself. That even further clashes with the idea of partitioning the program. Memory allocations are extremely slow on the GPU, and in addition to that, they block the whole device meaning they are not asynchronous.
* A minor point of friction compared to the above is that Futhark supports OpenCL which has minor market share instead of Cuda.
3) Based on the above, I question the current integration strategy by Futhark of making backend for different languages. It currently has a C and a Python backend, and an OpenCL CPU backend, and F# is planned, and you can imagine many different backends…
What might be worth trying instead would be to make Futhark an embedded interpreted language. This is not as crazy as it seems – it would define a natural API point for other languages to access it and the interpreter would be responsible for managing GPU memory. It would be a much better model than disposing everything once the program stops running and would allow for efficient intermingling of multiple Futhark programs that could reside in memory. Right now, that sort of thing would be very high friction.
I do not have much advice on how this could be accomplished and would no doubt require much design work, but I am going to try something like that in my own language at some point. I had this crucial insight when I was trying to use it from the language it was written in and realized that it is actually very difficult.
Scala, Clojure and F# in particular had the master stroke of latching themselves to already established ecosystems. Languages targeting the GPU cannot use that strategy directly and will need to be more inventive.