Hybridizer: High-Performance C# on GPUs
devblogs.nvidia.com
devblogs.nvidia.com
When it works. It works. When it fails you're in for a world of pain. A lot is going on behind the scenes. The data structures, classes need special annotations to carefully translate the c# structures into something that cuda understands. You have to pay special attention to your object hierarchy, you have to be aware of all C# keywords that are supported and not supported. The fields in your class may be misaligned by a few bytes of the wrong annotation is used.
Don't get me started on cuda pointers - you're not shielded from this. My experience with this has put me off cuda frameworks in general.
You're better off learning cuda or hiring a cuda Dev than investing heavily in stacks like this.
My background is in HPC in finance.
(Same story with those C# to Javascript compilers)
Very few quants know CUDA. My brother happens to be one of them but he does most of his development now in python/pandas. He gets someone else to do the heavy lifting to target the HPC platform of choice.
Now the motivation behind the layer or framework is to provide greater productivity, remove the step of pairing quants with cuda Devs when productionsing the code.
These are all valid reasons to adopt such frameworks. But in practice the cost of upgrades, strange pointer bugs and delays to production releases don't make it worth it. Also you need the CUDA Dev for situations that the framework fails. If such a framework does become available that removes these pain points, I'd adopt it in a heartbeat
This may vary in other industries other than finance.
Many times there is a need of ongoing investment, where things are failing pretty bad, until they actually take off.
Just yesterday I saw a documentary about Gotthard Tunnel, also considered impossible to achieve, with high human costs and risk of insolvency until they finally managed it.
John Harrison took almost his entire life to create the first usefull marine chronometer, amid disbelief and issues to get proper funding.
Or if you want to bring it closer to home, very few people believed JavaScript would ever become fast or even leave the browser.
[1] http://www.aleagpu.com/release/3_0_4/doc/introduction.html
Am I to assume the point here is to allow the C# programmer to run tight loops of parallel numerical code on the GPU, with minimal adaptation?
They support virtual functions, though.
Neat project either way, but the scope isn't clear.
I mean, it is surely nice to easily launch a bunch of threads within a single-source program, but there are already plenty of C++ libraries that let you do this, and it has not really led to an explosion of making efficient GPU programming accessible to the layman.
with the CUDA backend: https://hackage.haskell.org/package/accelerate-llvm-ptx
See http://chimera.labs.oreilly.com/books/1230000000929/ch06.htm... for a great introduction
I would be interested in anything that can serve as good building block material that has a good set of primitives built in.
However, is the matrix formulation the best one, if we have access to a programming language (like CUDA) where we have more flexibility in how we perform computation? For example, while we might be able to express k-means clustering as a matrix operation, we might express it more efficiently by programming it directly.
If there aren't any redundant operation, matrices give a good abstraction that does not leak much. Libraries take care of the best use of cache.
If you have to build castles out of individual grains of sand and burnt clay, it will limit how many of them you can build. That's why building blocks are useful. Where the building blocks don't quite fit, there is always sand, clay and mortar to fill the gaps.
This seems like the choice when you want to have custom kernels run on your large arrays.
https://www.roguewave.com/products-services/imsl-numerical-l...