There's also the issue the gpus are not premptable which kind of makes preempitve multi-tasking hard
There's also the issue the gpus are not premptable which kind of makes preempitve multi-tasking hard
Intel's latest gpu architecture has an embedded OS running on the gpu for scheduling command batches, I'm not sure what AMD and Nvidia do.
I still wouldn't write a general purpose OS for it.
It will also be sold in stand alone chips soon
Same on AMD and NVidia, except it's been like this for the past 10-15 years (depending if you count at the bottom or at the top of the hardware release pipeline).
If you're curious, lookup the Intel Broadwell GPU specs, there's sections devoted to the various levels of preemption. If you're really curious look up the workarounds needed for the finest grained preemption (this would be preempting a single GPGPU draw call).
Then decide enabling fine grained preemption should probably wait for Skylake, unless you took too much Adderall and no challenge sounds impossible. Do I speak from personal experience? I plead the fifth.
I've no experience with how fine grained nvidia's preemption is.
The advantage is that most of the silicon can go towards the actual computation rather than stuff like branch-prediction and out-of-order execution. The disadvantage is that branching and looping is problematic: when only one item wants to go down the other branch of an if-else-statement, the GPU has to run through both branches for all items (and execution is masked off on a per-item basis).
This works extremely well for graphics and high-dimensional numerics workloads (linear algebra, finite elements). It doesn't work at all for, say, spell-checking.