Modern GPUs have 32-64 independent cores, each can run different code path. Then each core runs the same code path on several different inputs in parallel.
It's probably not possible with current tools, but architecture-wise completely dividing the cores between VMs should be doable.