990 karma · joined December 2, 2019
I'm the the co-author of the GPU Virtual Machine (GVM project), and LibVF.IO. We just announced our enterprise product based on GVM called GVM Server. I'd love to hear what you all think of the work we've done and give suggestions on where we can improve in the future!
Ideally if GPU virtualization were sufficiently widespread as is support today for Intel VT-d, and AMD-v (IOMMU APIs for hardware assisted CPU virtualization) then software could make use of these functions without the user being aware of it. We're in a situation similar to that of CPU virtualization without hardware assistance with the early Xenoservers project from Cambridge (what would later become the Xen hypervisor and XenSource company). At that time there was not widespread support for virtualization assistance on most CPUs, and as a result Xen used methods like ring de-privileging to place the entire guest in ring 3 (userspace and kernel) while the hypervisor ran in ring 0 in order to virtualize any ordinary CPU model - my understanding is these were known as PV-guests (paravirtual guests). Over time however CPU companies began to introduce widespread support for features like VT-d and AMD-v to all of their models of CPU which enabled VM-exits/context save-restore with the use of shadow registers rather than ring de-privileging while Intel added new 'virtualization enhancements' through feature suites like vPro (SGX2 for example) which were only available on certain models of CPU (for example Xeon devices). Xen would adopt VT-d and AMD-v as HVM-guests (Hardware assisted virtualization) as they became more common on ubiquitous hardware and at the same time commercial forks of Xen would take advantage of these vPro features (like SGX2) for enterprise and high security government use-cases:
https://wiki.xenproject.org/wiki/Xen_Project_Software_Overvi...
Like before (around the time of the Xenoservers project) today we can effectively virtualize the GPU without hardware assistance mechanisms:
https://openmdev.io/index.php/GPU_Support
https://openmdev.io/index.php/Virtual_I/O_Internals#Mdev_Mod...
Since it's now practical to virtualize any GPU device (as was the case in the past with early Xen on CPUs supporting virtualization for various use-cases regardless of whether or not the hardware provided assistance mechanisms) it might then be time to start moving to a new paradigm of 'enterprise' vs. 'consumer' - in other words new 'virtualization enhancements' (similar to vPro on Intel's Xeons, ect..) are developed for enterprise GPUs (for example shadow page deduplication in VRAM, import/export of redundant objects between IO Virtual Address buffers, IOMMU protected balloon/deballoon, ect..) and basic hardware assistance mechanisms like SR-IOV & SIOV are enabled by default, across the board:
https://openmdev.io/index.php/Virtual_I/O_Internals#SR-IOV_M...
https://openmdev.io/index.php/Virtual_I/O_Internals#SIOV_Mod...
Here's the GPU Support page if you'd like to take a look:
https://lwn.net/ml/linux-kernel/20190222021927.13132-1-baolu...
This page has a comparison of the various IO assistance modes GVM can make use of (see comparison of assistance modes, the Mdev Mode section, and the SR-IOV Mode section):
https://openmdev.io/index.php/Virtual_IO_Internals
This will probably also play a role in future developments like SIOV (Scalable IO Virtualization):
https://lwn.net/ml/linux-kernel/20190222021927.13132-1-baolu...
I'll also be attending KVM Forum this year so I'd love to chat with folks there as well! :)
https://openmdev.io/index.php/AMDGPU
Ideally some folks who know about amdgpu might consider helping our open source community by adding similar information to that page to the information we added on the Nvidia Open Kernel Modules page:
https://openmdev.io/index.php/OpenRM
If that could be done then we would do our best to add in AMD support to GVM.
https://openmdev.io/index.php/Mdev-GPU#fbLen
Since this type of load balancing was originally used in the datacenter where virtual machine multi-tenancy was the use case the Quality of Service (QoS) functions here are fairly robust.
This can also be used for various high assurance use cases such as OpenXT / QubesOS (security by compartmentalization). For example a laptop computer being used by an automotive company for CAD to keep their information safe from other programs on the system (like a browser in another VM) the prying eyes of competitors. I made a video of that here:
https://news.ycombinator.com/item?id=32585333
We have also been working on something called "LIME Is Mediated Emulation" which is a Win64 binary compatibility layer similar to WINE but using vGPUs. You can read about that here:
I made this web page to try consolidate some information from various folks who have contributed a lot in this area of open source to show how it works (there are some novel additions we've made as well based on our own work with GVM):
Both VMs are actually based on Windows (one simply doesn't have explorer.exe running) and each has it's own virtual GPU attached.
This laptop has a Nvidia graphics processor inside but it also works with laptops that have an Xe Intel graphics processor. I'm hoping at some point it will work well on AMD too.
"GVM ... may be used in combination with KVM or other platform hypervisors such as Xen* to provide a complete virtualization solution via both central processing (CPU) and graphics processing (GPU) hardware acceleration."
Side note: We also added in support for Libvirt by allowing users to install standalone GVM components via ./scripts/install-standalone-gvm-components.sh.
Most package updates don't disrupt anything. Depending on the vendor driver you use (Intel i915, AMD GPU-IOV Module, Nvidia) kernel updates should be okay if we use DKMS. GPU-IOV Module works okay with DKMS if you know what you're doing. I can't speak too much to i915. On Nvidia DKMS is hit or miss but that depends on which driver version is used. Some folks report DKMS doesn't work at all, others say they've never had a problem with it. I think it also depends on the kernel version DKMS is compiling against as in some cases the driver hasn't been updated yet to support the latest kernel version.
We try to help with that by making patch files that run ahead of official driver/kernel version support (they basically just make the driver work on a newer kernel - that's all):
https://github.com/Arc-Compute/LibVF.IO/tree/master/patches/