Vgpu_unlock: Unlock vGPU functionality for consumer grade GPUs
github.com
github.com
I'm very grateful I wasn't required to figure that out.
What bugs me about companies like NV is that if they just sold their hardware and published the specs they'd probably sell more than with all this ridiculously locked down nonsense, it is just a lot of work thrown at limiting your customers and protecting a broken business model.
If Nvidia enabled all their professional features on all gaming SKUs, the only reason to buy a professional SKU would be additional memory.
Today, they make almost $1B per year in the professional non-datacenter business alone. There is no way they’d be able to compensate that revenue with volume (and gross margins would obviously tank as well, which makes Wall Street very unhappy.)
That’s obviously even more so in today’s market conditions.
Do you feel it’s justified that you have to pay $10K extra for the self-driving feature on a Tesla? Or should they also be forced to give away that feature for free? After all, it’s just a SW upgrade. (Don’t mistake this for an endorsement...)
I feel like i already paid for the hardware. If telsa says it's cheaper for them to stick the necessary hardware into every car, i'm still paying for it if i buy one without self-driving. Thus if i think the tesla software isn't worth 10k and i'd rather use openpilot, i feel like i should have the right to do that.
But nvidia is also actively interfering with open source drivers (nouveau) with signature checks etc.
The whole focus on hardware is just bizarre.
When you buy a piece of SW that has free features but requires a license key to unlock advanced features, everything is fine, but the moment HW is involved all of that flies out of the window.
Extra features cost money to implement. Companies want to be paid for it.
A company like Nvidia could decide to make 2 pieces of silicon, one with professional features and once without. Or they could disable it.
Obviously, you’d prefer the first option, even if it absolutely makes no sense to do so. It’d be a waste of engineering resources that could have been spent on future products.
Deciding to disable a feature on a piece of silicon is no different than changing a #define or adding “if (option)” to disable an advanced feature.
By doing so, I have the option to not pay for an advanced feature that I don’t need.
I don’t want the self-driving option in a Tesla and I’m very happy to have that option.
The issue is not that Tesla FSD should come with the hardware, the issue is that if I buy the hardware I should have the right to do whatever I want with it, and so we shouldn't leave aside that Tesla prevents us from running our own software.
This is relevant to the NVidia situation since their software doesn't add features, it limits things the chip is already capable of. Just like Tesla won't let you run Comma.AI or something similar on their hardware...
This doesn't make sense at all. In your scenario you always pay for what you get. and developing additional features has a non-zero cost asociated with it (Unless you download software which targets unskilled consumers, Like Chessbase's Fritz engine which was essentially stockfish but 100$ instead of 0)
>Extra features cost money to implement. Companies want to be paid for it.
This doesn't make sense in your scenario either. You already have the sillicon with the 'advanced features' in your hands. The reason they lock the feature is so that you have to buy a more expensive card, with overpowered hardware that you dont need, in order to use a feature that all cards have if it weren't disabled. The only reasonable explanation you could have at this point that doesn't involve monopolistic practices to make more money (Nothing wrong with that) is that the development of the feature itself was so prohibitively expensive that it required consumers to pay for much higher margin cards in order to offset the development costs. Which is what's happening
>A company like Nvidia could decide to make 2 pieces of silicon, one with professional features and once without. Or they could disable it.
That would cost alot of money. all the more reasons why it might have been done to upsell more cards instead of offering quantitative improvements for a different price.
Yes. Who actually likes being segmented into markets? We want to pay a fair price for products instead of being exploited.
> If Nvidia enabled all their professional features on all gaming SKUs, the only reason to buy a professional SKU would be additional memory.
So what? A GPU is a GPU. It's all more or less the same thing. They would not have to lock down hardware features otherwise.
> Today, they make almost $1B per year in the professional non-datacenter business alone. There is no way they’d be able to compensate that revenue with volume (and gross margins would obviously tank as well, which makes Wall Street very unhappy.)
Who cares really. Pursuit of profit does not excuse bad behavior. They should lose money every time they do it.
Would you prefer if companies made 2 separate pieces of silicon designs, one with virtualization support in HW and one without, even if it would reduce their ability to work on advancing the state of the art due to wasted engineering resources?
Or would you prefer that all features are enabled all the time, but with the consequence that prices are raised by, say, 10% for everyone, even though 99% of customers don’t give a damn about these extra features?
I'm against "licenses" in general. If your software is running on my computer, I make the rules. If it's running on your server, you make the rules. It's very simple.
> do you have the same irrational belief that only products that include a HW component shouldn’t be allowed to charge for advanced features?
When I buy a thing, I expect it to perform to its full capacity. Nothing irrational about that.
> Would you prefer if companies made 2 separate pieces of silicon designs, one with virtualization support in HW and one without
Sure. At least then there would be real limitation rather than some made up illusion.
> even if it would reduce their ability to work on advancing the state of the art due to wasted engineering resources?
The real waste of engineering resources is all this software limiter crap. They shouldn't even be writing drivers in the first place. They're a hardware company, they should be making hardware and publishing documentation. Instead they're locking out open source developers, adding DRM to their cards and blocking workloads they don't like.
> Or would you prefer that all features are enabled all the time, but with the consequence that prices are raised by, say, 10% for everyone, even though 99% of customers don’t give a damn about these extra features?
That is how things are supposed to work, yes.
> Today, they make almost $1B per year in the professional non-datacenter business alone. There is no way they’d be able to compensate that revenue with volume (and gross margins would obviously tank as well, which makes Wall Street very unhappy.)
You're looking at it wrong. If Nvidia were to enable all features on their hardware, they wouldn't be giving up that additional revenue, they would instead have to create differentiated hardware with and without certain features.
Their costs would increase somewhat (as currently their professional SKUs enjoy some economies of scale by virtue of being lumped in with the higher-volume gaming SKUs), but it would hardly be the catastrophe you're describing. The pro market is large enough to enjoy it's own economies of scale, even if the hardware wasn't nearly identical (which it still would be).
Also I have no doubt people will find other ways to unlock the hardware.
If the workaround results in enough money being left on the table, this might prompt 3rd party investment in open source drivers in order to keep the workaround available by eliminating the dependence on Nvidia's proprietary drivers.
It's currently impossible to find any nVidia GPU in stock because the demand far outstrips the supply.
Market segmentation is only helping their profit margins, not hurting it.
This is not to downplay the OP, of course - this is truly great and I'm sure it was a lot of work. But the hardware part is not new.
[0] https://web.archive.org/web/20200814064418/https://www.eevbl...
This not only enables the use of GPGPU on VMs, but also enables the use of a single GPU to virtualize Windows video games from Linux!
This means that one of the major problems with Linux on the desktop for power users goes away, and it also means that we can now deploy Linux only GPU tech such as HIP on any operating system that supports this trick!
If it's such a cool feature, why does NVidia lock it away non-Tesla H/W?
[EDIT]: Funny, but the answers to this question actually provide way better answers to the other question I posted in this thread (as in: what is this for).
GPUs are a duopoly due to intellectual property laws and high costs of entry (the only companies I know of that are willing to compute are Chinese and only a result of sanctions), so for NVidia this just allows for more profit.
This has the potential to challenge the cloud case for sporadic GPU use, since cloud vendors cannot buy RTX cards. But it would require that the tooling becomes simple to use and reliable.
The RTX 3090, with 24GB of VRAM is at USD 1499.
Customer dGPUs from other HW providers do not have virtualisation capabilities either.
Not anymore.
Nvidia used to detect if the host is a VM and return error code 43 blocking them from being used (for market segmentation between GeForce and Quadro). This is usually solved by either patching VBIOS or hiding KVM from the guest, but it was painful and unreliable. Nvidia removed this limitation with RTX 30 series.
This vGPU feature unlock (TFA) would allow GPU to be virtualized without requiring the GPU to first be detached from the host, vastly simplify the setup and open up the possibility of having multiple VMs running on a single GPU, all with its own dedicated vGPU.
[1]: https://wiki.archlinux.org/index.php/PCI_passthrough_via_OVM...
NVIDIA's stock price has doubled since March 2020, and most of these gains can be largely attributed to the outstanding growth of its data center segment. Data center revenue alone increased a whopping 80% year over year, bringing its revenue contribution to 37% of the total. Gaming still contributes 43% of the company's total revenues, but NVIDIA's rapid growth in data center sales fueled a 39% year-over-year increase in its companywide first-quarter revenues.
The world's growing reliance on public and private cloud services requires ever-increasing processing power, so the market available for capture is staggering in its potential. Already, NVIDIA's data center A100 GPU has been mass adopted by major cloud service providers and system builders, including Alibaba (NYSE:BABA) Cloud, Amazon (NASDAQ:AMZN) AWS, Dell Technologies (NYSE:DELL), Google (NASDAQ:GOOGL) Cloud Platform, and Microsoft (NASDAQ: MSFT) Azure.
https://www.fool.com/investing/2020/07/22/data-centers-hold-...
If you're brave enough, you can already do that with GPU passthrough. It's possible to detach the entire GPU from the host and transfer it to a guest and then get it back from the guest when the guest shuts down.
This could allow a single GPU with a single video output to be used to run games in a Windows VM, without all the hoops that GPU passthrough entails. I'd definitely be excited for it!
From experience not always. If the dedicated GPU gets selected as BIOS GPU then it might be impossible to reset it properly for the redirect. I had this problem with 1070.
I have to say vGPU is amazing feature, and this possibly brings it to "average" user (as average user doing GPU passthrough can be).
It doesn't. As I said, you can detach the GPU from the host and pass it to the guest and back again. I elaborated a bit more in another comment [0].
> This could be way more practically useful than GPU passthrough.
I think that depends on the mechanics of how it works. How exactly do you get the "monitor" of the vGPU?
Oh, i got sidetracked. I have a kernel command line that invokes IOMMU and "blacklists" the set of PCIE lanes that GPU sits on, the kernel never sees it, even when its in use. The next thing that i had to do was set up a vfio-bind script, that just tells qemu what GPU it's going to use. Thirdly, and this is the unfortunate part, since i forgot exactly what i did - there's some weirdness with windows in qemu with a passthru GPU - you have to registry hack some obscure stuff in to the way windows handles the GPU memory.
If i am not mistaken, 95% of all of my issues were solved by reading the ArchLinux documentation for qemu host/guests. My system is ryzen 3600, 64GB of ram, 2x NVME drives + one M.2 Sata drive, a gtx 1060 and a gtx 1070. Gentoo gets 16GB of ram (unless i need more, i just shut down windows or reset the guest memory) and the 1060. Windows gets ~47GB of ram, and the 1070, a wifi card, and a USB sound card. One of the things you quickly realize with guests on machines like this is that consumer grade motherboards and CPUs are garbage, there aren't enough PCIe lanes to, say, passthrough a bunch of USB or SAS/SATA ports, or a dedicated PCIe soundcard, or firewire. If you have an idea that you'd really like to try this out as an actual "desktop replacement" - especially for replacing multiple desktops, i recommend going to at least a threadripper, as those can expose like 4-6 times as many PCIe lanes to the host OS, meaning the possibility of multiple guests on multiple GPUs, or a single "redundant" guest, with USB ports, SATA ports, and pcie sound/firewire/whatever.
Why would anyone do this? dd if=/dev/sdb of=/mnt/nfs/backups/windows-date.img . Q.E.D.
I’m planning on switching to VFIO for my next rebuild, and was curious as to how stable the setup was.
I really hope Intel sees this as an opportunity for their DG2 graphics cards due out later this year.
If anyone from Intel is reading this: if you guys want to carve out a niche for yourself, and have power users advocate for your hardware - this is it. Enable SR-IOV for your upcoming Xe DG2 GPU line just as you do for your Xe integrated graphics. Just observe the lengths that people go to for their Nvidia cards, injecting code into their proprietary drivers just to run this. You can make this a champion feature just by not disabling something your hardware can already do. Add some driver support for it in the mix and you'll have an instant enthusiast fanbase for years to come.
You don’t need vgpu to get the job done. I’ve had two set ups over time: one based on a jank old secondary gpu that is used by the vm host, another based on just using the jank integrated graphics on my chip.
Even still, I dual boot because it just works. It always works, and boot times are crazy low for Windows these days. No fighting with drivers. No fighting with latency issues for non-passthrough devices. It all just works.
I consider true peripheral multiplexing with true GPU virtualization to be the way of the future. It's true virtualization and doesn't even require you to sacrifice and/or babysit a single PCIe connected GPU. Passthrough is just a temporary hacky workaround that people have to apply now because there's nothing better.
In the best case scenario - with hardware SR-IOV support plus basic driver support for it, enabling GPU access in your VM with SR-IOV would be a simple checkbox in the virtualization software of the host. GPU passthrough can't ever get there in terms of usability.
For the cloud, I could imagine wanting vGPUs so you can shard the massive GPUs that are used there. But in cloud, you would then have a single device be multi-tenant, which is a bit spicy security wise. Passthrough has a very straightforward security model.
saying that, these days I just have a second pc with a load of cheap USB switches...
Now I have additional RTX 2070 and it works ok.
I also have a macos vm, but I didn't set up gpu passthrough for that. Tried it once, it hung, didn't try it again. I use remote desktop anyway.
here are some misc links:
https://manjaro.site/how-to-enable-gpu-passthrough-on-proxmo...
https://manjaro.site/tips-to-create-ubuntu-20-04-vm-on-proxm...
https://pve.proxmox.com/wiki/Pci_passthrough
https://blog.konpat.me/dev/2019/03/11/setting-up-lxc-for-int...
What is the current licensing situation on this? Can I use it legally to build software for Mac?
- nvidia-patch [0] "This patch removes restriction on maximum number of simultaneous NVENC video encoding sessions imposed by Nvidia to consumer-grade GPUs."
- About a week ago "NVIDIA Now Allows GeForce GPU Pass-Through For Windows VMs On Linux" [1]. Note, this is only for the driver on Windows VM guests not GNU/Linux guests.
Hopefully the project in the OP will mean that GPU access is finally possible on GNU/Linux guests on Xen, thank you for sharing OP.
[0]: https://github.com/keylase/nvidia-patch
[1]: https://www.phoronix.com/scan.php?page=news_item&px=NVIDIA-G...
https://github.com/DualCoder/vgpu_unlock/blob/master/vgpu_un... (the comments behind the device ids)
If I go with vGPU solution, I don't need to turn on / off NVIDIA driver for these PCIe lanes when running Windows VM? (I won't use these GPUs on host machine for display).
The latter statement is correct. The GPU can be attached to the host but it has to be detached from the host before the VM starts using it. You may also need to get a dump of the GPU ROM and configure your VM to load it at start up.
Regarding the script, mine resembles [0]. You need to remove the NVIDIA drivers and then attach the card to VFIO. And then the opposite afterwards. You may also need to image your GPU ROM [1]
[0]: https://techblog.jeppson.org/2019/10/primary-vga-passthrough...
[1]: https://clayfreeman.github.io/gpu-passthrough/#imaging-the-g...
There's a lot of detail on the link which I appreciate but maybe I missed it.
Will definitely be checking this out!
I just want to send drawing commands from a node over network to another node with a GPU. Like "draw a black rectangle 20,20,200,200 on main GPU on VM at 192.168.1.102".
How would I do that in the simplest possible way? Is there some network graphics command protocol?
Like X11, but simpler and faster, just the raw drawing commands.
https://www.nvidia.com/en-us/data-center/virtual-solutions/
TL;DR: seems to be something useful for deploying GPUs in the cloud, but I may not have understood fully.
What would be the main use case?
What's the practical use case, as in, when would I need this?
[EDIT]: To maybe ask a better way: will this practically help me train my DNN faster?
Or if I'm a cloud vendor, will this allow me to deploy cheaper GPU for my users?
I guess I'm asking about the economic value of the hack.
Running CUDA in VMs
Running transcoders in VMs
Running <anything that needs a GPU> in VMs
Please see my edit.
Probably not. It will only help you if you previously needed to train it on a CPU because you were in a VM, but this seems unlikely. It will not speed up your existing GPU in any way compared to simply using it bare-metal right now.
> Or if I'm a cloud vendor, will this allow me to deploy cheaper GPU for my users?
Yes. This ports a feature from the XXXX$-range of GPUs to the XXX$-range of GPUs. Since the performance of those is similar or nearly similar, you can save a lot of money this way. It will also make the entry costs to the market lower (i.e. now a hypervisor could be sub-1k$, if you go for cheap parts).
On the other hand, a business selling GPU time to customer will probably not want to rely on a hack (especially since there's a good chance it's violating NVidias license), so unless you're building your on HW, your bill will probably not drop. But if you're an ML startup or a hobbyist, you can now cheap out on/actually afford this kind of setup.
The host still wants access to the GPU to do stuff like compositing windows and H.265 encode/decode.
However, I was under the impression - at least on Linux - that I could run multiple workloads in parallel on the same GPU without having to resort to vGPU.
I seem to be missing something.
What you're saying is true, but it's generally using either the API remoting or device emulation methods mentioned on that wiki page. In those cases, the VM does not see your actual GPU device, but emulated device provided by the VM software. I'm running Windows within Parallels on a Mac, and here[2] is a screenshot showing the different devices each sees.
In the general case, the multiplexing is all software based. The guest VM talks to the an emulated GPU, the virtualized device driver then passes those to the hypervisor/host, which then generates equivalent calls on to the GPU, then back up the chain. So while you're still ultimately using the GPU, the software-based indirection introduces a performance penalty and potential bottleneck. And you're also limited to the cross-section of capabilities exposed by your virtualized GPU driver, hypervisor system, and the driver being used by that hypervisor (or host OS, for Type 2 hypervisors). The table under API remoting shows just how varied 3D acceleration support is across different hypervisors.
As an alternative to that, you can use fixed passthrough to directly expose your physical GPU to the VM. This lets you tap into the full capabilities of the GPU (or other PCI device), and achieves near native performance. The graphics calls you make in the VM now go directly to the GPU, cutting out game of telephone that emulated devices play. Assuming, of course, your video card drivers aren't actively trying to block you from running within a VM[3].
The problem is that when a device is assigned to a guest VM in this manner, that VM gets exclusive access to it. Even the host OS can't use it while its assigned to the guest.
This article is about the fourth option – mediated passthrough. The vGPU functionality enables the graphics card to expose itself as multiple logical interfaces. So every VM gets its own logical interface to the GPU and send calls directly to the physical GPU like it does in normal passthrough mode, and the hardware handles the multiplexing aspect instead of the host/hypervisor worrying about it. Which gives you the best of both worlds.
[1] https://en.wikipedia.org/wiki/GPU_virtualization
[3] https://wiki.archlinux.org/index.php/PCI_passthrough_via_OVM...
Basically any workload that requires sharing a GPU between discrete VMs
I've been (more or less) accused of being an Nvidia fanboy on HN previously but this is an area where I've always thought Nvidia has their market segmentation wrong. Just wrong.
This is great work. (Period, as in "end of sentence, mic drop").
Does this work also on Xen? NVIDIA drivers were always non-functional with Xen+consumer cards.
My understanding is that CPU/GPU per application can make only single draw call in sequential manner. (eg. CPU->GPU->CPU->GPU)
Could vgpu's be used for concurrent draw calls from multiple processes of an single application ?
The limitation you're probably thinking of is in the OpenGL drivers/API, not in the GPU driver itself. OpenGL has global (per-application) state that needs to be tracked, so outside of a few special cases like texture uploading you have to only issue OpenGL calls from one thread. If applications use the lower-level Vulkan API, they can use a separate "command queue" for each thread. Both of those are graphics APIs, I'm less familiar with the compute-focused ones but I'm sure they can also process calls from multiple threads.
Threaded Computation on CPU -> Single GPU Call -> Parallel Computation on GPU -> Threaded Computation on CPU ...
I wonder if it can be used in such way:
Asyc Concurrent Computation on CPU -> Asyc Concurrent GPU Calls -> Parallel Time Independent Computations on GPU -> Asyc Concurrent Computation on CPU
hacks like nucleus coop[1] to run multiple instances of games on a desktop work via a shared screen & hacks, but I'd rather have a multi-head desktop with a bunch of gaming vms.
https://github.com/DualCoder/vgpu_unlock/blob/0675b563acdae8...
Yet AMD is no different there. They also locked SR-IOV for premium datacenter hardware only and they certainly want to keep some features away from consumers.