Januscape: Guest-to-Host Escape in KVM/x86 [CVE-2026-53359]
github.com
github.com
This is a very nasty vulnerability and risks any service that uses and allows nested x86 virtualization features at risk. Including those running VMs as a service.
> Running the PoC inside a guest VM can trigger a host kernel panic. A full escape exploit that works in a controlled environment also exists, but it is not released at this time and is planned to be released in the very distant future.
The first commit that introduced this vulnerability was in 2010. [1] So it was undiscovered for 16 years until now [2].
It was only a matter of time that a vulnerability in KVM would appear. This one is really not good as it is the first KVM guest-to-host exploit working on both AMD and Intel.
[0] https://github.com/V4bel/Januscape/blob/main/assets/write-up...
[1] https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...
[2] https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...
Publicly undiscovered
For what it's worth, this is a variant of a vulnerability discovered via fuzzing last April, CVE-2026-46113.
For everyone else it's dangerous enough to look seriously at getting an updated kernel or apply mitigations, especially since those are easy (disable nested virtualization, only requires restarting guests). Note that this is true even if you're not running guests, having a user running untrusted code and with access to /dev/kvm is enough.
If you're not running anything untrusted, you probably won't be affected but probably should still look at getting an updated kernel or apply mitigations.
It's the worst class of vulnerabilities for KVM in many years (this is the third variant, after CVE-2026-23401 which is a bit different and not guest teiggerable, and 46113 which I have already mentioned and is basically the same bug as this one), on the other hand it also says something about KVM that nothing similar was found in so many years. It's interesting that while this one was found with AI the first two were found with old school (albeit very sophisticated) fuzzing.
Does anything do a "real" nested virtualization in hardware? s390 might but I know ~nothing about it and it probably does not expose its bare metal hardware layers to Linux/KVM anyway.
Found this https://sixe.eu/news/kvm-nested-vms-ibm-power-linux-lpar
That's showing 3 layers of nesting: LPAR, hypervisor, nested hypervisor
Intel and AMD both have some small amount of acceleration of nested virtualization, respectively with shadow VMCS and virtualized VMLOAD/VMSAVE.
does this mean that you must have nested virtualization enabled to br vulnerable. does disabling this feature in the host os or bios, make you immune to this bug?
If you share resources, that reduces costs, but increases security risks.
choose whether to share a filesystem, an OS, a kernel, hardware, or just use a dedicated server.
The economics of sharing resources are all in a tiny sliver of the budget spectrum, the shoestring budget range :
0-1$/mo: serverless
1$-5$/mo containers
5$-200$/mo Virtual Machine(s)
200$-1Billion$/month , at least one dedicated server
So if your hourly is worth anywhere upwards of 5$/hr, and your project has any semblance of seriousness, just use a dedicated server, and avoid a whole class of LPE vulnerabilities just to save some $.
Businesses have expenses, let's stop pretending that all of these non dedicated server infrastructures are serious. Shell out 200$/month or stick to hobby status.
No, I don't sell dedicated servers, but I should
The reason “managed services” of all kinds, including cloud services, are so widespread in business is because someone else is managing things so that you don’t have to. This is as serious as it gets in business. Managing your own hardware makes very little sense for many, if not most companies.
It does not make sense for a lot of time and scaling. You need 3+ people maintaining it, you have upfront costs in the hundreds of thousands of euros on the very lower end. If you don't utilize that money spent, sucks to be you. You have planning times in the area of months, not hours, unless you keep capacity you don't use around (rackspace, cabling, power/cooling capacity).
On the other hand, if you have that hardware management running, it's very amazing. Before the AI nuke, We were looking at moving various systems fully bare metal, because it would simplify management on both sides a lot, and a common statement I heard is "We don't deal with systems that small. If we do bare metal container hosting, we don't measure in dozens of gigabytes of memory. Your business case validates that investment. Here is btw three test systems about double your requirement, just old".
Before the AI nonsense (HBM Memory Demand -> RAM & SSD prices), this would result in very competitive hosting costs after some scale, when amortized across 5 years and then tossed into the testing environment until it stops functioning. And these testing environments allow for a lot of experimentation and failover testing.
Though now it's all very different and not clear.
Also if you have several servers you do not need to hire a full-time sysadmin.
> Managing a non-trivial hardware fleet requires people, and people cost money.
People in AWS also cost money and guess who is going to cover this cost?
This means that you need sysadmins in close proximity to your hardware to do hardware swapping/troubleshooting. Or you need to engineer your system to not have a SPOF (which is not easy). So you're looking at employing at least 2 engineers near your datacenter.
What's changing is the scope of things that you can run on that one rack. 15 years ago, I was running clusters of 30 computers to do things that I now can do with 1.
Or because it's cost-effective to slice up your infrastructure and sell off the bits you're not using right now.
Even the most basic business app has an app server and a database server. If they have 6 business apps, they'd have at least 7 dedicated servers (assuming we're allowing a database server to have multiple app databases sharing it)
you can do a ton with just a couple of dedicated servers with the redundancy you need
I said dedicated servers, I never said anything about owning or managing the hardware. You can rent a dedicated server.
The decision to virtualize and the decision to own the hardware are separate decisions.
I also dealt with owned servers and I had to deal with power outages and gas based generators, internet outages caused by too high trucks taking out a data line, and UPS beeping because their battery life was nearing zero.
So there's a non trivial difference there.
All of these layers are a form of risk
By your own numbers that's 200x+ as expensive.
Really 20 containers is a pretty small app considering 5 app server containers, a DB, a cache, a load balancer, some monitoring/alerting crap 2x for redundancy.
(To disable nested virtualization on a per-VM basis. Only against exploitation from within that specific VM, obviously does nothing against users with access to /dev/kvm on the host.)
Wouldn't this also be a risk for people using VMs to sandbox untrusted code running on trusted hosts?
Also the vulnerability requires enabling nested virtualization on the VM.
Why on Linux device files are accessible by untrusted applications?
That's been the case forever: /dev/null, /dev/zero, /dev/stdin, ...
Would they potentially be a solution to sudo's all-or-nothing granularity in this domain?
So as a responsible user I am slowly writing my own sandboxes, struggling with lack of documentation and designing workarounds.
Which is precisely why many kinds of kernel feature should be exposed as operations on device nodes, not as system calls usable out of thin air. UGO and ACL permissions work on device nodes!
2: Because it's desirable for users to be able to run VMs.