It is the reduction to a smaller “kernel” that is responsible. If you applied the same design and operational model to running regular old processes instead of virtual machines you would also get a system with less security holes than the grossly insecure rat’s nest that is Linux, Windows, or whatever other commercial IT OS you have in mind.
Virtualization is almost entirely orthogonal, if not harmful, to security of the platform and operations. It is not magic pixie dust that makes your operational model more robust. You need a robust operational model, then you can have a robust operational model with virtual machines.
There is a reason why the most secure systems in the world are separation kernel architectures instead of hypervisors even though most of those systems do support virtualization as a feature, just not as the basis of their security propertys.
Let us review a standard operational model:
Virtual machines are usually pre-allocated their total RAM. Virtual machines are usually pre-allocated a number of cores and pinned to them. Virtual machines are usually only allocated a small number of devices such as a virtual block storage device and virtual network device upon which they implement a in-VM filesystem and network stack. Virtual machines usually have no access to shared services provided by the hypervisor.
So we have a operational model where you have to pre-allocate RAM to a process. You have to pre-allocate a whole core and pin the process to it. The process has no access to a global filesystem, network stack, or devices. The process has access to exactly one file, which is logically similar to a virtual block storage device, and a single raw network socket, which is logically similar to a virtual network device. The process has no ability to form a socket to another process, form a new file, or even have any way of interacting with other processes at all. The process has no access to shared services of any kind.
The chasm between that operational model and any commercial IT operating system is immense, being basically the polar opposite in every dimension in the direction of security. Default-deny instead of default-allow. Shared-nothing instead of shared-everything. What you have there is a system even more static and simple than what runs on most microkernels. That is the comparable class of platforms with a similar operational model.
To demonstrate that virtualization is the key factor, you need to demonstrate that actually comparable systems with similar operational models like microkernels have more platform vulnerabilitys than comparable KVM-based, or even just hypervisor-based, systems. Which, again, flies against the face of evidence as the systems that are actually used in high security applications designed to protect against state actors are separation kernels instead of hypervisors.
Since you are varying the security basis, implementation, and operational model simultaneously when you are comparing KVM to Linux to argue that the security basis is the important factor, I get to as well. Except mine is actually more fair because the design of a seL4-based system is actually much more similar to the design of a multi-tenant KVM-based system than the design of the KVM-based system is to the design of a Linux user environment.
https://news.ycombinator.com/item?id=41071954
Rather than replicating my response then, I'll just incorporate that link into my point.
I ported L4 to my ARM NUC a couple weeks ago. L4 is great. Of course, L4 is also principally a platform for virtualization, so it's a pretty odd bit of evidence to try to bring up. Where have you personally been using L4?
As independent evidence for this point, virtualization is not the basis of security/isolation in seL4. You have isolation without any virtual machines. Virtual machines are just a feature on top that can leverage the existing isolation functionality to also provide isolated virtual machines. Of course, a secure deployment then requires you to leverage this foundation with a good operational model and system design since you can always make a insecure system even atop a good foundation. This demonstrates that virtualization is not necessary for a secure base, nor sufficient to achieve highly secure systems.
Just to hammer in the point that you are heavily misinterpreting Theo's response, this is the full sentence at the start of the post that Theo was responding to:
"Virtualization seems to have a lot of security benefits. Rootkits can lie to DomU but not Dom0, and of course snapshotting, migration etc is really nice."
Wow, amazing, rootkits can never be in Dom0 because Xen has "virtualization" magic pixie dust. Theo is pointing out how this is nonsense and virtualization will only provide security if you can create implementations without glaring security holes. Furthermore, you should not just listen to the people who brought you insecure system 1 when they tell you that this time for sure they are going to give you secure system 2; maybe have just a little bit of cynicism and ask for some evidence first.
One of the points I am making is that when deploying on a modern multi-tenant VM-based platform, VM orchestration is analogous to process orchestration. However, you orchestrate the units, VMs, in a very different way to how you would orchestrate processes on say Linux. If your platform orchestrates, configures, and operates processes the same way a VM-based platform orchestrates, configures, and operates VMs and your implementation is solid then you would see similar security benefits. Virtualization is not the key. It just kind of looks that way because virtualization is usually paired with a fundamental re-architecture.
In greenfield application development or cases where you would do single-application VMs, you would target processes/platform directly. Only in situations where you are literally lifting code from a different OS environment that you can not or will not port to the native model would you need to go through a VM. If you orchestrate them the same way and your operational models for them are similar, then you will usually see similar outcomes. This is how it works in high security microkernels/separation kernels which simultaneously allow native processes alongside VMs. Virtualization is a feature to allow non-porting, it is not security/isolation; that is already provided underneath.
That said, I don’t know whether Theo has since “eaten crow” or has otherwise personally evolved.
And even smart people can be wrong. Sometimes a lot.
AMD released (ie commercially available) Pacifica on May 23, 2006 while Intel did released their Vanderpool a half of year earlier November 14, 2005. [0]
Windows Server 2008 was RTM'ed on February 2008 which provided Hyper-V as a first class component. [1]
Virtual Server 2005 R2 SP1 added support for both Intel VT (IVT) and AMD Virtualization (AMD-V) and was released 11 June 2007. [2]
https://en.wikipedia.org/wiki/X86_virtualization#AMD_virtual...
https://en.wikipedia.org/wiki/Windows_Server_2008
https://en.wikipedia.org/wiki/Microsoft_Virtual_Server#Versi...
The Xen hypervisor itself was pretty minimal, as I understand mostly serving to time-slice CPU cycles among guest domains and partition memory access. As a contrast to VMWare, device access and drivers were handled by the guests themselves.
As such, the attack surface of the Xen hypervisor itself is fairly minimal. Most security issues seem to be denial of service vulnerabilities, though there are some privilege escalation, access, information leak, and overflow issues listed:
<https://xenbits.xen.org/xsa/>
I generally respect de Raadt's expertise and instincts, though he may have been over his skis here.
I'm not sure what the purpose of revisiting this is beyond provoking a flamewar on a slow Sunday.
Regardless i do agree with you though, not sure what the point of digging up ancient skeletons is.
"Give me six lines written by the hand of the most honest man, I will find something in them which will hang him."
* https://en.wikipedia.org/wiki/Give_me_the_man_and_I_will_giv...
Is a Proxmox kernel that much smaller than a typical Linux kernel?
That the hypervisor is effectively an operating system/kernel I have always held, and that it is a smaller and thus less vulnerable kernel is an appropriate explication I think. It's very hard to secure an all purpose kernel like Linux without actually building it yourself (and even then..)