In the interim:
* The Linux virtualization interface has been standardized --- everything uses the same small KVM interface
* Security research matured and, in particular, mobile device jailbreaking have made the LPE attack surface relevant, so people have audited and fuzzed the living hell out of KVM
* Maximalist C/C++ hypervisors have been replaced by lightweight virtualization, which codebases are generally written in memory-safe Rust.
At the very least, the "nearly full kernel" thing is totally false now; that "extra" kernel (the userland hypervisor) is now probably the most trustworthy component in the whole system.
I would be surprised if even Theo stuck up for that argument today, but if he did, I think he'd probably get rinsed.
If I put up a 1 G$ bug bounty, do you think somebody would be able to claim it within a year? How about 10 M$? Please justify this in light of Google only offering 250 k$ [1] for a vulnerability that would totally compromise the security foundation of the multi-billion (trillion?) dollar Google Cloud.
Please also justify why the number you present is adequate for securing the foundation of the multi-trillion dollar cloud industry. I will accept that element on its face if you say the cost would be 10 G$, but then I will demand basic proof such as formal proofs of correctness.
[1] https://security.googleblog.com/2024/06/virtual-escape-real-...
My bigger problem though: I gave you a bunch of substantive, axiomatic arguments, and you responded to none of them. Of the three of them, which were you already aware of? How did your opinion change after learning about the other ones? You cited a 2007 Theo argument in 2024, so I'm going to have trouble with the idea that you were aware of all of them; again, I think even Theo would be correcting your original post.
later
You've written about the vulnerability brokers you know in other posts here; I assume we can just have a substantive, systems based debate about this claim, without needing to cite Theo or Joanna Rutkowska or whatever.
Furthermore, you have presented no empirical evidence that those changes actually result in meaningful security. No, I do not mean “better”, I mean meaningful, as in can protect against commercially-motivated hackers.
None of the systems actually certified to protect against state-actors used such a nonsensical process as imagining improvements and then just assuming things are better. Show a proof of correctness and a NSA pentest that fails to find any vulnerabilities, then we can start talking. Barring that, the explicit, above-board bug bounty provides a okay lower bound on security. You really need a more stable process, but it is at least a starting point.
And besides that, a 7-figure number is paltry. Google Cloud brings in, what, 11 figures? The operations of a billion dollar company should not be secured to a level of only a few million dollars.
So again, proofs of correctness and demonstrated protection against teams with tens to hundreds of millions in budget (i.e team of 5 competent offensive security specialists for 2 years, NSO group for a year, etc.). Anything less is insufficient to bet trillions of dollars of commerce and countless lives on.
Actual LOL at "an NSA pentest".
Slightly later
A friend points out I'm being too harsh here, and that lots of products do in fact get NSA pentests. They just never get the pentest report. We regret the error.
Common Criteria SKPP [1].
“AVA_VLA EXP.4.3E The NSA evaluator shall perform independent penetration testing.
AVA_VLA _EXP.4.4E The NSA evaluator shall determine that the TOE is resistant to penetration attacks performed by an attacker possessing a high attack potential.”
Which has been successfully certified against [2] for use in the F-22 and F-35.
Sneering dismissal of the existence of high security systems with actual empirical evidence while demanding recognition for unproven systems is ridiculous.
Oh, and by the way, the SKPP, like Level A1, required formal specifications and proofs. So, no, a demand to prove such systems are secure is not a winning move.
[1] https://www.commoncriteriaportal.org/nfs/ccpfiles/files/ppfi...
[2] https://www.commoncriteriaportal.org/nfs/ccpfiles/files/epfi...
Not for nothing but it’s hard to follow what you are even arguing.
2. We have always known that such attacks would eventually become feasible to execute once the hackers matured. This was so obvious that government security standards/certifications such as the Rainbow Series codified in the 80s already considered such threats and placed thwarting them as the middle-levels of security. So, we have always known that "adequate" security against commercially motivated attackers has always demanded protection against skilled teams.
3. The commercial IT systems in regular use have never once, over 40 years, demonstrated the levels of security needed to protect against such commercially motivated attacks. These techniques and processes have failed continuously to achieve adequate security despite claiming to solve it every year for 40 years. At some point you need to stop listening.
4. There are systems that did achieve adequate security against commercially motivated attackers and even security against state actors in those 40 years. These standards were not pie-in-the-sky unreachable goals. They were "practical" if you actually cared about security.
5. KVM is in the commercial IT system category, being derived and developed by people who have never once deployed, developed, designed, or likely even seen a system known to have adequate security. Such systems DO NOT get the benefit of the doubt. There is no metric of evaluation, no means of evaluating if the theorized improvements achieve adequacy. In fact, you would be hard-pressed to find literally anybody who would stand up and say: "KVM is unhackable by any team with a budget of 10 M$" (budget including the average salary for the team members so you do not get a talented team doing it just to prove it can be done). Nobody will vouch for the system claiming it achieves even the bare-minimum requirements for adequate security I stated above. That it might be "better" than the Linux kernel is irrelevant; bad is also better than terrible, but it is still bad. At the end of the day they are not meaningfully different; they are all inadequate and unfit for purpose.
So, you can disagree on two primary positions:
1. 10 M$ is too high of a standard for mid-size enterprise security.
2. Commercial IT systems, such as KVM, achieve the 10 M$ level.
You could also argue that I am being reductive by defining security as the cost for a successful attack. But that is silly because it is the one metric that exactly aligns with the operational goals. Every other metric is a means of helping quantify the cost of a successful attack. To put it another way, if you had a magic genie that told you that number, you would not even bother with any other metrics; you have a direct line to what matters.
In summary, it is none of the above. The betterness of KVM is irrelevant except if it achieves adequacy. There is no evidence of that and until there is, KVM is not categorically distinct from any other inadequate system. So, the original point still holds, the reason these "secure VM" techniques are not applied to make a "secure OS" is that there are no "secure VM" techniques to be found in that corner of the world.
I’m still not sure I understand what that means for your argument but a kvm exploit, especially a jailbreak, would be one of the highest value exploits in the world.
To quote the KVM escape Google Project Zero published in 2021 [1]:
"While we have not seen any in-the-wild exploits targeting hypervisors outside of competitions like Pwn2Own, these capabilities are clearly achievable for a well-financed adversary. I’ve spent around two months on this research, working as an individual with only remote access to an AMD system. Looking at the potential ROI on an exploit like this, it seems safe to assume that more people are working on similar issues right now and that vulnerabilities in KVM, Hyper-V, Xen or VMware will be exploited in-the-wild sooner or later."
A single, albeit highly capable, individual found a critical vulnerability in 2 months of work. KVM was already mature and the foundation of AWS at that time and people were already saying that it was highly secure and that it must be highly secure since it would be such a high value target, so logically it must be secure since only an incompetent would poorly secure high value targets, thus reverse logic means it must be secure. Despite that, 2 person-months to find an escape. What can we conclude? They actually are incompetent at security because they did poorly secure high value targets, and that entire train of logic is just wishful thinking.
Crowdstrike must have good deployment practices because it would be catastrophic if they, like, I dunno, mass pushed a broken patch and bricked millions of machines, and only an incompetent would use poor deployment practices on such a critical system, therefore they must have good deployment practices. Turns out, no, people are incompetent all the time. The criticality of systems is almost entirely divorced from those systems actually being treated critically unless you have good processes which is emphatically and empirically not the case in commercial IT software as a whole, let alone commercial IT software security.
That quote further illustrates how, despite how easy such an attack was to develop, no in-the-wild exploits were observed. Therefore, the presence of absence of known vulnerabilities and "implicit 7-figure bounty"s is no indication that exploits are hard to develop. The entire notion of some sort of bizarre ambient, osmotic proof of security just wrong-headed. You need actual, direct audits, with no discovered exploits to establish concrete evidence for a level of security. If you put a team with a budget of 10 M$ on it and they find 10 vulnerabilities, you can be fairly confidence that the development processes can not weed out vulnerabilities that require 10 M$, or possibly even 1 M$ effort to identify. You need repeated competent teams to fail to find anything at a level of effort to establish any sense of a lower bound.
Actually, now that I am looking at that post, it says:
"Even though KVM’s kernel attack surface is significantly smaller than the one exposed by a default QEMU configuration or similar user space VMMs, a KVM vulnerability has advantages that make it very valuable for an attacker:
...
Due to the somewhat poor security history of QEMU, new user space VMMs like crosvm or Firecracker are written in Rust, a memory safe language. Of course, there can still be non-memory safety vulnerabilities or problems due to incorrect or buggy usage of the KVM APIs, but using Rust effectively prevents the large majority of bugs that were discovered in C-based user space VMMs in the past.
Finally, a pure KVM exploit can work against targets that use proprietary or heavily modified user space VMMs. While the big cloud providers do not go into much detail about their virtualization stacks publicly, it is safe to assume that they do not depend on an unmodified QEMU version for their production workloads. In contrast, KVM’s smaller code base makes heavy modifications unlikely (and KVM’s contributor list points at a strong tendency to upstream such modifications when they exist)."
So, this post already post-dates the key technologies tptacek mentioned that supposedly made modern hypervisors so "secure" such as: "everything uses the same small KVM interface", "Maximalist C/C+ hypervisors have been replaced with lightweight virtualization, which codebases are generally written in memory-safe Rust".
KVM, Rust for the VMM, despite that one person on the Google Project Zero team invalidated the security in 2 months. Goes to show how effective and secure it actually was after those vaunted improvements and how my prediction that it would be easily broken despite such changes was correct, where as tptacek got it wrong.
[1] https://googleprojectzero.blogspot.com/2021/06/an-epyc-escap...
But he's right. And with the endless stream of leaky CPUs and memory (spectre, rowhammer, etc) he's even more right now than he was 17 years ago.
There are all kinds of things being done to mitigate multi-tenant security risks in the Confidential Computing space (with Trusted Execution Environments, Homomorphic Encryption, or even Secure Multiparty Computation), but these are all incredibly complex and largely bolted on to an insecure base.
It's just really, *really*, hard to make something non-trivial fully secure. "It depends on your threat model" used to be a valid statement, but with everyone running all of their code on top of basically 3 platforms owned by megacorps, I'm not sure even that is true anymore.
When did it become customary to defend people making claims of security instead of laughing in their face even though history shows them such claims to be a endless clown parade?
How about you present the extraordinary evidence needed to support the extraordinary claim that there are no vulnerabilities? I will accept simple forms of proof such as a formal proof of correctness or a unclaimed 10 M$ bug bounty that has never been claimed.
The history of KVM and hardware virtualization is not an endless clown parade.
Find a vulnerability researcher to talk to about OpenBSD sometime, though.
Notice that at no point does anyone actually show up with a working exploit.
I mean, jeez, even Joanna Rutkowska acknowledges the foundations are iffy enough to only justify claiming “reasonably secure” for Qubes OS.
You are making a extraordinary claim of security which stands diametrically opposed to the consensus that things are easily hacked. You need to present extraordinary evidence to support such a claim. You can see my other reply for what I would consider minimal criteria for evidence.
You have not even established what level of security you are arguing has been achieved. This is not even moving the goalposts, this is Calvinball.
I contend that a major cloud service, that runs trillions of dollars of commerce is at least as important as a fighter jet. The F-35 demanded a operating system certified according to the SKPP which follows in the heels of the Orange Book Level A1. That demanded a formal specification, formal proofs, and a failed penetration test by the NSA.
Do you contend that KVM has reached such a standard? Or do you argue that such a standard is too high? What standard should be expected? How do you verify such a standard has been achieved? How does that trace to operational security goals?
The operational security goal commerce needs is for the expected value of an attack to be unprofitable. How are you verifying your axiomatic arguments are moving that needle?
Thinking that everything is insecure and all that matters is “better” is not even binary thinking, it is unary thinking. There is no meaningful discussion to be had until you:
1. Establish a measure and level of security that matches operational goals.
2. Demonstrate proposed mechanisms empirically achieve such goals or have a track record of achieving the desired level of quality such that the reputation may provide some coarse substitute for evidence.
Until that point it is: “Dude, trust me. I have been wrong every time before, but I totally got it this time.”
"I can not evaluate my work, I demand you do not evaluate my work, and I do not even know what my goal is, but I can tell you hot or cold."
What you have there is not engineering, it is art and is why commercial IT software security is a joke.