A Look at the CPU Security Mitigation Costs Three Years After Spectre/Meltdown
phoronix.com
phoronix.com
Vulnerability Spectre v1: Mitigation; usercopy/swapgs barriers and __user pointer sanitization
Vulnerability Spectre v2: Mitigation; Full AMD retpoline, STIBP disabled, RSB filling
By having those it will affect my CPU perferomance? Will applying "mitigations=off" help with perfermance? Also by having "mitigations=off" what are the real security implications? Thank you for your help!The security risk part is hard for most people to understand and make an informed decision. (Mitigation s ON is obviously safer, but even if I decide I want to go faster, I still don't understand the risk in taking.)
This is obviously a niche case. But IMO, I can see similar calculus for other more mainstream use cases. I wouldn't turn migrations=off on a public facing web server. But if you have a cluster that does ETL data processing, it's probably CPU-bound and has very little externally exposed surface area. Arguably, even backend SQL servers can benefit, assuming the set of clients that directly access the dbase are tightly controlled and audited.
When every cycle counts, wasting them in the kernel is not a good idea anyhow. So, isolcpus and kernel bypass FTW.
(EDIT) Honestly I did not notice any performance issues by NOT using "mitigation=off". Default kernel 5.10 settings seem to be working well enought. At least for my tasks :-) But than again, I'm not playing any games on it LOL
If you have something that's just making a huge number of syscall's in a tight loop the mitigations absolutely dumpster performance.
There are more granular parameters that are supported by all kernels implementing CPU mitigations: noibrs noibpb nopti nospectre_v2 nospectre_v1 l1tf=off nospec_store_bypass_disable no_stf_barrier mds=off mitigations=off
https://unix.stackexchange.com/questions/554908/disable-spec...
I do lots of compiling at work and our compiler hasn't changed in 5+ years.
pre spectre my 8700k could do 20k lines/sec. Post spectre with mitigation it's about 10k, with them disabled it's about 15k.
There's clearly been some under the hood changes to windows beyond these optional? changes.
The more difficult/interesting problem is being future-proof to whatever else undiscovered. Many Spectre attacks are heavily timing based, and even a single cycle variation in pipeline stages or flushing structures will spawn a new variant (see, MDS vs. Spectre). This is actually partially why a current trend in hardware research is trending towards fuzzing-type stuff [1].
Something also worth noting is that it's incredibly difficult to quantify "leakage". There's a pretty big difference between vulnerable and exploitable-- e.g. original Spectre papers had 10KB/s of kernel dumping, which is a big reason it was scary, but would it really be a big deal if it had <1b/s? Not going to explicitly name and shame, but there've been a handful of reasonable high profile "vulns" with cute domain names that I'm shocked to even see accepted at conferences due to how contrived the exploit was and tiny their leakage rates were.
I personally don't really know how to address the quantification problem, but I very much think it's necessary in any discussion of a bug's impact/severity. Definitely gets exhausting when every cute name gets a headline, and it's easy to blow things out of proportion without some grounding in reality.
* NOTE: I do think script kiddies have their place still. Definitely important to have automation to determine whether a system has updated security [2], it's just that I doubt such a scenario is applicable to Apple.
[1] https://www.usenix.org/conference/usenixsecurity20/presentat...
[2] https://owasp.org/www-project-top-ten/2017/A6_2017-Security_...
Oh, do tell how spectre is solved. I would love to know, defending against it is a real problem I currently have. If you allow untrusted code to execute on your machine (e.g. JavaScript) then you're vulnerable to it. There are no practical attacks in the wild that I'm aware of, and it's tricky to do, but it's not impossible and the only defense really is to make it harder and more time consuming. This is the approach that I and others have taken.
Most of those CPUs are designed by CS or EE students taking a computer architecture class, so... in a sense... one can argue that defending against Spectre-like attacks is actually super simple: just use a simple CPU design.
To actually become vulnerable to Spectre, you need a very complex CPU design, so in the same way, it can be argue, that making a CPU vulnerable to Spectre is actually hard, since it takes a lot of work to create such a CPU design.
Now, if what you want is a CPU that's both fast and secure, then I'm sorry to tell you that such thing cannot exist. Those two goals are at tension. You can either get a F1 or a tank, but no vehicle that offers the same amount of protection as a tank is going to be able to compete against a F1 car, and vice-versa.
I guess if you remove the branch predictor, you might avoid spectre while keeping caches, but I think you can keep the branch predictor and remove caches to also avoid spectre.
The downside I see in keeping the caches is that you keep the _source_ of the timing differences, so an attacker just needs to find a different attack vector to create a new timing attack.
If you remove the caches, you kill the source of most timing attacks.
I'm not an expert on this though.
If it works, flush your caches or just update your kernel. If it doesn't, you checked off an item on your list.
I misspoke, silicon should be more like "system". Said it in this comment more about silicon fixes https://news.ycombinator.com/item?id=25665276
[1] https://gist.github.com/anonymous/99a72c9c1003f8ae0707b4927e...
Also your references to flushing caches and updating the kernel underlines to me that you don't know what you're talking about.
>> The more difficult/interesting problem is being future-proof
Obviously read the source first but there are a few implementations of the basic ones on GitHub, not sure about the higher versions of spectre (I don't know whether CVE's require an implementation to be published).
If I can find I'll link a paper that describes a system to try to automatically characterize these types of side channels.
It is a memory leakage problem and in the world of today with apps stealing info left and right, I am almost unsure if I should care about it that much. Maybe attackers might be able to steal a key here and there, but I managed to stay quite cool when the architectural flaws were published.
I believe I saw a demonstration about the M1 not being affected by these side channel attacks, but I have no source.
But I don't think anybody managed to, due to the amount of noise the browsers sandboxes add into the necessary syscalls.
Now the question would be are those attacks still viable given the additional hardening browsers have done independent of the kernel mitigations?
Although many pages do not work with it, so I disable it quite often
I wonder if it is possible to activate the mitigations only for the browser?
Or only for one user? I have created a separate user for the browser, so it cannot change my actual files
I'd have preferred a hardware flag that says "I accept that if untrusted code runs on this CPU, I've already lost" and keep the performance boosts of speculation, but alas.
1) It takes a long time for desired design changes to make it to fabs, so it was always going to take years for the classes of vuln that were discovered to get mitigated in hardware.
2) More have been discovered in the intervening period.
3) Speculative execution really speeds things up, so the fixes are removing/changing as little spec. ex as you think you can get away with, and leaving the rest.
CPUs are extremely expensive to design and verify, and astronomically expensive to make - these things take time.
Given common high perf microarchitectures, that's quasi-impossible. You can use somewhat efficient mitigations though.
Then again, if the solution was simple, I'm sure it would have already been implemented.
The best system design change IMO would be to tag at programming language level (to systematize the introduction of the needed barriers, without needing too many of them, thus minimizing the overhead -- which is a must otherwise you can as well remove OOO entirely), but I suspect that's not going to happen for most systems. Or even any of them.
For Spectre, you can't really have a high performance chip that isn't vulnerable to it in theory. A processor can't read your mind and know which code within a process should be able to communicate with which data within the process unless you tell it. You could just go to strict in order but that's taking a huge performance hit. Programmers cominging secure information and JITs running untrusted code within the same process will have to work to make sure they're not vulnerable to Spectre going forward.
You could if you had a fully transactional cache such that branch misses could be rolled back. It's highly non-trivial and likely far too expensive in die space to justify, but theoretically solvable.
But I think realistically CPUs are just going to say "processes are the only security boundary we offer" and leave it at that. Which puts things like WASM in a very questionable spot (particularly things like WASM in the kernel), but Intel & AMD probably don't care about that too much. For web browsers all the major ones just gave up on in-process sandboxing and that's why we have things like per-iframe process sandboxes now ( and then reduced privileges in-process iframes with the 'sandbox' attribute: https://caniuse.com/?search=sandbox )
I do understand the need for top security in enterprise applications. But for personal use, all of this seems a bit ridiculous.
I've disabled mitigations on every computer in my home, will enable if I get burned. ~10-20% performance difference - I consider it a decent tradeoff.
But then again, I don't stress much about being hacked or privacy. It's not worth the mental effort, plus I'm a nobody and I accept that.
The literature is still very thin too, so we're still very much in the early days of these side channels.
Just look at the mitigations list of i9 10900k or Ryzen 9 5950X as used in the article, both released in latter half of 2020.
>> itlb_multihit: Not affected + l1tf: Not affected + mds: Not affected + meltdown: Not affected + spec_store_bypass: Mitigation of SSB disabled via prctl and seccomp + spectre_v1: Mitigation of usercopy/swapgs barriers and __user pointer sanitization + spectre_v2: Mitigation of Full AMD retpoline IBPB: conditional IBRS_FW STIBP: always-on RSB filling + srbds: Not affected + tsx_async_abort: Not affected
Many of these are still by and large "disable this optimization" and "barriers", all done in software, with the potential exception of EIBRS which just essentially tags the branch predictors (which, in my opinion, they don't do in a particularly effective way, but stay tuned for Research tm). Also Meltdown, which I think is a pretty easy fix, and even Intel agreed [1].
Given typical design --> market time might be ~2 years, designers only had a few months to 1) finish currently pipelined work and 2) work on mitigations. On top of that, given that this is a pretty hard problem, a few months (not even > 1 year IMO) is definitely not enough time.
I'd guess that this problem won't be "solved" in a sense for many many more years, much in the same way that the rich history of buffer overflows has been a cat and mouse game for decades, with each mitigation coming with its own tradeoffs and potential performance penalties. On the "Hardware Mitigations" side, think of ARM's pointer authentication-- it's only just now that some pretty nice hardware support has been spawned, decades after buffer overflows were "known".
Still though, I personally think some more silicon mitigations will come in the next 2 years. Part of the reason "Pointer Auth" came "decades after buffer overflows" is due to the rise of cloud compute, and security matters much much more to the general public these days. I just mentioned current research generally has penalties on the order of 10%-- While pretty distasteful, it can definitely be (and has been) gradually improved on and is a far step from initial estimates of ~33-50%.
[1] https://www.rockpapershotgun.com/2018/01/29/intel-cannon-lak...
With the death of Moore's law [1], it's not too surprising that newer CPUs (or at least x86 ones) are still recoiling from Spectre and Meltdown et al.
[a]: Obviously, for absolute security, a hammer and/or a shredder does much better.
[b]: This also ignores that the Gutmann method was designed for hard drive encoding methods that aren’t used anymore.
This still leaves large pieces of magnetic plates intact. At uni we could read some data from floppies that went through shredder.
Hammer and shredder is NOT "absolute security".
https://www.phoronix.com/scan.php?page=news_item&px=Spectre-...
I mean, imagine coming upon a system where you want to add some other kernel option. You see that the system has all those options already added to the default kernel options. Are you going to research and clean up those options and remove the irrelevant ones? Or are you going to punt, add your own option to the end, thereby adding to the chaos?
Or do you mean "please, not in a production system where you share a host"?
If that's the case, I think you have it backwards. In production, please, don't share a host.
Turning off the mitigations in the kernel sure speeds things up a lot on older machines, but if browsers don't do a proper job of mitigating those attacks then someone could extract data through JavaScript.
Only because this list happened to pick an expensive chip as the representative model for Ryzen 5000. For an even competition, look at the numbers for the 5800X and 5900X, which are respectively $50 cheaper and $50 more expensive. Compared to those, the 10900K isn't impressive. https://openbenchmarking.org/embed.php?i=2011098-FI-AMDRYZEN...
And the reason people got excited is not because Intel parts became useless, it's because AMD finally managed a generation of chips that flat-out beat Intel's, even in single core performance. Especially in light of how important a role AMD played in breaking the status quo of quad cores, putting fear in Intel's eyes is great for competition.
As far as stuff you can buy today there's Apple way out in front (if you can live with their RAM configs) followed distantly by last year's i9, then there's AMD.
> nobody owns one.
https://i.imgur.com/8EMeiCZ.png
This is only one source, but the 10900 numbers since launch are dwarfed by the 5800X numbers since launch. For the people building PCs with this site, Zen 3 chips are currently outselling all of Intel combined.
> Using the benchmarks in this article, the Ryzen 5950X is 2% faster.
I think this article is all single core stuff. Which is the aspect they "finally managed" to win at. It's not their strength at all, it's the thing Intel was able to lord over them. Should I have said "barely" in addition to "finally managed"? I thought it was clear enough.
It took me a while of waiting but I finally got mine a couple of weeks ago. It is pretty nice.
Since the chips sell out in the same day they arrive at the store you won't ever see them "in stock" but obviously, out of all the chip models at least 100 new PC enthusiasts every week own one just from that one store.
I also know several friends of mine who finally received their prebuilt gaming systems with Ryzen 5800 or 5900 chips and Nvidia 3080's. There was a lot of delay but the OEMs are shipping a lot of boxes.
As for the i9-10900K, if you check out a review from launch [0], you'll see nothing but good things to be said about the performance, with a very large trade-off of being a ridiculously power hungry, heat-producing CPU. If you're OK with that, then you're probably perfectly happy with the Intel CPU. For me, a cool and quiet CPU that can still smash my parallel processing needs and meet my gaming needs is a better sweet spot.
[0] https://www.tomshardware.com/reviews/intel-core-i9-10900k-cp...
If you're doing media encoding, it would be hard to measure. If you're doing an haproxy tcp proxy, it's big. If you're doing https, it's probably not so big (encryption eats enough cycles that you wouldn't necessarily see the slowdown).
Considering the millions of chips running worldwide and the CPU cycles / energy wasted on this, this really is absolutely horrendous.
cat /proc/cmdline