Firmware Updates and Initial Performance Data for Data Center Systems
newsroom.intel.com
newsroom.intel.com
well, that's expected! its _context switching_ that causes the slowdown -- it seems, one cannot trust intel's PR on meltdown/spectre issues.
further down:
> For FlexibleIO, [..] When we conducted testing to stress the CPU (100% write case), we saw an 18% decrease in throughput performance
well, that's more like it.
But in this specific case I don't really mind them publishing an "expected" result of one benchmark together with a lot of other ones. Not everybody understands the problem as well and is able to tell which use cases are affected or not. If you send them to this site and all they see are negative numbers they might not be able to tell that nevertheless there are other use cases without significant impacts.
And I find it completely fair for them to show that not everything is affected the same, and show the whole spectrum starting with zero impact.
The irony is of course that the software and hardware in use is only as powerful and useful as it is because a billion people wanted processors to take, share and look at cat pictures...
(That, and programmers with "real jobs" [see what I did there] happen to think the opposite, that cashing your government grant isn't a real job!)
http://web.mit.edu/humor/Computers/real.programmers
In this case, there would be lots of programmers in HPC field alone who want the micro and macro benchmarks to be good since they're squeezing every bit of number crunching they can out piles of machines they spent a fortune on. When I studied supercomputing, they were timing everything from latency of memory operations (esp NUMA) to context switches on CPU's to raw MIPS. The suppliers were competing on that stuff, too.
See: GPU drivers checking running process, bumping voltages and all sorts of other shenanigans.
[1] https://www.anandtech.com/show/7384/state-of-cheating-in-and...
The right way to do things is have a private test that matches a slice of your real workload. That way the vendor can't tailor their chips/drivers to it.
> Red Hat is no longer providing microcode to address Spectre, variant 2, due to instabilities introduced that are causing customer systems to not boot.
[0]: http://www.theregister.co.uk/2018/01/18/red_hat_spectre_firm...
I've blacklisted the microcode update from my systems for now. I understand this is being rushed out due to security issues, but the risk posed by complex local exploits like Spectre is substantially less than the risk posed by system lockups/reboots due to broken microcode. It appears that even ultra-conservative Red Hat is being forced into that conclusion.
All of the mitigation stuff needs at least 4-6 more weeks in the oven before it's anything near production-ready, and in the case of Intel, probably more like 3-6 months before they have a semi-stable microcode, if ever.
Disclaimer: I say this as an outside observer with no direct knowledge.
Means neither '90% of Intel CPUs sold in the past five years' nor '90% of the Intel CPUs currently in use'.
At least they took five years and not the usual two years...
edit: And if its not turning off speculative execution, how is it addressing Spectre? Because I thought that was the only way.
What they are doing is flushing (some of?) the btb on privilege level change.
Spectre comes in two varieties, the generic branch avoidance "boundary check bypass", and the BTB poisoning one.
The solution to boundary check bypass is to just surrender and document branches as unsuitable for providing security boundaries. Going forward, "branch on out of bounds" is going to be replaced by using unconditional math to clamp access to the array boundary. In any case, this is of very marginal utility to an attacker, because it's only useful if there is some privileged information within an address space where the attacker gets to write code. Really only useful in JIT situations, and those will be quickly fixed in software.
The other half of spectre, the BTB poisoning, is much more scary, as it allows you to inject arbitrary code to an arbitrary process (or kernel!) running on the same CPU. (The limitation is that you only get to run until the branch reaches retirement, and you can only communicate with the rest of the world through cache timing.) This one will be hotfixed by retpolines in software, then fixed by ucode changes that provide options to flush the BTB, and in the long term fixed in hardware by tagging BTB entries better.
Yeah, BTB tagging is probably a lot better than flushing performance wise.
For example, if the code of a JS array does p%arraylength before using p, it makes the spectre 1 vulnerability impossible to exploit. Browsers with JIT engines are also very quickly patched software, and afaict all the major browsers have fast-tracked changes to prevent spectre 1. At this point, spectre 1 is no longer a major threat.
In any case, many people are misinformed that disabling speculation is a viable fix. It really isn't -- completely disabling speculation means that every branch has a cost of ~20 cycles and serializes execution around it. Current normal x86 code executes a branch every 5-10 instructions (generally, more when using dynamic languages, less when using static compiled-to-metal languages). Executing branches so often doesn't ruin performance because branch predictor hit rates are >95%, as most of those branches are basically guards, type checks and the like which are almost never taken. Disabling speculation would make modern high-end CPUs spend the vast majority of their time just waiting for the branches to resolve.
There is no, and can be no hardware fix to this. The only solution is just to accept that you cannot use a branch alone to protect secret data.
What's being measured here must be mainly the impact of the Meltdown fixes.
This is why benchmarks which use more heavily the operating system are affected the most, while benchmarks which stay in user mode doing computations are affected the least.
Anyway, Meltdown is only fixable by an os update (that software patch which causes the massive 5-20%).
The microcode updates give os developers a few extra tools that allows them to build Spectre migrations, like temporarily disabling indirect branch detection while kernel code executes, or flushing the indirect branch entries on switch to kernel mode.
Yes, currently the OS update with the performance degradation is all we have. But could there be other future solutions that work differently and thus have lower performance impacts?
I think there could be. AMD is not affected, so it is not at all impossible to have a CPU behave "correctly". Wether Intel is able to correct their behavior only in microcode is of course a different question, that I'm not really able to judge.
But it could still be possible for them to add special CPU instructions that allow the kernel to explicitly protect it's address space and go back to the previous memory mapping.
I'm not super hopeful since they already had a lot of time to look into that and did not come out or announce such a solution, but maybe they deferred that in the light that KPTI works and is "good enough" for a first mitigation.
Everyone else will switch to the same model.
I suspect that cache changes will need to be tagged with the reorder buffer slot and rolled-back on mis-predict. It also means a hit to the N-way scheme because you must be able to hold multiple instances of the cache line for the same address.
I also worry there are undiscovered side channels lurking in arch-specific registers or status bits.
Instead, hold the newly loaded cache lines in a "cache line buffer", the same way how stores are held in a store buffer. For any reads, the CPU will check the cache load buffers before L1 cache.
Then once the instruction which triggered the cache read completes, the new cache line will finally be applied to the L1 cache and the old line evicted.
In the case of a misprediction, the cache load buffers can be discarded instantly.
In the future, we might see OS optimisations that work around the slowdowns by doing even less syscalls, but KPTI will stay until meltdown is fixed in silicon (which won't happen for 2-4 years)
Granted, I don't know why they couldn't just put it in HTML.
At the very least they should change the link text and/or add some alt text on what exactly they will be clicking on
This is an attack that lets an attacker read all of memory from user space. Maybe even from Javascript in the browser. Remember, serious attackers don't want to take over your computer and send spam. They want your data.
[1] https://www.bloomberg.com/news/features/2018-01-18/intel-has...
Edit: I'm not sure this is right. RHEL/Centos kernel 3.10.0-693 is vulnerable but 3.10.0-693.11.6 is patched.
> Over the past several days, Intel has made further progress to address the exploits known as “Spectre” and “Meltdown.”
Then it goes on to say:
> Generally speaking, the workloads that incorporate a larger number of user/kernel privilege changes and spend a significant amount of time in privileged mode will be more adversely impacted.
All of this implies that they're testing a fixed kernel.
There are also Spectre fixes landing in kernels. E.g. Linux 4.14.14 added initial retpoline support:
https://lwn.net/Articles/744621/
The current LWN has very good coverage on the latest work on Spectre/Meltdown mitigation in the kernel:
https://lwn.net/SubscriberLink/744287/d868ef1ac3f68d70/
(Posting a subscriber link in good faith. If you like such content, please subscribe to LWN.net, they are excellent!)
Which has its own performance drawbacks, but the microcode update itself has even more. And you need the microcode update for Broadwell and newer for retpolines to work.
It does go on to say they have been able to reproduce the issue and are making progress towards finding the cause.
Let's say it like it is.
Random reboots should really never happen and the fact that Intel is trying to imply otherwise is deeply worrying.
We all know that things go wrong. The problem is rarely that mistakes are made, but rather that people aren't open about them and don't simply provide concise technical analyses.
It's just embarrassing.
* Domain registered only 8 days after meltdownattack.com yet is "based on the work highlighted by Meltdown and Spectre". Hardly seems like enough time to have come up with something significant enough to give a name to. Goes out of its way to copy the font used by meltdownattack.com and advertises itself with the names of meltdown and spectre, and their CVE IDs without listing its own. Given what they said it should have its own CVE IDs reserved by now. Just looks like a cheap grab for attention as it is.
* Unlike meltdown: Where are the mysterious Linux patches being speculated about if it's going to be announced when "operating system vendors have prepared patches." Is it so early that noone's begun work on it? Did noone invite Linux to the party?
* If it's actually important enough to be under "embargo", why are they hinting details on a public website about it at all?
* Its current icon[1] is a really cheaply made recolour of the Intel logo. Worst of any "hip and cool vulnerabilities with a name, logo and a website" yet, if real. Seems like the kind of thing I'd expect someone who doesn't understand meltdown/spectre to create because they saw people shitting on Intel, and definitely not the creation of someone who is supposedly working with chip manufacturers and following "embargos".
* Also apparently there's a second icon[2] based on the solaris logo. If one vulnerability is intel-related (i.e. a general purpose attack) and the other is solaris-related (i.e. a specific attack on solaris), why would they be bundled together? It's either inconsistent or the logos have nothing to do with the vulnerabilities which would make even less sense.
Please disable Intel boot guard for coreboot and work with the open source community.
That way we will have more secure systems.