I know that bugs happen and that there was nothing intentional on this one, but at times like this is hard to held at bay the temptation of claiming for a class lawsuit against Intel....
I know that bugs happen and that there was nothing intentional on this one, but at times like this is hard to held at bay the temptation of claiming for a class lawsuit against Intel....
It's, however, really bad if you sell CPU cycles for a living. You just lost between 5 and 30% of your capacity. If you have a large building, you just lost part of your parking lot to the Intel Kernel Page Problem building.
Honestly, I'd just make sure the server firewalls are super tight and not take in the future patches. At least for now.
I know I’m not alone.
Then again. Think of microservices, Kubernetes for instance; Network requests are system calls.
(In my case anyway)
30% overhead might be inscentive to revisit the assumption we can’t rewrite it for Linux.
I don’t know very much about computing on that scale, but I wonder if all the people selling off Intel stock are thinking this story through.
It's possible that the patches applied to fix this bug will cause some single-threaded benchmarks to change from Intel being the fastest to AMD being the fastest.
Not trying to kill expectations. This decision isn’t mine alone. You know the old saying “nobody got fired for buying Cisco” that applies to Intel too.
That's a good description of basically every cloud environment out there, from AWS on down.
In other words they are extremely common.
We'll start to get conscious about the number of syscalls we use on each operation, start using large buffers, start buffering stuff user-side...
Good security is about layers. No one layer can be assumed to be watertight, but with enough layers you hopefully get to a good place.
What do you mean by "compressible"?
Kernel ABIs will eventually reflect that and crop up higher level expensive calls that replace groups of currently cheap syscalls (that will become expensive after the fix).
And Intel will profit handsomely from next generation CPUs that'll get an instant up-to-30% performance boost for fixing this bug.
Who really sells CPU cycles? Cloud providers sell instances priced per core. So the real hit is by the customers since they have to shell out for more instances for the same amount of computing power.
The hit I see is by providers of 'serverless' computing, since they charge per request and have their margins reduced.
AWS, Azure, and GCP all bill serverless with a combination of per-request fees and compute (GB-seconds), so I'd expect the entire hit to be passed on to the user since this will cause increased compute time for each request. N requests that used to average 300ms each will now be N requests that average, say, 400ms, so the per-request billing remains the same and the compute billing will increase by approximately 30%.
30% is a big hit. I'm wondering if that isn't a bit exaggerated, or perhaps the consequence of a poorly optimized workarounds that will rapidly improve. I recall seeing figures on the order of 3% only a few days ago.
How big it will be for your workload is a function of what your workload is. Benchmark if it is important to you.
Also a 30% decrease is also equivalent to setting Moore's law back 7 months. A 5% loss is only setting it back 1 month. I know that's a bit of a naive calculation. But the point is computing power has long operated in an exponential domain. So big differences in absolute numbers aren't necessarily a big deal.
Maybe a lot more now?
Well... Everyone who bought AMD. Some people managed to see beyond the hype and go for the optoon that made sense.
What hype are you referring to? Are you suggesting the people who bought AMD knew this was a problem for Intel?
True, sometimes you will leave boxes at low utilisation for various reasons, e.g. to deal with traffic spikes. But those reasons have not gone away. So now instead of heaving a predictable increase in CPU cost, you have an unpredictable increase in performance snafus.
The only good news is that the real performance hit will be less than 30% on many workloads. Especially once the providers start juggling and optimising.
TL:DR - queue behaviour gets nonlinear as you approach the theoretical max load. If you are running your processors at a high load, even a small change in code throughput makes a huge difference to real world behaviour.
Not to be mean, but that's not what is being changed.
You're right on the bug - userlevel code can now read any memory regardless of privilege level. However the fix isn't to manually check the privileges on each access - that would be extremely slow and wouldn't actually fix the problem.
The fix is to unmap the kernel entirely when userspace code is running. Because the kernel will no longer be in the page-table, the userspace code can no longer read it. The side-effect of this is that the page-table now needs to be switched every-time you enter the kernel, which also flushes the TLB and means that there will be a lot more TLB misses when executing code, which slows things down a lot.
So, to be clear, it is not accessing pages that is being slowed down, it is the switch from the kernelspace to the userspace.
All your points are right though. Page access times will in general be slower because of all the extra TLB flushes, leading to more TLB misses when accessing memory.
And how often the kernel services interrupts.
Or am I completely off the mark?
If people who received written assurance from Intel that their hardware is 100% bug free can form a legal class, sure. I highly doubt there is even a single one such customer.
Yes and no. Yes, Intel would get a chance to claim that the case should be dismissed out of hand. To do that, they have to prove that, even assuming all the claimed facts are true, the people suing still don't have a valid case. That's a high bar. It can be reached - there's a reason that preliminary summary judgment is a thing in court cases - but it takes a really flawed case to be dismissed in this way.
How flawed? SCO v. IBM was not completely dismissed on preliminary summary judgment, and that was the most flawed case I've ever seen.
> It may very well be a question of who has the better legal team.
Well, Intel can afford to hire the best. A huge class-action suit can sometimes attract the best to the other side as well, though. (There's not just one "best", so there's enough for both sides of the same court case.)
IANAL, but it looks to me like there's at least the potential for a valid court case. CPUs are (approximately) priced according to their ability to handle workloads; if they can't provide the advertised performance, they didn't deserve the price they sold for.
What I meant was that the presence of the bug itself is not a valid cause, for example you can't claim that due to the error you lost 1 trillion dollars via a software hack - even if it's true. If Intel can prove they acted ethically when disclosing the bug and that they replaced / compensated users up to the value of the CPU, they are in the clear.
The question is did anyone receive performance assurance from Intel? Probably not.
Some cloud providers or compute grids just lost a lot. Maybe they will find an angle to claim compensation.
Agreed!
(We should probably also stop overgeneralizing about the nature of computational workloads.)
Incorrect. It also affects interrupts and (page) faults.
Any usermode to kernel and back transition.
Hosting on bare metal will become more attractive. Too bad you can't long OVH and Hetzner.
What does that even mean?
Also Hetzner just introduced some AMD Epyc server.
As opposed to "shorting" a stock, which means making a bet that it will go down in value.
HN doesn't let you do this to new comments to avoid back-and-forth commenting that is typical in flamewars.
You can reply anyway, but you have to click on the timestamp ("X minutes ago") to do it.
The other benchmark that has generated some consternation is running 'du' on a nonstop loop.
Both of these situations are pathological cases and don't reflect real-world performance. My guess is a 5-10% performance hit on general workloads. Still significant, but nowhere near as bad as some of the numbers that are getting thrown around.
And, databases are the worst case scenario, most real-world applications are showing 1% performance impact or less.
https://www.computerbase.de/2018-01/intel-cpu-pti-sicherheit...
https://www.hardwareluxx.de/index.php/news/hardware/prozesso...
Your last link is all gaming benchmarks, which as the article mentions are not affected much.
[1] http://lkml.iu.edu/hypermail/linux/kernel/1801.0/01274.html [2] http://lkml.iu.edu/hypermail/linux/kernel/1801.0/01299.html
This isn’t an excuse for Intel consistently having terrible verification practices and shipping horrendous hardware bugs. From 2015: https://danluu.com/cpu-bugs/ There have been more since then.
I’ve talked to multiple people who work in intel’s testing division and think “verification” means “unit tests”. The complexity of their CPUs has far surpassed what they know how to manage.
Found a quote:
"We need to move faster. Validation at Intel is taking much longer than it does for our competition. We need to do whatever we can to reduce those times… we can’t live forever in the shadow of the early 90’s FDIV bug, we need to move on. Our competition is moving much faster than we are".
Overall it’s a depressing story of predictable market failure as well as internal misbehavior at Intel, if true. Few buyers want to pay or wait for correctness until a sufficiently bad bug is sufficiently fresh in human memory. And if you do want to, it’s not as if you’re blessed with many convenient alternatives.
The same reason could have been used to give the NSA some legroom for instance, but tell everyone that's why they won't do so much verification in the future.
As other comments suggest, there might be a third stage, completely forgetting how to design and validate chips properly.
Furthermore, I just a read an article (can't find the link) that certain ARM Cortex cores have this same issues as Intel.
More likely "good enough" is much lower because ARM users aren't finding the bugs. The workloads that find these bugs in Intel systems are: heavy compilation, heavy numeric computation, privilege escalation attackers on multi-user systems. Those use cases barely exist on ARM: who's running a compile farm on ARM, or doing scientific computation on an ARM cluster, or offering a public cloud running on ARM?
Vendor, in conversation: "We're pretty sure we can make the next version do cache coherency correctly."
Me (paraphrased): "Don't let the door hit you in the ass on the way out."
Management chain chooses them anyway, I spend the next year chasing down cache-related bugs. Fun.
(I should remark that there are good reasons for this effort. Such as: It boots in under 500ms, it's crazy efficient, doesn't use much RAM, and your company won't let you use anything with a GPL license for reasons that the lawyers are adamant about).
So now you get to find all the places where the vendor documentation, sample code and so forth is wrong, or missing entirely, or telling the truth but about a different SOC. You find the race conditions, the timing problems, the magic tuning parameters that make things like the memory controller and the USB system actually work, the places where the cache system doesn't play well with various DMA controllers, the DMA engines that run wild and stomp memory at random, the I2C interfaces that randomly freeze or corrupt data . . . I could go on.
It's fun, but nothing you learn is very transferrable (with the possible exception of mistrust of people at big silicon houses who slap together SOCs).
There are hardware manufacturers that are better than others at being open and providing documentation. My minimal level of required support and documentation right now is mainline linux support.
Can you document your work publicly, or is there something I can read about it? I'm very interested in alternative kernels beside Linux.
When you buy an SOC, the /contract/ you have with the chip company determines the extent and depth of their responsibility. On the other hand, they do want to sell chips to you, hopefully lots of them, so it's not like they're going to make life difficult.
Some vendors are great at support. They ship you errata without you needing to ask, they are good at fielding questions, they have good quality sample code.
Other vendors will put even large customers on a tier-1 support by default, where your engineers have to deal with crappy filtering and answer inane questions over a period of days before getting any technical engagement. Issues can drag on for months. Sometimes you need to get VPs involved, on both sides, before you can get answers.
The real fun is when you use a vendor that is actively hiding chip bugs and won't admit to issues, even when you have excellent data that exposes them. For bonus points, there are vendors that will rev chips (fixing bugs) without revving chip version identifiers: Half of the chips you have will work, half won't, and you can't tell which are which without putting them into a test setup and running code.
My favorite ARM experience was where memcpy() was broken in an RTOS for "some cases". "some cases" turned out to be when the size of the copy wasn't a multiple of the cache line size. Scary stuff.
> The AMD microarchitecture does not allow memory references, including speculative references, that access higher privileged data when running in a lesser privileged mode when that access would result in a page fault.
Out-of-order processors generally trigger exceptions when instructions are retired. Because instructions are retired in-order, that allows exceptions and interrupts to be reported in program order, which is what the programmer expects to happen. Furthermore, because memory access is a critical path, the TLB/privilege check is generally started in parallel with the cache/memory access. In such an architecture, it seems like the straightforward thing to do is to let the improper access to kernel memory execute, and then raise the page fault only when the instruction retires.
If it, like it seems, is just an attack on OS kernels and PV hypervisors, you can simply turn off the mitigation, since nowadays kernel security is mostly useless (and Linux is likely full of exploitable bugs anyway, so memory protection doesn't really do that much other that protecting against accidental crashes, which isn't changed by this).
Even if it's an attack against hypervisors any large deployment can simply use reserved machines and it won't have a significant cost.
Well, if I rent a VPS with x performance, I still expect x performance after this flaw is patched. The company providing the virtual machine will perhaps have to pay 30% more to provide me with the same product I've been getting.
Since most VPS offerings arbitrage shared resources, this will not increase costs of providing VPSes by the full performance penalty.
So you may suddenly find that your own performance requirements, that were previously satisfied by "2x m5.xlarge" are no longer being met by that configuration, and I doubt AWS will just provide you with more resources at no additional charge.
Are there any providers that state you will get x performance? Most that I've seen say you will m processors, n memory, and p storage but don't make any guarantees about how well those things will perform.
On the other hand, shrinking Intel's market share due to bad PR and thus adding some competition into the industry could actually foster that progress.
The bigger issue is for things that don't scale easily. That sql server that was at 90% capacity is suddenly unable to handle the load. Sure that could've happened organically, but now it happens (perhaps literally) overnight for everyone all at once.
Expect a bunch of outages in the next few weeks as companies scramble to fix this.
It seems you need root or physical access to the system as a prerequisite for the attack.
Where that gets tricky is when everyone's using cloud hosting solutions where the physical machines are abstracted away, and a given physical server may be running multiple virtual servers for different customers.
Think of it like this:
* Somewhere in a data center at a cloud provider is a physical server, wired up in a rack..
* That server runs virtualization software, allowing it to host Virtual Server 1, Virtual Server 2, and Virtual Server 3.
* Virtual Server 1 belongs to Customer A. Virtual Servers 2 and 3 belong to Customer B.
* Normally, Virtual Server 1 can't access any memory allocated to Virtual Servers 2 and 3.
* BUT: Customer A can now use Meltdown to read the entire memory of the physical server. Which includes all the memory space of Virtual Servers 2 and 3, exposing Customer B's data to Customer A.
That's the threat here.
Just wanna point out that a 30% performance hit means a 43% cost increase.
For those confused: the math here is a 30% decrease puts you at 70%. To go from 70% back to 100%, 30% only gets you to 91% (0.70*1.3). 1/0.7 = 1.43 means you need 43% to recover.
Our ElasticSearch nodes all had 32GB of ram and we had 10 of them and they were all being pushed to the max.
Something like this would be a massive hit, requiring a lot more work into identifying new bottlenecks and scaling up appropriately.