Meltdown, aka “Dear Intel, you suck”
marc.info
marc.info
I mean I get it: large corporations (1) don't necessarily have my interests at heart; (2) are not able to perfectly execute (on extremely large and complicated) products and systems.
um. This is as radical as taking a stand that the sun sets in the west.
I like HN because of the promise that people think just a little bit before just typing something that strokes their feelz-good neurons. Yet, here we are on the front page with "In tel suxxx!" (and "Apple suxxx! in the htop thread, etc., etc. etc.)
sigh I guess it's inevitable.
People just don't like paying for intangible things, so the only way to pay for the internet is clickz/eyeballz. And OUTRGE! clickz/eyeballz are the cheapest and easiest.
I know, I know. There was a bug from vendor X that really irked/irks you and it feelz good to express your outrage.
I just wish there was just some corner of the internet where this wasn't the dominant trend and I had hoped it was here on HN. Well, at least reddit has extremely cute dog and kitten videos, and some chuckle-memes.
You're not going to find that, because taking the path of least intellectual and emotional resistance is just human nature. There is a positive feedback loop with outrage narratives that feeds people's ego and sense of in-group superiority, where one either feels empowered by agreement, or feels catharsis through disagreement.
The only thing HN has over Reddit in that regard is a culture that puts up the pretense ... sometimes.
I agree that both of those appear to be fundamental elements of human psychology. Fortunately, we're also have the ability to counteract this, either individually via self-reflection, or more easily, through feedback from others. We can be aware of these tendencies and work to counteract them. If the community has a common goal to do so, to encourage the "better angels of our nature" as it were, we can do better together than we may find it easy to do on our own. And that doesn't have to be a pretense. It can be an explicit, earnest value.
It didn't really trigger my outrage button, and did inform me of something that had not actually crossed my mind--who got embargo information or not.
So, with a lot on HNers interested in open source, and using open source software, I see value in raising awareness that we should let these vendors know we value this. Some of us are making purchasing decisions at our companies and can steer our dollars toward good corporate citizens.
I find the outrage rather incidental to the post, regardless of even the email list author's intent in that regard.
IMO, it's a good thing that this issue gets publicity... if only so that the 'lesser' open source projects don't get ignored in the future.
I think it was the WPA2 handshake vulnerability.
I remember people saying that BSD would probably no longer receive information about embargoed vulnerabilities because of that.
BSD is not an operating system unless your referring to the original Berkley Unix. Are talking about FreeBSD?, OpenBSD?, NetBSD?, DragonflyBSD? Remember they are all different operating systems.
I think the outrage is justified.
...I've no idea what your ARM comment has to do with it though. Meltdown is Intel specific.
A relevant question however might be who coordinated the work between june 2017 and now for Meltdown mitigations? Project zero at google discovered it, but did they hand over the responsibility to Intel or decide who else to inform themselves?
It's an odd world where Linux is in the "big three" and the BSD counterparts are more on the fringe. It's a far cry from the early world of the 2000s.
At the end of the day, we are all humans and the chip designers screwed up in a big way. But this issue is so complex and obscure that it's been unknown for 20+ years.
Perhaps we can create a new place?
> um. This is as radical as taking a stand that the sun sets in the west.
If all you're picking up from this is the emotional component (aka "outrage"), I think you've missed the point. The larger problem now is that Intel (in this case) consciously chose to ignore serious risks for their customers and is currently doing their best PR job to downplay the impact of those risks now that they have been exposed.
The reality is that these bugs are potentially catastrophic. But maybe we are supposed to simply accept that that is part of the risk of living in this modern world.
When large corporations make serious mistakes that are predictable and that others publicly warned about, they deserve all the criticism that comes their way.
My question is why anyone who doesn't work for the PR department of one of these corporations feels the emotional need to defend them from justly deserved criticism? They didn't include all the responsible stakeholders in their mitigation efforts, that decision put the users of those OSes at risk, and so the developers of those OSes are pissed about it.
Why is that unacceptable to you? Why do you consider that to be "wrong"?
“Intel sucks” except arm cpus are affected too. “Intel sucks” except there’s a network of entities involved in disclosure. “Intel sucks” except OpenBSD previously opted out of this disclosure mechanism. I don’t see how anything will move in useful direction from this. Perhaps Intel will find a way to write statements with better feelz. That’s nice, but it’s not substantive and doesn’t deal with any of the underlying issues here.
To be fair, you look at those CPU utilization graphs, and the optimization bonus for this sacrifice of security is notable. The benefits are real, for these computation strategies. The reality is that yet again, we must sacrifice convenience in the name of security. This is why we can’t have nice things.
Ultimately, though, your pining for dignified silence and stoicism as the one true lowest common denominator is to be unrequited. Civilization among the civilized is not always a good thing. We all know that somewhere deep within the unknown recesses of some massive social planning department, circuits are being designed to lock us in, and make us pay for stepping out of line. Keep that fact in the front of your mind. Not as a goof. Not as a pacifier or placebo.
It’s all bullshit. It’s a valid gripe. It bears repeating.
The irony of course is that this was posted on a mailing list (so clicks are meaningless) by a developer of free software who has to spend his time working around these bugs that Intel caused. He has every right to be outraged and vent all he wants. He is actually doing the hard work while you watch kitten videos and chuckle-memes.
It might be argued that the risk of the exploit making it into the wild would have been higher if the BSDs were notified, on account of the sort of accidental premature disclosure that actually happened. This would appear to be self-serving if stated by the embargo insiders, but it may still be valid, especially if my guesses that a) this problem's biggest potential impact is in cloud computing, and b) there is relatively little use of BSDs in cloud computing, are accurate.
From https://www.krackattacks.com
> We notified OpenBSD of the vulnerability on 15 July 2017, before CERT/CC was involved in the coordination. Quite quickly, Theo de Raadt replied and critiqued the tentative disclosure deadline: “In the open source world, if a person writes a diff and has to sit on it for a month, that is very discouraging”. Note that I wrote and included a suggested diff for OpenBSD already, and that at the time the tentative disclosure deadline was around the end of August. As a compromise, I allowed them to silently patch the vulnerability. In hindsight this was a bad decision, since others might rediscover the vulnerability by inspecting their silent patch. To avoid this problem in the future, OpenBSD will now receive vulnerability notifications closer to the end of an embargo.
Note the date there that de Raadt was commenting on the discouragement of sitting on a fix for a month. What is the likelihood that he would be very discouraged to sit on it for six months? What if it was a three month embargo that changed to a six month embargo - when would the fix be released?
I would assume that those are questions that need to be asked prior to notifying a project.
There was another "incident" with the KRAK embargo, where OpenBSD got permission to silently patch it early and then the researcher who found it regretted giving them permission.
I think people put these two incidents together, combine it with the developers' attitudes towards embargoes and come out with: OpenBSD doesn't honour embargoes!
It's the technological version of insider trading
'We don't respect embargoes' -> 'Company only releases to individuals who respect embargoes' -> repeat
With this recent news, I feel like part of the reason that Intel has been doing so well is that they have been _cheating_.
All processors (including AMD/ARM) since the early 90's have been using instruction pipelining. It's a universal performance improvement, not Intel cheating.
If you are open to arguments, there are many good reasons to take a negative stance towards Intel.
Fun fact: I sent an overview of unknown instructions in compressed text form to a gmail address, and Google rejected it citing potentially malicious content.
Was it ever?
The details really matter for making an exploit: so the amount of speculation and what gets speculated, how caches work, how good timing is (and how large the difference between cache and memory is) etc.
Merely knowing that the combination of speculation, caching, and timing have the potential to break through memory protection barriers is a far from enough to actually exploit that weakness.
For comparison: it was widely known that sha1 had weaknesses, yet it took many years for somebody to construct two pdfs that actually demonstrate a hash collision.
Now i know it would take a lot for a major company to delay/cancel a release like that, but i think it could be argued that they were dishonest if they knew it would have a performance impact.
If so, I hope the market punishes that decision mercilessly.
[0] Caveat: Skylake and later are extra-vulnerable to Spectre because they have even more aggressive speculative execution, but fortunately, the IBRS microcode update is roughly equally performant as retpoline on Skylake+ only.
It's really not. It's the sort of thing you wonder after first learning about out-of-order and speculative execution in a computer architecture class, but your professor assures you that implementors have been very careful to ensure any partial execution is properly flushed and rolled back. Then it turns out that, nope, no one's actually been keeping an eye on this after all.
Life-safety work has no business in a public cloud.
Anthopogenic climate change falls into this category.
I have to wonder, first, what sort of obvious stupid self-destructive things am I doing right now but ignoring. And second, how can we build systems that systematically take this into account? Is there some way to short-cut the process and find and fix problems in the early stages? Or better yet, design our systems so that they don't have the problems from the start?
There is one attitude that exacerbates those problems that are real - the tacit assumption that concerns do not count for anything until an exploit has been demonstrated.
And, at the time, your professor was correct, to the extent of what he/she defined as "any partial execution".
But his/her definition of "partial execution" only considered state changes to the programmer visible CPU architecture state (i.e., the user level register set and the flags register). Their definition ignored the cache, because at the time the cache was considered simply a transparent optimization system that did not effect the values of contents of the CPU architectural state.
And in a way, that definition is still valid, even after the knowledge of these exploits. The CPU architectural state (registers/flags), even after executing exploit code, is exactly what it would have been had sequential execution happened. The exploit takes advantage of the fact that you can arrange the right set of code to run just the right way to leave a different state in the _cache_ (i.e., that thing believed to have been merely a transparent optimization).
It turns out now that the belief in the cache being transparent and only providing optimizations was the flaw in the belief system at the time.
You're putting words in other people's mouths here.
If it was such a mythical attack why doesn't it work on anyone else's chips? Intel screwed up hard on this and I have no sympathy. Hopefully the incoming lawsuits will make up for the massive amount of money wasted for the performance losses
I've got a friend in CPU design and he's only got about 50 companies he can work for in the world where he could do the same job he does now
I’ve written RTL that speculatively fetches data from memory in order to avoid bubbles in a pipeline. Not a CPU, but the concept is exactly the same.
If somebody had assigned me to a CPU project without guidance from a security architect and ask me to speculate reads, I’d probably have done the same as Intel.
The chance that the same guy did both CPUs is small. It’s just that it’s not an unreasonable way of doing thing if you’re not familiar with these kind of attack.
And if anyone even considered the results of the reads, they likely saw them as nothing more than free cache pre-fetch instructions that would enhance performance should the speculative path turn out to be the correct path after-all.
And because the push was for yet more performance, free cache pre-fetch operations were likely viewed as a great bonus.
> We also tried to reproduce the Meltdown bug on several ARM and AMD CPUs. However, we did not manage to successfully leak kernel memory with the attack described in Section 5, neither on ARM nor on AMD. The reasons for this can be manifold. First of all, our implementation might simply be too slow and a more optimized version might succeed. For instance, a more shallow out-of-order execution pipeline could tip the race condition towards against the data leakage. Similarly, if the processor lacks certain features, e.g., no re-order buffer, our current implementation might not be able to leak data. However, for both ARM and AMD, the toy example as described in Section 3 works reliably, indicating that out-of-order execution generally occurs and instructions past illegal memory accesses are also performed.
theo de raadt of openbsd called this out, 10 years ago: https://marc.info/?l=openbsd-misc&m=118296441702631
> I should note that we kernel programmers have spent decades trying to reduce system call overheads, so to be sure, we are all pretty pissed off at Intel right now. Intel's press releases have also been HIGHLY DECEPTIVE. In particular, they are starting to talk up 'microcode updates', but those are mitigations for the Spectre bug, not for the Meltdown bug. Spectre is another bug, far more difficult to exploit than Meltdown, which leaks information from other processes or the kernel based on those other processes or kernel doing speculative reads and executions which are partially managed by the originating user process. Spectre does NOT involve a protection domain violation like Meltdown, so the Meltdown mitigation cannot mitigate Spectre.
> These bugs (both Meltdown and Spectre) really have to be fixed in the CPUs themselves. Meltdown is the 1000 pound gorilla. I won't be buying any new Intel chips that require the mitigation. I'm really pissed off at Intel.
[1] http://lists.dragonflybsd.org/pipermail/users/2018-January/3...
However for unknown reasons it’s worse in Intel. They were not able to exploit the vulnerability to dump the kernel memory on AMD and ARM.
So I find the whole “Intel you suck” really out of line. I would not be surprised someone clever will come and figure out a Meltdown style attacks which works on AMD and ARM CPUs.
As Bruce Schneier says: “attacks always get better they never get worse”
No ME, no PSP, no meltdown, and no spectre.
Still, no Meltdown and no Spectre on any RPi
So AMD looks like a fair choice if this is true.
Basically the choice between Intel and AMD in the context of in-chip vulnerabilities comes down to whether one feels better choosing a chip with known vulnerabilities + significant effort to mitigate or a chip without known vulnerabilities and without significant efforts at mitigation. In both cases, unknown vulnerabilities are probably equal.
To put it another way, building a computer is largely a consumer process not a technical one. The biggest security risk is software you download and run, not hardware flaws (e.g. rowhammer)
This is a bogus criticism. The issue was successfully embargoed for something like six months(!) before being published a week early.
Due to how i386 TLB works (ie. no ASID) doing that is essentially required to get reasonable performance, micro kernel or not.
edit: apparently it has been:
http://www.qnx.com/developers/docs/6.5.0/index.jsp?topic=%2F...
Wikipedia states that supported platforms historically were i386, PowerPC, MIPS, SuperH and ARM. And I would not be too surprised if there was support for Altera Nios-II with MMU in the times when QNX was owned by Harman (As at least the touch screen part of some versions of Harman made Audi MMI is based on NiosII).
By the way, here's a short script that will download a (full) page with wget, add it to IPFS (if you have a running daemon) and copy the Eternum link to clipboard:
[1]: https://www.eternum.io/ipfs/QmQU1bCPsg7VY5puKqH6wZfAv1ZWBteK...
Any other IPFS node will work just as well, such as https://ipfs.io/ipfs/QmQU1bCPsg7VY5puKqH6wZfAv1ZWBteKTCgHRKz....
Large open source projects tend to have significant amount of funding - think about Chrome, OpenJDK, the Linux Kernel...
Seems to have worked for the IPFS guys (I'm talking about existing projects that only launched a cryptocurrency later on, rather than from the start). Brave's BAT qualifies for that, too.
This is the sparc and sun or MIPS and Irix all over again... except this time without any viable business model to keep moving forward and progressing on hardware or software.
So what if they make their processor that is free from all known processor level bugs. Would that mean dropping support for all the other processors? or would they still need to be aware of the processor bugs that exist in other architectures that they'd still be expected to create patches for? Would they have the resources to fix processor bugs in their own design when they are discovered?
My read on this is that it doesn't make any sense for an operating system organization that lacks any financials to try to pursue custom hardware.
it's no surprise they're not high up on the list of important people to tell
* https://news.ycombinator.com/item?id=16074531
That is hardly enough time, considering that it is likely a holiday period for most of the people concerned, to prepare what needed to be prepared. Google and Intel gave themselves six months, in contrast, and they gave Ubuntu since November 2017.
I still haven't nailed down how much advanced notice Matthew Dillion, author of DragonFly BSD, was given; although his apparent level of preparedness with patches on 2018-01-05 can be attributed to the fact that he actually talked publicly about the problem in April 2017, months before Google Project Zero secretly reported it to processor manufacturers in June 2017 in the first place.
Correct, but it’s significantly better than OpenBSD’s timeline (which learnt about it from the media). Mitigation for Meltdown is possible within of a week, or two. It requires working through the night, but it’s possible.
Mitigation after it’s in the media is a whole different situation.
From https://www.krackattacks.com/#openbsd :
> As a compromise, I allowed them to silently patch the vulnerability.
Receiving permission to patch is the opposite of breaking an embargo.
> As a compromise, I allowed them to silently patch the vulnerability. In hindsight this was a bad decision, since others might rediscover the vulnerability by inspecting their silent patch. To avoid this problem in the future, OpenBSD will now receive vulnerability notifications closer to the end of an embargo.
For example, from the conversation on the KRACK embargo:
>Q. What is the rationale on extending the embargo so long? To give vendors a chance to patch? It seems like once people in the know know about a vulnerability, the shorter an embargo time the better.
> A. There’s a couple of different interests at play.
> For instance:
> Researchers want to make a timed media splash.
> Security agencies want to evaluate the problem and make sure they get patched before anyone else.
> Vendors want time to prepare patches, yes, but that alone does not justify such a long delay.
> Reviewing the patch, testing it, and preparing it for commit and publishing erratas took only a couple of hours of my free time.
Apparently Meltdown was discovered back in June, or seven months ago. Only a select few were told about it and have been working on it. It is apparent that some are more equal than others in the "free and open source community," which ironically apparently includes Microsoft, Apple, and the three-letter agencies.
To quote that last paper:
> We show how processor architecture features such as simultaneous multithreading, control speculation and shared caches can inadvertently accelerate such covert channels or enable new covert channels and side channels. We first illustrate the reality and severity of this problem by describing concrete attacks.
So not just were side-channel attacks on CPUs documented; specifically the dangers of speculative execution and caching on intel cpus (there: the Itanium) was documented.
It's inconceivable that this comes as a surprise to Intel. They're playing the public for fools.
Of course some of this is intrinsic: speculative loading, caching, and accurate timers are an intrinsic problem: an JIT running in the same process as the JITted code will not be able to keep secrets from that JITted code while all three are in play. But intel specifically also allowed this across memory protection boundaries, which isn't necessary.
Quote from the meltdown paper:
> We expect that Meltdown and Spectre open a new field of research to investigate in what extent performance optimizations change the microarchitectural state, how this state can be translated into an architectural state, and how such attacks can be prevented.
So we won't see the last of Spectre.
Not to be cynical or downplay the research, but that's the kind of thing you (have to?) put into papers to make getting research funding easier in future. Hardly a new field, just very difficult to do research on.
Why does that matter? Because people seem to imply intel isn't to blame here; that since the attacks were unforeseen, it's essentially just bad luck. But that's not the case: the risky bits of architecture were known. It's like finding a memory corruption vulnerability and then deciding not to fix it since you don't also know of an exploit - that's just being careless.
Furthermore; finding holes like this isn't some divine act; although there is luck and skill involved there's no question that others can do so too. Even if an attack were completely new, you have to ask yourself how hard it was to find. Does a supplier have any responsibility to provide secure things? Then they should be looking for these kinds of holes. Obviously to a novice any hack like this looks almost magical - but there's a risk that by interpreting any novel attack as not-their-fault that the vulnerable supplier never bothers trying to make anything secure.
Intel was being careless. They should definitely have known this was exploitable even without knowing how to exploit it. And it's at least conceivable that they could have found this or an exploit like this if they had tried.
https://www.ibm.com/blogs/psirt/potential-impact-processors-...
Out-of-order execution is an indispensable performance feature and present in a wide range of modern processors.
quote
However, for both ARM and AMD, the toy example as described in Section 3 works reliably, indicating that out-of-order execution generally occurs and instructions past illegal memory accesses are also performed
end of quote
it just need more work to make it work their
Meltdown involves crafting some code yourself that speculatively accesses data it does not have access to before the CPU rejects the access due to permissions.
In other words, Meltdown involves your code accessing the data, but Spectre involves getting the kernel to do it for you. Spectre is a neat clever side-channel attack that is hard to protect against, while Meltdown is a blundering error in the way the CPU works.
There are 3 vulnerabilities.
Meltdown is 1 of the 3. Meltdown is pretty much Intel only. Some ARM SoCs are also affected, but these are relatively rare. AMD64 is unaffected by Meltdown.
Spectre are the other 2 vulnerabilities. Spectre affects pretty much everyone.
Meltdown is more severe, and more of a blunder.
For a technical explanation, see [1]. Was recently referred to at HN as well.
[1] https://www.raspberrypi.org/blog/why-raspberry-pi-isnt-vulne...
From the Meltdown paper[0] section 6.4 it would seem that out of order execution referencing illegal memory locations still occurs unless of course I'm misunderstanding something.
AMD's claim is that the result of the faulting instruction can never end up affecting a later speculated instruction - ie, for Intel the bad access says "here's the value you asked for" and later says "actually you don't have permission, forget you saw that". For AMD the bad access says "you don't have permission, so no result available"