How Screwed is Intel without Hyper-Threading?
techspot.com
techspot.com
If you're using an overclockable workstation CPU, the good news is that turning SMT off gives you additional ~200Hz of overclocking headroom. Modern Intel chips are thermally limited when overclocking and turning SMT off makes them run a bit cooler. My i9-7900X overheats over 4.4Ghz, but with SMT off it's comfortable at 4.7Ghz on air (Noctua D15).
[1] https://www.guru3d.com/news-story/intel-core-i7-8700k-benchm...
Here's another problem with the article:
> Microsoft is pushing out OS-level updates [...] However, this doesn’t mitigate the problem entirely, for that we need motherboard BIOS updates and reportedly Intel has released the new microcode to motherboard partners. However as of writing no new BIOS revisions have been released to the public.
That's wrong - microcode updates can be successfully applied by either the OS or by the early-boot firmware, you do not need a BIOS update (just apt update intel-microcode).
Sure it is. I've done it myself.
Microcode is stored in volatile memory on the CPU. Updates are applied on boot, every boot. "Downgrading" is as simple as not applying updates, or applying an older update.
That'd be 200 MHz, not 200 Hz?
https://www.geeks3d.com/20190306/tested-cinebench-r20-cpu-be...
So Techspot numbers are fine, just a typo on version of benchmark, should say R20 on graph not R15.
In any case, the article is worst-case FUD-like nonsense. No idea what the eventual perf loss will be, but it's a safe bet it'll be less than that.
Consider that it's likely safe enough for 99% of usages for the OS to allocate CPU cores to processes in pairs, and that would incur only a small penalty. And something smarter is almost certainly possible, too.
"First off we have Cinebench R20 results ..."Personally I wonder whether the cost of mitigation is worth it. According to the article (and their simplified HT methodology) certain workloads experience a 25% performance hit.
The only cases I currently consider as exploitable are VMs in the cloud (perhaps a reason to own dedicated servers after all this time) and running JS in the browser (perhaps a reason to disable JS).
There will always be side-channel attacks. Our devices are physical devices and will leak side-effects, like electromagnetic radiation ( https://en.wikipedia.org/wiki/Tempest_(codename) ). This recent spate of MDS flaws don't necessarily fit in my threat model.
Maybe instead of slowing everything down for the sake of JavaScript, we need a special privilege mode for untrusted code that disables speculative execution, sharing a core with another hyper-thread, etc.
You are exactly correct. However, the advent of "The Cloud" changed that.
"The Cloud" by definition runs untrusted machine code right next to yours. IBM and other players have been screaming about Intel being insecure for decades--and it fell on deaf ears.
Well, suck it up buttercup, all your data are belong to us.
While I hear lots of whingeing, I don't see IBM's volumes rising as people move off of Intel chips onto something with real security. When it comes down to actually paying for security, people still blow it off.
And so it goes.
Did someone already demonstrate that speculative-execution bugs are observable in Google's cloud stack?
In particular, disabling JS would be pretty disabling for an average modern web user, so an easy, practical attack through this vector is especially relevant.
If we were able to run without JavaScript and have basic use of websites again then these attack vectors wouldn’t be so scary.
Maybe it’s a personal failing of mine; but I don’t understand why people don’t consider JavaScript untrusted code when other vectors of running software on my machine come from a gpg-signed chain of custody.
Sure, passive content such as nytimes.com can work without JS (although some of their interactive articles would not), but anything more complicated will often be done with client side rendering these days.
Not true, latency scales with CPU capacity on the client. SPAs now exist with client side rendering to mask the bloat in the site plus its dependencies.
If you have a SPA arch but you did a page load per click, you would die. But all sites creep up to 500ms page load times regardless of their starting stack and timings.
Speaking from about a decade of experience with progressive enhancement and all the other things. It is 'much' more. There is an expectation of equivalent functionality/experience in these cases, and you just can't spread your resources that thin to get half a product out that works without Javascript. You're literally development a completely independant application that has to sit on top of another application. Everything comes out 'worse'
These days we invest in ARIAA and proper accessibility integration and if you run without Javascript that is going to get you next to nothing in a 'rich' web application.
Interpreting the code instead should completely kill any chance of exploits of this class. It will also completely kill the performance though. Even tor dropped that idea https://trac.torproject.org/projects/tor/ticket/21011
But later on, Firefox has streamlined its architecture and made it more difficult to disable JIT. After one update, disabling JIT causes a broken build that crashes instantly. I've spend a lot of time reading bug trackers and looking for patches to unbreak it. The internal refactoring continued, after Firefox Quantum, JIT seems to be an integral part of Firefox, and the Gentoo maintainers eventually dropped the option to disable JIT as well.
I wonder if an ACL-based approach could be used, as an extension to NoScript. The most trusted website has JavaScript w/o JIT, the less trusted website has JavaScript w/o JIT, the least trusted website has no JavaScript. Tor Browser's security lever used to disable JIT at the "Medium" position. But rethinkig it, this approach has little security benefit, it's easy to launch a watering hole attack (https://en.wikipedia.org/wiki/Watering_hole_attack) and eject malicious JavaScript from 3rd-party elements.
I wonder if extending seccomp() and prctl() based approach can be a solution. SMT can be enabled but no process is running on a SMT thread by default. Non-confidential applications such as scientific computing or video games can tell the kernel to put their processes on SMT threads.
A valid option, though in general I'd rather allow it for everything except browsers.
So instead of trying to achieve the impossible (perfect safe code that still has unlimited access) the direction is stricter sandboxes. Then at least you only need one near perfect piece of safe code (the sandbox) instead of tens of thousands.
Instead, the only way of having a performant, secure system today is to disable hardware mitigations and ensure you only run trusted software, the opposite of your proposal. The sandbox still helps for other issues (e.g. buffer overflows).
I think that ship has sailed..
We should work to build secure sandboxes instead.
Also, to keep this all in perspective, we are talking about one company that made egregious engineering decisions to maintain their market leadership which has put their customers at risk. To say we should do away with JS in the browser because Intel made some really poor decisions to undermine their competition is just crazy. As far as I'm concerned, this is typical Intel and they are getting what they deserve yet again. I just feel bad for their customers that have to deal with this.
Fixed it for you, the current list I have is AMD, ARM, Freescale/NXP, IBM both POWER and mainframe/Z, Intel, MIPS, Oracle SPARC, and maybe old Fujitsu/HAL SPARC designs for Spectre, with at least four of those CPU lines also subject to Meltdown.
Given that AMD is fully vulnerable to Spectre, there's absolutely no reason to believe it isn't similarly vulnerable to microarchitectural detail leakage if people were to look. And going back to what I was replying to:
> we are talking about one company that made egregious engineering decisions to maintain their market leadership which has put their customers at risk
We demonstrably aren't, seeing as how ARM, and IBM, both POWER and mainframe/Z CPUs are also vulnerable to Meltdown. That and the significant prevalence of Spectre says this is not "typical Intel" but "typical industry", a blind spot confirmed in 7 different companies, and 8 teams to the extent the IBM lines are done by different people.
The "Intel is uniquely evil" trope simply doesn't hold water.
That said, in the consumer space, being (this) vulnerable to Javascript attacks is catastrophic. My original point is that we should not be crippling something very useful (javascript in the browser) because of flawed architectures that mostly affect one company in a way that decimates performance.
Lots of us have a different opinion on the usefulness vs. risk of running random and often deliberately hostile JavaScript in your browser, see the the popularity of NoScript, and how many of us use uMatrix with JavaScript turned off by default. Most of the time I follow a link where I don't see anything, I just delete the tab, most of those pages aren't likely worth it.
"Mostly affect one company" is something completely unproven, since AMD is not getting subjected to the same degree of scrutiny, AMD has a minuscule and for some reason declining in 19Q1 market share for servers, while desktop and laptops are modest but showing healthy market share growth: https://www.extremetech.com/computing/291032-amd-gains-marke...
While ARM is announcing architecture specific flaws beyond basic Spectre: Meltdown (CVE-2017-5754) and Rogue System Register Read (CVE-2018-3640, in the Spectre-NG batch but by definition system specific): https://developer.arm.com/support/arm-security-updates/specu...
This may be popular among neckbeards, but regular people could care less. Regular people care about being tracked and that's about it.
> "Mostly affect one company" is something completely unproven
You speak like a person of authority on this subject, but AMD has come out and said they are specifically immune to these threats: https://www.amd.com/en/corporate/product-security
AMD has "hardware protection checks in our architecture" which disputes your assertion that AMD just isn't a targeted platform. The reality is any computing platform can be vulnerable to undiscovered vulnerabilities so making that point is kind of pointless.
Also, disabling JS on the browser pretty much completely eliminates e-commerce on the web. Again, I can't fathom the masses seeing any benefit in this.
The exploits that do worry me, are ones that can be done remotely without any action from the user. Heartbleed is a good recent example. Fortunately those tend to be rare.
Security is always relative to a threat model, and not everyone wants perfect (if that can even be possible) security either, contrary to what a lot of the "security vultures" tend to believe.
All the more reasons to run your own servers and compile your own binaries. Like we used to do.
Compiling once from a tarball and reusing that can definitely reduce the number of times you would need to trust something from a third party.
The real problem is why I mentioned auditing: the attacks we’ve seen over the years have been updates from maintainers who followed the normal process. Auditing is the most reliable way to catch things like that because the problem isn’t the distribution mechanism but the question of being able to decide whether you can trust the author.
This is fucking scary. If it's a package used by Wordpress you could end-up with 30% of the web open to an attack.
1 - See "Synchronized ring 0 entry and exit using IPIs" in https://software.intel.com/security-software-guidance/insigh...
> There is no hardware support for atomic transition of both threads between kernel and user states, so the OS should use standard software techniques to synchronize the threads.
Ha, “standard software techniques”. Let’s see: suppose one thread finishes a syscall. Now it sets some flag telling the other core to wake up and go back to user code as well. But the other core is running with interrupts on. So now we sit there spinning until the other core is done, and then both cores turn off interrupts and recheck that they both still want to go to user mode. This involves some fun lock-free programming. I suppose one could call this a “standard software technique” with a somewhat straight face.
Meanwhile, the IPI on entry involves an IPI and the associated (huge!) overhead of an x86 interrupt. x86 interrupts are incredibly slow. The IPI request itself is done via the APIC, and legacy APIC access is also incredibly slow. There’s a fancy thing called X2APIC that makes it reasonably fast, but for whatever reason, most BIOSes don’t seem to enable it.
I suspect that relatively few workloads will benefit from all of this. Just turning off SMT sounds like a better idea.
Intel themselves said [1] "Intel is not recommending that Intel® HT be disabled, and it’s important to understand that doing so does not alone provide protection against MDS."
[1] https://www.intel.com/content/www/us/en/architecture-and-tec...
[1] https://www.intel.com/content/dam/www/public/us/en/documents...
The "don't apply the updates" argument in security is basically philosophically comparable to the antivax movement
As long as you also apply the quarantine principle when relevant. And herd immunity, both applied backwards:
Quarantine for when you run code that's under your control, like an HPC cluster; the biggest dangers from these vulnerabilities are when you're running completely random code supplied by others, JavaScript in browsers for normal users, and the multi-tenant systems of cloud providers.
Herd size matters because the bigger the herd, the larger the number of potentially vulnerable systems. IBM mainframe chips are vulnerable to Meltdown and Spectre, but there aren't hardly as many of them as x86 systems. Although the payoff of compromising them is likely to be disproportionately high.
Is this situation so unexpected that major companies don't even know how to react? If I were Amazon or Google, I'd be requesting my 30% discount, retroactively. Perhaps they are?
IIRC Vulkan was designed to be more suited to modern multi-core processors, and I believe that the "HT-agnosticism" it exhibits is good proof that they managed that pretty well...
Wouldn't that win half the battle, at least on server-side? It wouldn't apply to workstations, of course. Malicious/untrusted processes running there would still need to be isolated (malicious JavaScript in the browser, say).
Also, am I missing something, or are these concerns really not applicable to gaming applications (which the article emphasises)? That's a highly trusted codebase, so securely containing it isn't worth a 15% performance hit, no?
plenty of places rent bare metal machines, and the price doesn't seem quite that world-ending. see packet.net for instance.
and it only gets better if on the bare metal machine you get to turn off all of the performance-draining fixes.
I don't care whether Amazon use virtualisation to get me my 'instance'. They're always careful to stick with 'instance', and never to commit to 'VM'. Wise move, now that they're offering physical ('dedicated') instances [0]. Time will tell whether they take an interest in low-horsepower dedicated instances.
They're probably a bad match for when you have a single batch of work to do, like when the New York Times converted their archives, or every time Netflix has to reencode their video library.
I'm not convinced the cost difference would be a showstopper. More little computers rather than few big computers. Same total horsepower. More motherboards, sure.
Not if they can be made smaller and packed more densely.
> more hardware management overhead
I'm not sure this is a significant issue.
> more and longer wiring
I don't think this follows. They could be clustered. A cluster could be made to look a lot like a conventional big-iron server.
We're talking about replacing a single large machine running, say, 16 VMs (each with an IP), with 16 small machines (each with an IP). Put an Ethernet switch at the last-mile and it's virtually the same, no? Different number of MAC addresses, sure.
What issues? If your cloud architecture can't save power when demand is low, how can that be a positive?
> Go too far and you won't be able to also run a spot market in CPU instances like AWS does.
As I understand it, a spot market isn't ever an architectural goal, it's a way to exploit the practical realities of unused resources/capacity. If that necessity goes away, that'd be a win.
I see at least 3 possibilities:
1. The spot market is about filling temporarily-vacant VM slots on host physical machines
2. Additional physical machines must be kept online so that VMs can be quickly spun-up, and the spot market is about capitalising on this necessity (this is really a variation on Option 1)
3. The spot market is about deriving at least some revenue from the unutilised fraction of an always-on fleet, encouraging customers to run heavy jobs during times of low load on your cloud
I doubt Amazon's fleet is always-on -- I imagine they have fairly significant capacity in reserve to keep up with growth and spikes -- so I doubt it's Option 3. (This is a perfectly obvious question but I really can't find a definitive answer. Maybe my Google-fu is weak today.)
I wonder about live migration. I imagine there must be some brief interruption even there.
> Many VMs are run indefinitely, so whatever machine they’re on has to stay powered on
Amazon must be adequately motivated to find a way to power-down unused capacity. My guess is they have some kind of system to maximise this.
[0] https://www.theregister.co.uk/2012/12/20/aws_ec2_servers_ret...
Plus Xen has live migrate now, so I assume AWS’s frankenxen would have it as well.
There is indeed a brief interrupt, which is why I said they can’t be too aggressive with it.
Anyways, I’m sure Amazon does what it can to power down machines, I just think they’re stuck leaving much the data center running because they can’t kick off running VMs.
That doesn't sound right. A large number of weak computers, should be easier to cool. A cluster of smartphones can get by with passive cooling, but my laptop cannot.
Apart from cooling, I'm not an expert on this but I don't know that performance-per-watt has to improve as overall power increases. Going with my idea you'll need to power more motherboards, but they'll be less powerful motherboards. (The total number of CPU cores being powered remains the same. Same goes for RAM modules.)
Is that true? My phone gets quite hot if I run a taxing load on it for too long. My, perhaps poor, intuition tells me that phones can rely on passive cooling because of the burst nature of phone use. If you’re utilizing your smartphone cluster, I can imagine your fleet overheating.
iPhones survive long sessions of 3D gaming even when wrapped in protective plastic shells. If it was an issue, use a bigger passive heatsink than the one they cram into the smartphone form-factor. You could never do that with a server-grade CPU.
With a high-power CPU, the heat emissions are more concentrated than if we use many low-power CPUs.
The Raspberry Pi would have been a better example.
If cooling systems for machine rooms haven't caught up with modern systems that dissipate more heat, the direction you're thinking has merit. But I doubt there are serious limiting factors here, besides cost.
I think we agree, but I wasn't really thinking of power-savings as the real motivation here, I was thinking of security.
It seems likely we're going to see a steady steam of Meltdown/Spectre-like security concerns. We can have our cake and eat it -- isolation and all the CPU cleverness there is -- if the isolation between instances is done physically rather than with VMs.
There's clearly a market for weak, cheap, machines -- Amazon didn't originally offer a 'Nano' instance-type, but they do now.
So: even if I built that machine to not be exclusively used for gaming (which would be an outlier), anyone who has a machine of that kind could easily motivate making it exclusive for gaming, because there is so much money in that 10-15% that we can buy a second much cheaper machine (e.g. a low end laptop, chromebook etc) to use for other things.
* Gaming-only machines, like my RetroPie example.
* Media servers, like for Plex or Kodi.
* NAS servers, powered by software like UNRAID.
Is there any way to mitigate this performance drop, disable the fix?
Then when your fuckups are discovered, other people not on your payroll expend their own resources to fix it.
And you don't even have to compensate customers who paid full price but now find 40% of their compute evaporated.
And that massive performance hit needed to fix your fuckups stimulates incremental demand for your product because customers now need more cores to finish their compute.
This isn't an excuse for Intel and how they (sometimes poorly) handle issues that come up. And it is well known that many of these vulnerabilities are currently specific to Intel and how their subsystems are implemented, or only exist on Intel processors. But side-channel attacks are a nearly universal vulnerability.
SMT of which Hyper-Threading is an implementation of is used in many chips from many companies, AMD's Bulldozer, IBM's POWER, Sun/Oracle's Sparc, etc.
These exploits are not a monolithic process, but many techniques chained together to get most of these exploits to work, and this common knowledgebase which has been refined by the researches will most certainly be recycled and reused on other processors.
If anything, the "capture 90% of the market." by Intel as you put it only ensures that these problems will usually be discovered on their processor's first.
If as much effort was put towards other cpus as it is towards Intel, you would see significant performance hits as well with them.
> If anything, the "capture 90% of the market." by Intel as you put it only ensures that these problems will usually be discovered on their processor's first.
Precisely. I haven't looked at the MDS set of flaws yet, but to take the first version of Foreshadow/L1TF, it targets the Intel SGX enclave, by definition the designs of other companies that don't have SGX won't be vulnerable. But as you note this doesn't mean they don't have their own design or implementation specific bugs.
The software stack is at fault for building their security model on assumptions about what the hardware did that aren’t guaranteed in the spec. The software assumes the existence of this magical isolation that the CPU never promised to provide.
But Intel cut corners and offloaded most of the downside onto their customers and partners while pocketing the upside for themselves.
Vulnerabilities like Meltdown and the latest "MDS" vulnerabilities are absolutely a design decision.
Think about Meltdown, for example. The problem here was that data from the L1$ memory was forwarded to subsequent instructions during speculative execution even when privilege / permission checks failed. Now the way that an L1$ has to be designed, the TLB lookup happens before reading data from L1$ memory (because you need to know the physical address in order to determine whether there was a cache hit and if so, which line of L1$ to read), and the TLB lookup gives you all the permission bits required to do the permission check.
Now you have a choice about what to do when the permission check fails. Either you read the L1$ memory anyway, forward the fact that the permissions check failed separately and rely on later logic to trigger an exception when the instruction would retire.
This is clearly what Intel did.
Or you can treat this situation more like a cache miss, don't read from L1$ and don't execute subsequent instructions. This seems to be what AMD have done and while it may be slightly more costly (why?), it can't be that much more expensive because you have to track the exception status either way.
The point is, at some place in the design a conscious choice was made to allow instructions to continue executing even though a permissions check has failed.
The MDS flaws seem similar, though it's a bit harder to judge what exactly was going on there since it's even deeper in microarchitectural details.
And ARM, and IBM, both the POWER and mainframe/Z teams.
Although this design decision of Intel's was in the early 1990s, shipping in their first Pentium Pro in 1995. When AMD's competition was a souped up 486 targeting Intel's superscaler Pentium, although I assume they were also working on their first out-of-order K5 at the time, it first shipping in 1996.
https://en.wikipedia.org/wiki/AMD_K6
The AMD K6 is a superscalar P5 Pentium-class microprocessor, manufactured by AMD, which superseded the K5. The AMD K6 is based on the Nx686 microprocessor that NexGen was designing when it was acquired by AMD. Despite the name implying a design evolving from the K5, it is in fact a totally different design that was created by the NexGen team, including chief processor architect Greg Favor, and adapted after the AMD purchase.
Is the point of your argument that others did it too, so Intel shouldn't be accountable?
Or that others did it too, so it's reasonable to believe that the vulns were impossible to prevent?
This argument sounds a lot like the "deflect" part of Facebook's strategy to "delay, deny, deflect".
With VIPT caches (which AFAIK both AMD and Intel use for the L1 data cache), the TLB lookup and the L1 cache lookup happen in parallel. The cache lookup uses the bits of the memory address corresponding to the offset within the page (which are unchanged by the translation from virtual to physical addresses) to know which line of the L1 cache to read. Only later, after both the TLB and the L1 cache have returned their results, is the L1 tag compared with the corresponding bits of the physical address returned by the TLB.
This is clear from the fact that a page is 4KB while the L1 data cache is 32KB, so the page offset cannot contain enough information to tell the processor which part of L1 to read.
(Somewhat ironically, I think "Hanlon's razor is just a saying, not an argument" is just a saying, not an argument.)
Charity (which Hanlon’s razor is a part of) only applies to individuals.
Why?
Do you think that the only reason for applying Hanlon's razor is some sort of moral principle (that only applies to individuals)?
I think the reason to apply Hanlon's razor is that stupidity is far more widespread than malice, and so that should be my prior. This is a purely epistemic argument, with no moral component.
I also think that larger groups always contain more stupidity than smaller groups, and I think it grows super-linearly. On the other hand, the effect of group size on malice (and on benevolence) is very complicated and unpredictable. Do you disagree with one of these beliefs?
If you agree that groups are almost always stupider than people but only sometimes more malicious, I think it's clear that I should be at least as eager to apply the razor to groups as to people.
In conclusion, an apposite quotation:
> Moloch! Nightmare of Moloch! Moloch the loveless!
Take the 737 Max as an example: It's absolutely malice, because the level of incompetency you'd have to have to let a plane into the air that can tilt all the way down by means of a single(!) faulty sensor would disqualify anyone from ever building a plane.
An aviation problem such as that is still comparatively easy for a layman to understand. When it comes to CPU-microarchitectures, I'm not so sure. But I trust they have professionals designing their chips, and my default is to assume they are competent and that malice has occured, and rather they'd have to prove the opposite.
This is extremely disingenuous phrasing. You seem to be trying to craft a narrative that makes putting security vulnerabilities in products an intentional thing to sabotage people, time, and resources.
It's intentional business decision to maximize performance and minimize costs paid by Intel, making their products faster and generating higher margins.
They achieved this by gambling that vulns wouldn't be exploitable. But Intel externalized that risk onto their customers and partners without their knowledge or consent.
So when the gamble failed, the rest of the ecosystem bears the cost while Intel harvests the gains for themselves.
Externalizing risks and losses on others without their knowledge or consent is a form of theft.
That's a huge leap, and seemingly without merit. You're still just arguing that Intel knew this would create huge vulns and they went ahead regardless, which is baseless.
The first scientific paper about such vulnerabilities namely in Intel processors is from 1995 [1]. They went ahead regardless. Processor technology has developed a lot. Many things that were not practically feasible in 1995 are possible today. That holds for the good and the bad. There must be some people inside Intel who must have understood that. Otherwise all that progress would not have possible. Some experts might have blind spots, but I doubt Intel microarchitecture is developed by just a handful of people.
[1] https://en.m.wikipedia.org/wiki/Meltdown_(security_vulnerabi...
I lack the technical expertise to honestly critique the architecture. My assumption is that the folks behind the design of some of the more advanced technology on the planet aren’t morons or seeking to defraud the market.
All of the noise here is about significant vulnerabilities that haven’t been exploited in public. It is a serious defect, but not the end of the world.
And the decision to implement it without that study, or despite it, was made. Because implementing it increased performance which made the numbers which captured market share.
It's not like the processor designed itself that way. Humans made choices along the way, and every choice is a tradeoff. Execution-time attacks and other sidechannel leakage have been well-known for years, and I can't imagine that at a place like Intel, nobody had heard of that.
More than the following big things, in both directions?
- A consistent fabrication advantage until very recently? That's said to have wiped out a generation of clever hardware architects, when Intel beat their best efforts with their next process node.
- Falling behind AMD with the Netburst "marchitecture", the front side bus memory bottleneck, and only supporting Itanium as a 64 bit architecture?
- Reversing the above by licencing AMD64, reverting to the same Pentium Pro style design AMD was using, and copying their ccNUMA multi-chip layout (each is directly attached to memory, with fast connections between them).
- AMD losing the plot after the K8 microarchitecture, including putting a huge amount of capital into buying ATI instead of pushing their CPU advantage? And then having to sell off their fabs, putting them at a permanent disadvantage until Intel's "10nm" failed? (And what happened to the K9??)
- Anticompetitive marketing and sales?
> Execution-time attacks and other sidechannel leakage have been well-known for years, and I can't imagine that at a place like Intel, nobody had heard of that.
The strange thing is that security researchers assumed for years that Intel was accounting for this, when it turned out not a single one of AMD, ARM, IBM POWER and mainframe/Z, Intel, MIPS, or SPARC did?
Because we should assume all your past missteps were due to willful negligence? Yeah, let's not just start making shit up to support the group think outrage.
and either no one thought about the security implications, or no one could come up with a security problem at design time - likely just no one thought about side channel problems.
It obviously wasn't obvious to the very, very many people worldwide who know about execution-time attacks and other sidechannel leakage, as it took more than 10 years of almost every chip worldwide being vulnerable until it was noticed. It's not as if some 2016 design tradeoff suddenly enabled Meltdown. Engineers were carefully studying competitor's chip designs, and none of them noticed that flaw for more than a decade.
Those professors aren't doing too bad now either, with plenty of papers to be published outlining the flaws and proposing solutions.
/s
Really, though, humanity is still pretty new to this whole computer architecture thing in general, and running untrusted code securely in particular. Im skeptical that someone at Intel was aware for the possibility of vulnerabilities for 10+ years before any of them were made public.
To be honest, I'm a little perplexed by the animosity directed towards Intel after all these vulnerabilities. Vulnerabilities in software are most frequently the result of ignorance, but everyone seems to assume the engineers and managers at Intel knew better (or should have know better). Who would have told them?
http://www.daemonology.net/hyperthreading-considered-harmful...
That is, these earlier timing side channels can be avoided, even in the presence of SMT/HyperThreading, by not doing data-dependent branches or memory accesses, and not using instructions with data-dependent timings. For instance, instead of "x = cond ? a() : b();", do something like "mask = cond - 1; x = a() & ~mask | b() & mask;".
That is not the case with MDS: even if you carefully avoid data-dependent branches, memory accesses, and instruction timings, you are still vulnerable. It's really a new vulnerability class.
Anyway, that's not the same vulnerability as any of those that have become public in the last 18 months.
The issue with MDS is not which CPU resources each thread is using. It's something much more subtle. The issue with MDS is data left over by each thread within internal temporary buffers, which shouldn't be an issue since these internal temporary buffers can't be read by the other thread before being overwritten. However, in the phantasmagoric world of Spectre, these residual values can leave an impression which can be picked up through careful manipulation.
Notably, MDS can also be exploited by the same thread. This is what the microcode updates are about: they hijack a useless legacy instruction from the 90s, adding to it the extra behavior of clearing all these internal temporary buffers. Operating systems are supposed to call this instruction whenever crossing a privilege boundary. The problem with SMT is that it's really hard to make sure both threads are on the same side of all privilege boundaries all the time.
Also note that, as far as we know, all AMD processors, even the ones with SMT, seem to be immune to MDS. This shows that this vulnerability is not a fundamental property of SMT, unlike other side channels which depend on the resource usage by each thread.
Or maybe it was an industry blind spot, now a very big deal because so many of these CPUs run other people's random code, JavaScript in browsers, and anything and everything in shared cloud hosts.
Followed the Wikipedia link to the 1995 Fujitsu/HAL out-of-order design (https://en.wikipedia.org/wiki/HAL_SPARC64) and it says it's superscalar, but then that the first version can execute as many as 4 instructions at once and out-of-order. And from the very first version had branch prediction, which equals speculative execution as far as I know.
In my case, we observed very marginal performance impacts from last year’s mitigations.
So I think this is some sort of misguided hit piece against intel.
Everyone knows pcs are riddled with security flaws less obscure than this. People who run their business on cloud servers might care. Gamers though? No.
https://react-etc.net/entry/javascript-spectre-meltdown-vuln...
Gamers are are the 5 dollar whores of the security world, they are riddled with bugs and the dont give a crap.
The only two mitigations is either disabling hyper-threading or disabling JIT.
IMO Intel customers should not be indifferent, what you thought you bought was not what you got latter, you lost performance and you have to do advanced stuff to disable the security patches and be unsafe
Edit: My bad, I considering buying i7 and got an i5 , I got confused(my i5 is a K CPU 6th gen).
Haven't kept up on Intel well enough for the last two years or so to be sure, but I seem to remember there being a generation where that happened.
Your argument is invalid.
> Note that n1-standard-1 (the GKE default), g1-small and f1-micro VMs only expose 1 vCPU to the guest environment so there is no need to disable Hyper-Threading.
I'm wondering if they've decided to just eat the loss on the single vCPU nodes, and for vCPUs >= 2, they pass the decision on to the customer.
Or is there someone else's VM running on the other hyperthread?
Here's their guidance on machine types: https://cloud.google.com/compute/docs/machine-types
> For the n1 series of machine types, a vCPU is implemented as a single hardware hyper-thread on one of the available CPU Platforms.
I wonder how much these continuous performance degradations are costing bigger customers, like cloud operators. This shit can’t be cheap.