Intel Has a Big Problem
bloomberg.com
bloomberg.com
Just look at their product release lifecycle: In years past, they'd get maybe one extra product release off each new arch (tick/tock); for example, Sandy Bridge bore Ivy Bridge and Haswell bore Broadwell.
Skylake has born SIX new product lines; Goldmont, Goldmont Plus, Kaby Lake, Kaby Lake Refresh, Coffee Lake, and the upcoming Cannonlake. Their failed 10nm shrink has forced product delays; remember, Cannonlake (the 10nm shrink of Skylake) was supposed to be released in 2016, and its not even out yet. Just at CES this week they said they've shipped mobile Cannonlake CPUs.
They have zero presence in mobile. Their best efforts involve competent Y-series processors. Then Apple comes around and, seemingly without even trying, destroys them [1] with a product that's more thermally efficient and, in some ways, more powerful than Intel's best mobile processors, not just their thermally efficient ones.
They have little presence in HPC/AI, where Nvidia is slaughtering everyone and its not even close.
Its completely inevitable they're going to lose Apple as a customer for consumer products; its just a matter of time. AMD is gaining traction with Zen, and they're moving in the direction enterprise cloud provider want (lots of cores, not much $$). How much longer can Intel keep holding on? Do they have an ace they've been hiding? Will people even trust their ace after Meltdown?
[1] https://9to5mac.com/2017/06/14/ipad-pro-versus-macbook-pro-s...
I would be absolutely shocked if Apple did not already have MacBook-level ARM cores running macOS in their development labs. Just as Apple had OSX running on Intel hardware for years while developing/selling PPC.
I think there's a very good chance we see a MacBook form factor laptop running an ARM chip in the next few years. I'd buy one, having a laptop with multi-day battery life would be awesome. I don't need super performance on my MacBook.
That being said Apple might find x86 emulation useful too.
Intel seems to be way above everyone in pushing performance/TDP. Apple might have comparable CPU's in iPhones, but they are only comparable for first 10 seconds - then they overheat and throttle. If we put good cooling on their CPU's then you have TDP in 15+ watt class which means poor battery life regardless of who made the CPU (Intel or Apple).
The only reason why all of this works with iOS in a first place is super insane power awareness on OS level: as soon as an app is not in view it basically stops almost all work (except for background handlers), and also their browser capabilities are very limited. To achieve 20h+ class of performance on laptops we either must accommodate the same principles (but hey that will probably kill almost any background web apps, like soundcloud) or invent better battery tech :(
Then you have to realize on Intel's roadmap, there is nothing substantial changes in the next two years at those TDP SKUs apart from 10nm. While Apple already has A11x lined up and 7nm from TSMC. First time Apple and Intel SoC has Fab features size parity.
https://www.digitaltrends.com/laptop-reviews/hp-envy-x2-2017...
I think the smartphone market shows that multi-day battery life is not something the average person will trade a decrease in device performance for. And that's on phones, where you also don't have much practical opportunity to use the device plugged in, unlike a macbook. I'd be really surprised to see this tradeoff made on macbooks.
Apple wouldn't use the technology for a multi-day battery life, they'd use it to make the laptop thinner/lighter (they've already started making MBP batteries smaller than they need be, leaving empty space in the case)
Previously [1], from 2012 (starting around t=49:52): "Lots of problems here [at 11nm]. Intel and IBM have publicly discussed solving the problems with 11nm by skipping it", i.e., going straight to 7nm with never-been-used-before EUV lithography.
[1]: https://news.ycombinator.com/item?id=16175949
Were people realistically expecting Intel to hit deadlines at 10nm? This sounded more like a research project than a product line.
As GP mentioned the main problem with 10nm is the photolithographic process. Currently they're using 193nm ArF lasers. EUV "light" is very difficult to generate and handle. There aren't even lenses for it, only lossy mirrors, so each optics step will lose you a fraction of the input power and heat the mirrors.
It's only a matter of time before they do. Nvidia's Drive PX platform is Intel-free and something derived from it might look interesting for HPC and/or high-end workstations in the next few years.
It’s a controller, but if risc-v become established, then the acquired knowledge might be applied in CPU development.
On POWER9, both NVLink 2.0 and OpenCAPI are implemented using the same PHY. However, above that level they're completely separate protocols.
(I work on OpenCAPI drivers and firmware at IBM)
They will fall back on their old tried and tested tactic of anti-competitive practices.
They do have an ace. $17 billion in operating income the prior four quarters. And about $17 billion in cash on hand.
Who knows if they'll put it to appropriate use to recover from their present mess. They very clearly have the resources to do so. Intel's $17 billion in operating income is over 4x the revenue of AMD, and 2x the revenue of nVidia. It's a pretty fantastical premise to be already counting Intel out, they've recovered from dramatically worse business situations in their past.
You know what was missing the last five years for Intel? Desperation. The need to fight back from a situation that threatens their well-being. They've done it before, usually when they were put in a desperate situation. That's also typical of human behavior in general (which is where the corporate behavior derives from).
- they need to capture/create product niche with volume of sales in XX billions per quarter, otherwise it is not worth the effort
- in their traditional niche they have 90+%, the only way is down, basically
we'll see, but signs are not good. F.e., this idiot (https://www.fool.com/investing/2018/01/19/intel-corps-upcomi...) is hyping some 6/12 cores notebook CPU, while for any reasonable analyst this should be sign of desperation - Intel does NOT KNOW WHAT TO DO WITH SILICONE, so it's just slapping more cores together
Microsoft has been reinventing themselves; partly successfully, partly failing (Windows Phone / Nokia acquisition was arguably a recent example of a failure in this regard). Windows and Office are still the primary income for Microsoft, but cloud has gained traction, percentage-wise. Xbox is profitable (it wasn't for a long time). The Surface line is another example of Microsoft reinventing themselves.
> IBM should not exist at all
IBM has been relatively marginalised, and could've been so much more rich if the IBM-PC didn't backfire them with the clones getting popular, plus them getting tricked by Bill Gates with MSDOS and Windows (the failure to collaborate with Microsoft on OS/2). But as the B in IBM suggests, IBM never cared much about the consumer space.
Note, I'm not sure about POWER, how competitive that is with Intel's high-end products. Perhaps someone can comment on that.
I think what a company like Intel needs (or any company, really) is direct competition. Ie they need AMD. The competition with ARM and such is far more vague, less opaque. To be fair, Intel did try to compete on low end, with Atom. I remember they were busy with Moblin/MeeGo in the mid '00s. Never took off AFAIK.
Such as?
This situation may become worse for them over time due to the attitudes of their partners in response to this fiasco, but as it stands (AT THIS POINT IN TIME) I think the FDIV bug was probably worse financially.
That said : I hope it hurts them just to help turn the industry away from such single-vendor dependence.
I lived through both of those, they did not come even close to the sense of brand damage that I have right now regarding how I personally perceive Intel, they have mis-handled this from day #1 and they are not doing much better than when it started.
https://www.realworldtech.com/intel-dram/
Here's a pretty good cover of it from 2012 by NPR, interviewing Grove, in a series they did:
https://www.npr.org/2012/04/06/150057676/intel-legends-moore...
Aside from those, I added a link to Intel i960 that spun out of the BiiN project since it was a nice, little RISC that you might like. It had object-based security and fault-tolerance built into it. There was still potential for reviving that before RISC-V ecosystem, CHERI, and so on obsoleted it. Too bad.
https://en.wikipedia.org/wiki/Intel_iAPX_432
https://en.wikipedia.org/wiki/BiiN
https://en.wikipedia.org/wiki/Pentium_FDIV_bug
Huh? The Xeon Phi is used regularly in the industry. Nvidia is dominant but to call Intel's presence small is very odd.
The memory bandwidth is just crushingly bad.
This is the most important thing. AMD has already made the switch to Multi Chip Modules which makes it much much easier to produce chips for 10nm (what TSMC/GF/Samsung call 7nm).
Right now Intel cant even make a dual core low speed mobile chip on 10nm. How are they going to make a giant 30+ core server processor? This is extremely bad news for them that has not been fully realized by the markets because they have faith that Intel will figure it out, but they may not.
AMD may ship 7/10nm server chips before Intel and Intel may never ship them before switching to MCM themselves.
MCM is as revolutionary as AMD64 was but most people dont realize yet how important it is and how much of an advantage AMD has because of leading with it.
What’s the reason for this? From my understanding of numbers, 10 is not 7
From https://en.wikipedia.org/wiki/7_nanometer#7_nm_process_nodes
> The naming of process nodes by different major manufacturers (TSMC, Intel, Samsung, GlobalFoundries) is partially marketing driven and not directly related to any measurable distance on a chip – for example TSMC's 7 nm node is similar in some key dimensions to Intel's 10 nm node.
I'm led to believe that if we're going to recompile The World as well as build new JIT javascript++ compilers, it will be for ARM64?
Neither. Multi-chip modules simply make it more economical to manufacture large processors. When you can make a 32-core CPU by connecting two 16-core silicon dies together, you'll have higher yields than if you try to make a monolithic 32-core die. There isn't necessarily any consequence visible to software, with the possible exception of having multiple NUMA nodes per physical socket.
Once Intel's fab advantage is actually eliminated, they'll be forced to make multi-die CPUs to compete against AMD's multi-die CPUs. But every indication is that Intel will be ready to do multi-die better than AMD is doing multi-die. How much better is up for some debate, but it does look like Intel will have a leg up on EPYC's biggest weakness.
The advantage of this architecture is that smaller chips are easier to manufacture because the probability of a random defect rendering the chip useless is lower and when a chip is bad it is not as great of a loss.
Epyc chips run all existing software. There is extra work involved with Non-Uniform Memory Access (NUMA) where if you want software to actually use 32 or 64 cores simultaneously it has to be purposefully designed to take into account the architecture but that is not really any different than developing for multi-processor systems in the past.
The point is that we have hit the ceiling on frequency scaling, which means more cores, and now we have hit the ceiling on core count scaling, so now it is going to be more CPUs, and the best way to do more CPUs is with MCMs.
When Intel gets around to making a multi-die CPU using EMIB, it'll blow away AMD's EPYC in terms of bandwidth between dies.
[0] https://www.intel.com/content/www/us/en/foundry/emib.html
[1] https://www.anandtech.com/show/12003/intel-to-create-new-8th...
It is Intel's CPU using AMD's GPU. I am not sure how this is evidence that EMIB is superior tech. Of course Intel is going to use their own tech wherever possible.
> When Intel gets around to making a multi-die CPU using EMIB, it'll blow away AMD's EPYC in terms of bandwidth between dies.
Do you have any bandwidth numbers on EMIB to support this? There are other sources saying that EMIB is crap and doesn't interface with HBM without using a bridge provided by TSMC. [1]
[1] https://semiaccurate.com/2017/12/19/intels-claims-fpgas-hbm-...
That article doesn't say EMIB is broken, it just says that the FPGA product Intel/Altera announced wasn't actually natively designed to use EMIB to interface with HBM. This is not relevant to a discussion of whether two Intel CPU dies could be connected over EMIB.
We do know that the Intel CPU with an AMD GPU on-package will have a much faster interconnect between the dGPU and its HBM2 over EMIB than AMD's EPYC and Threadrippper manage between CPU dies with a conventional MCM. It's also faster than Intel's old Crystalwell parts that used a conventional MCM to add an eDRAM L4 cache to a consumer CPU.
The limits of conventional MCM packaging are well-known across the industry and are a problem for everyone who's trying to use HBM-style DRAM or otherwise make high-speed inter-die links. Full-scale silicon interposers with TSVs are really expensive and would not have been adopted in high-end GPUs and FPGAs over the past several years if there was an alternative. Likewise, AMD wouldn't have adopted EMIB for their project with Intel if conventional MCM packaging were sufficient, because EMIB is still more expensive than not using two kinds of interconnects within the package.
Once again, you are talking about an Intel CPU, so what you are saying makes no sense. It is not AMD's project. It is Intel's CPU that happens to use an AMD component.
The only EMIB in that product is between AMD's GPU and that GPU's DRAM; the Intel CPU die in that product is not using EMIB. The fact that the AMD GPU is using EMIB is significant—it's a somewhat external validation that EMIB has value as an alternative to conventional MCM packaging or large scale silicon interposers, both of which AMD uses for products of their own.
If EMIB wasn't saving money while offering sufficient performance, then Intel would have integrated the AMD GPU into this product using the off-the-shelf silicon interposers that AMD's own products use to connect the GPU to its HBM DRAM.
Correct. There's a conventional link through the package substrate carrying the PCIe signals from the CPU to the GPU. The only EMIB on that product is between the GPU and its DRAM. That's why the GPU and its DRAM are adjacent, while the CPU is about 1cm away from them: https://images.anandtech.com/doci/12003/intel-8th-gen-cpu-di...
Intel would need to redesign the CPU die to accommodate an EMIB connection to the CPU. As it stands, an EMIB link would probably get in the way of lots of contacts that need to go to something other than the GPU, and EMIB wouldn't be worth the trouble for a relatively narrow and slow link like PCIe. The point of EMIB is to enable very wide and fast links between dies, so that inter-die communication is almost as fast as communication across a single die.
That is a nice claim, based on no evidence.
Intel spends billions subsidizing non-competitive products in order to try to build market share all the time. (mobile, ultrabook, etc, etc)
Unless you can present any evidence to the contrary (still no bandwidth numbers) it is just as likely that this product has EMIB because they couldn't get anyone else to use it and they need some shipping products with the tech to try and sell it.
The concept of market share does not apply to EMIB. It's an internal implementation detail, and Intel gains nothing from using EMIB in one of their own products over a cheaper solution that would still result in the end product being an Intel-branded part.
The only motivation that Intel could possibly have to use EMIB as any kind of loss-leader would be to entice third-party foundry customers to use Intel for the sake of EMIB. But attracting foundry customers is clearly not a priority for Intel, and Intel's primary goal with EMIB is to produce Intel products that may incorporate some third-party silicon, not to produce silicon to be incorporated into third-party products.
Fantastic claims require fantastic evidence.. history shows that there are many ways of fixing or designing-around manufacturing yield problems. Do you have any actual evidence that their 10nm fabs will never yield a competitive monolithic server part in a financially-viable way?
> MCM is as revolutionary as AMD64
I would disagree somewhat, as AMD64 clearly benefits every customer with very little downside, while it is fair to say that MCM CPUs can create or exacerbate irritating performance problems for the customer by being "much more NUMAed". [All else equal, most customers would prefer a monolithic part, or 2 NUMA domains instead of 8-16].
Because our engine took advantage of memory-mapped unnamed file handles to cache frames, we didn't really need the extra address space of larger pointers. I was able to exhaust system memory by manually managing whether or not buffers were mapped into our address space.
Though it had been made for something else, I lucked into being able to use that same engine for interprocess legacy plug-in support. AMD64 support also wasn't necessary to get more physical memory support out of the chips of the era. Operating systems supported 36 and 40 bit physical address spaces, and applications could use higher (would-be negative) addresses by indicating support for it. Those wanting to really cut loose could turn to manual mapping. (See: PAE, LAA)
In the end, we shipped AMD64 support because there were substantial performance wins. And, yes, I profiled every possible software/hardware combination to determine where gains and losses were coming from.
I'm not saying that things were universally faster/better. RIP-relative addressing, though generally quicker than absolute 64-bit, is a headache, and the loss of 80-bit floating point intermediates broke some rare code (which would have broken in strict mode, anyway). I'd love to see some evidence of slow-down now, but, until then, I'm unconvinced. I had a dev bring me "evidence" of huge performance differences between ARMv7 and ARMv8 a couple of years ago (and there are some if you dig). I asked him if he had compiled with or without thumb in the linked modules.
He let me know of his new results a few days later, a little red-faced.
I'm not saying that your detection of slowness is the same thing, but many aspects of performance are measurable (and measured). If it's slower, we can probably measure just how much.
MCM has been relatively unpopular because of cost and that's why most of its applications could justify it only to put chips from different silicon proceseses in same package (cache or EDRAM + CPU usually). Except for some high end chips like the POWER stuff.
The interesting bit is how and why AMD put logic chips in low cost MCMs. Probably involves the time or $$ budget to make and validate additional high end variations of the silicon.
Furthermore, the site you link to, which is clearly rather gung-ho about macs, still shows that the ipad was not able to keep up with the intel chips, not even in multithreaded mode, even though the intel has 2 cores to the a10x's 3+3. The ipad beat the MBP on the GPU tests, and I don't think anybody is going to dispute that intel's gpu's are... not exactly record-breaking.
Still, you're totally right that apple demonstrated that intel's tech isn't the uniquely fast chip you might think it to be! It's not a small achievement to come this close, especially if the ipad is using less power (which isn't actually clear - battery capacity is roughly similar, at any rate).
* The reason I don't like geekbench as a cpu test is that it includes a great many tests like FFTs, gaussian blurs, jpeg with DCTs, image processing, and crypto (even in the noncrypto segment!) etc - i.e. workloads that are very, very vectorizable. The problem with that is that such workloads are really tricky to get right - small differences in code can make considerable difference in perf, and what's right for one CPU isn't for another (notably, geekbench is necessarily running completely different code on these processors). Secondly, those kinds of workloads are very, very amenable to special-purpose instructions, which leads me to the related point that I don't think they're representative of what is typically annoying slow today, and much less what is likely to matter tomorrow: these kind of workloads are either good enough on the CPU as-is, or they're going to be moved to the GPU (or even more specialized hardware), which happens to excel at that.
> Skylake has born SIX new product lines; Goldmont, Goldmont Plus
Goldmont, Goldmont Plus are Atom products and are not "Skylake-derived" in any really major sense.
> zero presence in mobile
This is true only if you disregard the traditional laptop and newer convertible and chromebook markets. More importantly, Intel has an admittedly-indirect but financially very significant presence in the mostly-ARM cell phone and tablet markets, in that the backend for pretty much every cellphone app runs on an x86_64 server core somewhere. Every cell phone and tablet sold helps sell a modest fraction of an x86_64 core too.
> little presence in HPC/AI
Again, every major HPC/AI deployment that I've heard of still has a huge number of Intel cores. Though - as with the cell phone and tablet markets - you can definitely fault Intel for not capturing these markets in their entirety; they certainly had the talent, IP and the capital to do so, but they screwed up strategically in a way that will almost certainly become a cliched business school study at some point, if it's not already.
Intel is not the entirety of x86_64 servers, and as the GP notes, AMD is going in the better direction with Zen for serving all those requests than Intel is at the moment.
As an aside, because it doesn't really matter to the point, x86_64 is actually an AMD extension that Intel licenses from them.[1]
The way Intel tries to fight back I find questionable. Both their MIC lineups and many-core Xeons suffer from a crucial flaw: Almost laughably low memory bandwidth when compared to NVIDIA, even compared to Epic. They rely so heavily on caching that it would require tremendous software efforts to get even in the range of 50% of the performance of GPU ports - I‘d argue it‘s significantly more difficult and less performance portable to do that rather than just do a basic GPU implementation, even with manually handled data transfers.
Add to this the increasingly successful efforts by GPU makers to automate data handling and Intel‘s attractiveness only dwindles further.
And now people are told that they might loose 20% performance? There‘s applications that use these chips like it were real-time systems (because it‘s the only thing financially possible for them). For these use cases, 20% will hurt, a lot.
/rant
In any case it would IMO be better to just now license 3D memory at market cost rather than just sitting on the hands and let HPC markets fade away.
Remember, it has historically only taken 8-10 years for HPC developments to scale down to people's pockets - it could well be that the next iPads outperform what Intel can offer as x86 based laptops, which will not be a good look for software developers deciding which platform they should choose for the next potential killer app.
[1] http://www.wipo.int/edocs/plrdocs/en/lexinnova_plr_3d_stacke...
You mean RDRAM?
Applications that use these chips in a realtime capacity don’t share hardware with other apps, so it will be possible to deploy them on unpatched machines, even if inconvenient. Those workloads are also not big on context switching and so not affected much by the patch.
I think cloud hosting providers will hurt the most, because that tends to be a context-switch-heavy workload, and their pricing model assumes a certain level of performance per core. They’ll mix in AMD in significant quantities just to diversify and avoid getting burned in the same way. They’re probably also the biggest market for intel’s server line.
This incident will seriously hurt intel even if they handle it perfectly, which they’re not doing.
Right, you can very easily imagine a machine whose CPU merely marshals and dispatched work to compute elements, be they GPUs, FPGAs, ASICs, TPUs, whatever. You don’t need an all-singing all-dancing Xeon for that...
Intel has moved away from the tick/tock release cycle. They now work to what they call process/architecture/optimization, or in more familiar unofficial terms tick/tock/tweak.
Moore's law is essentially broken now, and this isn't just an Intel problem, this is a problem for all the major chip manufacturers. Furthermore, the challenges and cost of future node shrinks is also causing a slowdown in progress in integrated circuit manufacturing. We may get a couple more node shrinks, but we should prepare ourselves for the brick wall that we're likely to hit in the next decade.
For this cycle it's been tick/tock/tweak/tweak. Sounds like a broken clock.
They switched to git because free is better (we were a rounding error on a rounding in terms of cost to Intel, the most they ever paid us in a year is .00000004 of their revenue).
But that was too much so they switched to git and it's been downhill for them ever since.
I'm not an idiot, I don't think that the switch to git is the cause of their problems, their problems are self inflicted. I'm just one of many many vendors that Intel has fucked over. So I like seeing them squirm.
Karma is a bitch Intel.
Edit: yup, knew I'd get down voted. Don't care. Try being an Intel vendor and get back to me about how much you like that.
However, as poorly as they treat their vendors, it's nothing compared to what they do to their employees and sadly large segments of their first line and middle management not only buy into that but have spent time making that into an art form.
If you are an Intel vendor you will like this tidbit, they were on our paper. Not their agreement, they agreed to our paper. I think we are the only small vendor that they ever did business with where they didn't force their terms into the deal.
For this cycle the 'tick' (Cannon Lake) was pushed back a year (with two 'tweaks', Kaby Lake Refresh and Coffee Lake, taking its place). Cannon Lake devices are due to be released to the mass market this year, and according to Intel it was able to manufacture some Cannon Lake devices last year:
https://en.wikipedia.org/wiki/Cannon_Lake_(microarchitecture...
"At CES 2018 Intel announced that they had started shipping mobile Cannon Lake CPUs at the end of 2017 and that they would ramp up production in 2018."
Many industries are held back by lack of interoperability, rather than raw computer power.
There are performance ceilings on those as well. Best case scenario we come up with a new architecture that is more efficient than the current ones (something built around reconfigurable chips could be ideal). However, we will eventually reach a point where computers don't get significantly faster. Maybe it takes 20 years, maybe it takes 30 years, but there will come a point where that happens.
Help me understand why. Even if individual computational units cannot get faster, can we not benefit from additional computational units? I know that these don't scale linearly but I suspect there remains work that can be done to improve scaling. Isn't that where GPUs get their computational power?
The most significant cost in making an ASIC is the silicon wafer, which is extremely expensive, so anything that uses it more efficiently [ less space] makes things cheaper, faster (easier to keep a faster clock synced over a small area), and use less power (less power means you can go faster too, because you have more power/head headroom for cranking clock speeds).
Scaling by going smaller is extremely synergistic, and has super exponential performance impact, whereas in the very best case adding more cores is linear, and in almost every real world case, sub linear. It also costs more, not less.
The fact is baring major breakthroughs, we are stuck with roughly today's level of performance for the foreseeable future.
Maybe finally we stop wasting hardware and use the supercomputers we have in our pockets for greater things.
Another thing to consider is that memory access is still the most prevalent speed problem. If we can work on these bottlenecks we can still get significant performance increases without actually increasing CPU performance that much.
> The fact is baring major breakthroughs, we are stuck with roughly today's level of performance for the foreseeable future.
That's obviously false. Otherwise graphics wouldn't get any faster when you add more GPU cores, to name one common embarrassingly parallel problem.
Our software is just really bad at making use of those extra cores. Measure real-world use cases with a well-tuned work stealing engine and you'll see how much performance we're leaving on the table.
(Source: I wrote the first version of possibly the largest consumer deployment of a major software component backed by a work stealing engine.)
GPUs get their power from being able to execute multiple calculations in parallel. Rendering a 3D scene is something that lends itself to being processed in parallel, for example on a simple level you can have different cores each rendering a different section of the overall image (tile-based rendering is one example of an approach that benefits from this: https://en.wikipedia.org/wiki/Tiled_rendering ).
The GPGPU uses of GPUs (i.e. non-graphics uses) also take advantage of this parallel processing power.
The issue is, not all computing workloads are easy to split up into smaller workloads that can run in parallel. Some workloads are easiest to manage sequentially. For example, consider some code that had a lot of conditional logic (such as "if" statements), where the code executed depended on validating a condition was met. What advantages would you gain from running this code in parallel?
They could have been saved, and still might.
I gave them more than they deserved, better than they could have asked and they turned it to crap.
Pundits want to say that there's a huge thing here, because pundits don't optimize for the truth. They optimize for clicks. So you really need to be careful looking to their writing for the truth.
I think the major difference between this and, say, the Equifax blowup, is that Intel's institutional clients are affected by this.
I'm not sure what they're thinking internally, but it stands to reason that they're probably a bit upset at least: Their CapEx just went up to maintain the same level of computing power. I'd be surprised if internally Google is buying the "AMD is just as affected" line that Intel's been throwing out.
So, I wouldn't be surprised if they're at least evaluating AMD.
Or, again, maybe Intel just totally has them over a barrel and transitioning isn't feasible at all. It certainly doesn't paint a great picture of Intel's future if AMD does catch up, though.
I am deeply skeptical of the commentary on Intel's attitude and press releases. I really doubt that matters much to most buyers.
Correct me if I'm wrong: It's been mitigated by applying a patch that has fairly severe performance implications, no? How does this not affect the institutional clients' bottom line in that case?
1 - 30% impact range for best / worst case. On the kernel mitigations. So really workload dependant.
Then there are companies like Epic Games who reported horrific numbers. It seems if you do lots of simple communications (eg websockets or UDP), you can expect a huge slowdown.
https://www.epicgames.com/fortnite/forums/news/announcements...
First, ARM is doing to Intel what Intel did to the Unix workstation vendors in the 80’s.
Second, given that they’re being cornered into the server business, they need to have products that are rock solid there until they can regroup. This is one of a long parade of recent screwups with their big bets in this space:
(1) A while back, all their server atom chips (tons of crypto and I/O with piles of ECC DRAM and cores for < $1000 and < 20W) had a bug where they stopped booting af 18 months of uptime. These compete exactly in the space server-ARM has a chance, so many affected vendors were already dual sourcing.
(2) NVIDIA crushes them for AI, and Intel is a distant third for graphics in general
(3) Samsung SSDs generally trounce Intel ones.
(4) They’re rapidly losing client device share. Their big recent innovation there is AMT, which is increasingly considered an anti-feature.
That leaves conventional IT compute, (web services, DBMS, etc) for their core business, but even on-prem stuff is moving to private cloud, which needs multi tenancy, and they’re looking pretty risk for that use case too (vs AMD?)
They’ll certainly be around for a long time, but it’s not clear how long they’ll keep their “no one gets fired for buying IBM”-level of dominance.
This is true only for the consumer SSD market, where Intel outsources large portions of the product development. It's also a market that Intel may abandon completely in the next few years as Intel and Micron start to pursue separate flash memory development. If Intel doesn't score a solid win with a consumer SSD in the next two generations, it would be reasonable for them to pull out and focus solely on enterprise SSDs, where they have no trouble winning.
They don't have to buy it. Whether affected or not, AMD is a non starter at this moment for those things.
Speaking of Dell, they are launching some EPYC stuff: https://blog.dellemc.com/en-us/poweredge-servers-amd-epyc-pr...
But again, getting the ball rolling might take a couple of years. Look at what happened with Opteron as an example.
That was my hypothesis for why their stock didn't drop much. The problem is very bad but intel's quasi-monopoly and the very high switching costs involved will let them weather it.
For this incident, I got an email a few hours after the embargo was lifted, that essentially said that it was no big deal and referenced public information. The purpose of the communication was to have people like me message up the chain that this was no big deal. That misdirection is inexcusable, particularly when they could have given meaningful guidance under NDA.
We had some follow up questions, which weren’t really answered. We were directed to hardware OEMs, as ETA for microcode updates are out of their control and according to Intel are the full responsibility of the OEM. In reality, Intel was struggling to deliver the code, and the OEMs we deal with issued patches in hours, and had to pull back updates due to Intel code revisions.
Personally, I do have alternatives for strategic parts of the business that drive high margin Intel sales. Many critical aspects of my business can run on Intel or Power platforms, and we can engineer solutions either way in similar cost footprints.
Less strategic aspects of the business, like end user compute now have niche competitors that can gobble up Intel business very quickly. Half of my desktop users run on VDI, mostly with AMD thin clients. 50% of my constituencies can run their core line of business functions on iOS. iPad with a keyboard could reduce my Intel desktop spend by 50-75% for 2-3 years.
https://twitter.com/avtargill/status/951195158229463046?s=17
At this point it’s up to users to scour motherboard manufacturers’ clunky forums to determine if their platforms will ever receive patches. Given the severity of the issue this really should be handled with more accountability and with a greater sense of urgency.
It’s particularly obnoxious when you read about the heroic efforts that were put in place to put AWS, Azure and GCP right. If it was no big deal, why go through that?
I’ve read Andy Grove’s account of the thinking behind the response to the Pentium math bug. I expect better from Intel. As a customer somewhere between an individual PC builder and Amazon web services, I don’t think they handled this incident well at all.
The only useful thing in these documents was a timeline/detailed list for the microcode patches, all of which should be public.
They also claim that Spectre/Meltdown are "not a bug or flaw in Intel products" and their slide deck has a whole slide dedicated to forward-looking statement disclaimers. Sigh.
Needless to say, we're not impressed.
But perhaps this was no big deal. We've seen years of research suggesting that modern CPUs are full of issues like this. There's probably a good decade worth of papers on cache side channel attacks. See https://eprint.iacr.org/2013/448.pdf for example
Perhaps the big deal is that there are still people who think they can safely run multiple different things on a single machine?
Both attacks, in their practical form, use cache timing as the side-channel to extract information. But the surprise is the control over (as the paper calls them) 'transient executions'.
like check this out: https://www.tau.ac.il/~tromer/acoustic/
> Here, we describe a new acoustic cryptanalysis key extraction attack, applicable to GnuPG's current implementation of RSA. The attack can extract full 4096-bit RSA decryption keys from laptop computers (of various models), within an hour, using the sound generated by the computer during the decryption of some chosen ciphertexts.
So the seriousness of a side-channel attack is determined as a function of the impact of the exploit as well as how easy it is to carry out the exploit.
Meltdown in particular is nasty because it is relatively easy exploit and is undetectable and affects a ridiculously large range of hardware. So it is actually a pretty big deal. Yes there are side-channel attacks against Intel CPU's, but this isn't just any old side-channel attack.
Here's just one of the more practical attacks http://palms.ee.princeton.edu/system/files/SP_vfinal.pdf
>but this isn't just any old side-channel attack.
It isn't, but you were already screwed. Now you're just slightly more screwed.
Not some artificial poc that already knows an address to attack and needs to be helped by continually pulling the data into the L1 cache.
If it is so easy to do then why has nobody written anything that can read a password from a browser or sudo?
We might instead point to facts. Major tech companies are getting into chip design. Apple's foray into fabless last year almost destroyed the value of Imagine and Dialog shares. If they and other companies are successful, we can see the same happen to Intel.
The company has no competitor in server chips at the moment, but this episode could change that. Microsoft and Google have publicly praised Qualcomm Inc.’s first server chip, which went on sale in November, and Apple, Google, Microsoft, Amazon, and Facebook all have internal divisions working on chip designs.
AMD people should be proud, as customers we're really happy. I hope that GCP would have EPYC-based platform at some point too.
In two weeks, it's Dell. R6415, R7415, R7425.
It's way, way, way too soon to judge the long-term implications of Meltdown and Spectre on Intel. If their clients want to switch, it'll take months and years to do that. That doesn't mean they won't do it, but it means we won't really know the full extent for a while.
The stock price is a really crude metric. For judging the long-term implications of this, we can't look at how the stock has performed in the last two weeks alone and extract any meaningful information.
My prediction: nobody big is going to switch away from Intel entirely, but they will start to prioritize investments in technology built on its competitors, as a way to hedge their future risk. That's definitely bad for Intel, because over time, it'll reduce their lock-in.
The reason people are saying Intel stock isn't down is because of that, when what HN considers the worst thing ever is indistinguishable from fluctuations for the last 3 months the market doesn't think it's a big deal.
Of course, the stock market by no means knows everything. But the aggregate prediction of traders is that this doesn't matter very much to the bottom line, and I tend to agree.
For example a while back amd announced in an earnings call: we are in the black and reduced out debt subtantially. The stock tanks by 15-20 percent
In the longer term (1-3 years) Intel has a very large problem. Especially as they are being eaten alive in non-desktop class cpus right now.
These issues may not be a knockout punch, or even have them on the ropes, but it made them stumble and they look vulnerable.
Botched micro code updates, Intel engineers arguing with each other on the linux kernel mailing list, etc. Intel's best and brightest have had 6+ months to work on proper mitigations in secret and this is the result.
https://marc.info/?l=linux-kernel&m=151559244214217&w=2 https://marc.info/?l=linux-kernel&m=151559367514704&w=2
Yeah, handling of mitigations for Spectre has been awful during the embargo period, but I am very happy with the result we're getting now. It's taking less than three weeks to get everything sorted out.
It's down several percent since the announcement in a rising market.
More generally they have underperformed the SP500 over the past 2 years and AMD in particular is blowing them away.
Yes they have a problem. The current CEO seems more interested in politics than technology. http://www.breitbart.com/big-government/2015/09/10/intel-cut... (inb4 I don't like Breitbart)
Intel has as much of a problem as VW had after diesel-gate: none. Same will be valid for Apple's throttling scandal.
Big corps like this may experience some little storms here and there, but there is no iceberg big enough for them.
Articles like this exist just for the sake of writing something and making some money.
The article itself states: "So far, Meltdown and Spectre probably pose less risk to the average person than, say, a simple phishing attack in which a hacker tries to send you to a malicious website. They won’t lead to the kind of widespread panic that resulted from the 2017 hack of Equifax’s customer database.
But that could change. Hackers who hadn’t tried to break into Intel’s hardware, believing there was no way it would leave a side door open, are now seeking ways in."
"But that could change" is a vague term that doesn't mean much to me. Again - not trivialising this nor saying Intel shouldn't do some soul searching, but I'd like to better understand the justification for the apparent hysteria - unless, of course some people more experienced in security would care to explain what it is I'm missing here.
P.S. One thing I wanted to check was whether Spectre/Meltdown breaches could somehow be caused by manipulating a web browser. Some searching revealed that this is indeed a possibility so at this point, everyone feel free to panic :-)
What I see happened to Intel is that once they consolidated their monopoly in the late 2000s, they lost the healthy management practices that tend to come from being in a competitive industry.
All this talk from upper management about velocity was about trying to find a way to make more money when you've mined out your current niche completely. It ended up instead opening the door for AMD to make a comeback on x86
This way they can pretend its an industry wide problem and they are not to blame, great.
It was Google who found the exploits, and Google who published the exploits in the same document & their press release.
Google is the one who packaged these exploits together. How does Intel PR take credit for this?
They weren't though. The technical publication by project zero [0] did distinct between the two, but Google's PR oriented article [1] didn't bother to do that. The PR oriented article packages the two exploits together saying "These vulnerabilities affect many CPUs, including those from AMD, ARM, and Intel, as well as the devices and operating systems running on them."
[0] https://googleprojectzero.blogspot.com/2018/01/reading-privi...
[1] https://security.googleblog.com/2018/01/todays-cpu-vulnerabi...
Instead no one seems to grasp what meltdown is, intel is feeding the media with benchmarks and says there is no real impact, users believe intel.
I read somewhere that all big companies Google, Intel etc. knew about these security vulnerabilities, especially meltdown months before. Intel made a deal with Microsoft to disclose the vulnerability on a Microsoft Update date and tell everyone that there was an issue and it was fixed. Microsoft was also supposed to make all processors including AMD to get affected by the slowdown. However Google reported 1 week earlier said it couldn't keep this a secret with good conscience. But even if that plan failed there is still a lot of misinformation caused partly by intel.
Even if you visit sites such as meltdownattack.com (first result on google with some deeper info), it says we do not know in which aspect Meltdown affects processors.
I'm on the verge of buying a new desktop, I will go with intel because I have no other cheap choice. Intel still delivers the cheapest option for me (i3-8100). But I would appreciate if Intel played fairly. All this manipulation behind the scenes is corrupting the market and intel is responsible for it.
I don't know what we got ourselves into. The RAM folks built a cartel and are keeping the prices high. Microsoft is deliberately crippling third party antivirus software. Intel is shipping us CPUs with backdoors (Intel ME), there are serious vulnerabilities that get swept under the rug, yet we have to buy intel because we are locked into the x86 platform.
At least microsoft is teaming up with Qualcomm and Apple is rolling out its own ARM processors. We need to have an alternative. IMHO Intel has been playing an unfair and an unethical game.
Is there no working AMD exploit by design, or did AMD just get lucky? I bet they just got lucky.
It is also due to memory access checking and out-of-order execution, since those aren't orthogonal to "speculative execution" (and in fact are tightly related).
I can't speak to your computer architecture course and I'm not sure what literature you've read (but it's easy to get the impression that it only relates to branches since a lot of literature might only be addressing that aspect), but the Meltdown authors are clear at least (quoting from the pdf):
In practice, CPUs supporting out-of-order execution support running operations speculatively to the extent that the processor’s out-of-order logic processes instructions before the CPU is certain whether the instruction will be needed and committed.
They go on to note that for the remainder of this paper their use of the term will refer to a more restricted definition related specifically to branch speculation, since that's what they care about. That's fine, and it's good they are clear about it - but it doesn't change the recognized meaning of the term (and indeed their narrowing of the term helps confirms the general definition).
They aren't as clear in the Spectre paper, and they focus on branch-related speculative execution since that's what they care about for the purposes of their description, but they don't contradict the idea that speculative execution is limited to branch prediction.
Not that the Spectre/Meltdown authors are a particularly authoritative reference for CPU architecture terminology: these are, after all, software guys peeking into the hardware world for the purpose of putting together these attacks.
Modern "big" cores are, conceptually, executing most of their instructions speculatively, since wide out-of-order execution windows that that there is a large-degree of divergence from pure in-order and any time any earlier instruction can fault, the remaining instructions are speculative (and the CPU mostly doesn't care: the infrastructure such as the ROB are going to be used regardless of whether the current head of the instruction stream is speculatively or not).
The crux of it comes down to how their TLB, L1 and out-of-order engine interact. Can a load Y whose address depends on an earlier load X whose entry exists in the TLB but only with the S-bit (supervisor access only) end up actually using the true value of X before it is squashed?
Clearly Intel allows it to occur, but it's not obvious that this has to be the way. The TLB, which presumably contains the S-bit, is already on the critical load-to-load dependency path (since the physical tag needs to be used to select the way) so it isn't obvious that the S-check can't happen at the same time and essentially squash the result of the load X before it ever shows up on the bypass network for consumption by Y. On the other hand, it might be slightly easier to let the X result show up up on the bypass network and do the S-check in parallel and then only flag the ROB-entry as squashed after the fact[1] - perhaps it saves a MUX in the critical path.
This is very much like unlike Spectre, where it is more or less obvious by the way that almost everyone does branch prediction and speculative execution that you can probably pull of the attacks on modern cores (perhaps with the exception of branch predictors that do a full address check to access the BTB). Here AMD has also tried to claim to some invulnerability, but IMO these claims are much weaker since it seems unlikely that the predictors cannot be trained. AMD is probably just relying on the fact the the predictors are harder to train.
[1] It's worth noting that everyone is saying that Intel only applies the security check at retirement - but we don't really know this: we only know they apply it "too late" in that the Y load can consume the result, but it could still be applied before retirement, as little as a few cycles later.
Actually not. The toy example shows that speculative execution occurs past a fault. Every big CPU does that and it's not news (and it's the basis for the Spectre stuff, except with "branch" replaced by "fault"). It should probably be in the Spectre paper too or instead of the Meltdown one. This also isn't news to anybody (if CPUs couldn't execute past faulting instructions, OoO would be almost useless).
The really specific interesting thing that makes Meltdown work is that supervisor-only flagged lines can be immediately used by a subsequent dependent load. That could entirely be specific to a particular architecture decision, and so it could be that some designs are immune while still being highly speculative in the general sense.
I'm not saying that AMD had some brilliant foresight to do it this way, but it might have been a design decision for unrelated reasons or just fallen naturally out of other constraints.
As I mentioned in another reply to the OP, there is some small advantage in delaying the squash of the first load (slightly reduced complexity on the L1-load-hit path), but maybe there are some small advantages to doing the early squash too (you don't use the "wrong" value in subsequent instructions and so avoid wrongly evicting useful lines from the cache, etc).
I think we can safely assume that AMD wasn't like "Oh, we should squash disallowed loads early to avoid cache timing side-channels" - not because they thought of it (they should have), but because they would have already tested Intel chips for it and published for the huge PR boost it would give them.
Why is it that if the disclosure (after it was found by someone else) was scheduled for early January that everyone was still not ready and the patches still aren't released and they are still buggy and broken?
It is clear from the fact that Google broke their own disclosure rules that they were colluding with Intel to a certain extent to manage the fallout from this.
Google had an enormous self-interest to avoid disclosure. At the 90 day disclosure date, their Google Cloud product was vulnerable to all three exploit variants. From their blog post, it looks like at that point they didn't have acceptable mitigations for variant 2. It took Google until December, or nearly 3 months after the projected disclosure date, to fully patch their Cloud product. Google only announced the exploits to the public after they had mitigations in the pipeline for all their products.
Reviewing the timeline, isn't it more convincing that Google tried to protect themselves first, rather than Intel?
Furthermore, if Google really cared so much about 'colluding with Intel' and having disclosures 'waived for Intel', why did they announce these exploits before Intel could release a microcode patch for their CPUs? It does not add up.
It is true that both would suffer from disclosure.
> why did they announce these exploits
Because they got outed and so they had to develop a cover story.
> It does not add up.
Many parts of the story do not add up, probably because the parties involved are lying.
Your claim is at odds with Google's explanation of why they announced the exploits prior to the coordinated disclosure date. In their press article, they state:
We are posting before an originally coordinated disclosure date of January 9, 2018 because of existing public reports and growing speculation in the press and security research community about the issue, which raises the risk of exploitation. The full Project Zero report is forthcoming (update: this has been published; see above).
Recall Google's announcement came the day after multiple reports speculating the existence of a CPU hardware bug. See https://news.ycombinator.com/item?id=16046636 and https://news.ycombinator.com/item?id=16052451 .
Google's statement on why they announced before the coordinated disclosure date is valid. There was an unprecedented amount of embargo violations prior to the coordinated disclosure date. There were multiple high-profile reports correctly speculating the nature of the vulnerabilities. Where is your evidence this is a 'cover story'? I can't see it.
> probably because the parties involved are lying.
It is trivial to accuse someone of lying. Specifically, what did Google lie about, and where did they lie about it? Most importantly, where is your evidence that Google is lying?
And Intel did not have a saying in packaging both vulnerabilities together. Researchers at Project Zero did, and with good reason, since both exploit the same function through different means.
Datacenters should probably be worried, but what about the hundreds of millions of users out there? Doesn't seem like a big deal, tbh - until an actual exploit is out there, why should they worry?
For a typical kernel without Meltdown mitigations, the entire kernel, including that window into all of physical memory, is in the page tables of every 64 bit process at all times.
Amazon, Google, Microsoft, and basically everyone else is scrambling to fix these issues because of the potential. Basically, it’s going to take years to cover the long tail for this issue, and waiting until exploit kits are commonplace isn’t necessary to understand the potential impact.
The closest thing I can find is Intel Architecture Reference Manual Volume 3, Section 5.1.1:
> With page-level protection (as with segment-level protection) each memory reference is checked to verify that protection checks are satisfied. All checks are made before the memory cycle is started, and any violation prevents the cycle from starting and results in a page-fault exception being generated. Because checks are performed in parallel with address translation, there is no performance penalty.
If you read section 11 on caching, the terminology "memory cycle" seems to exclude cache access. Indeed, Volume 3, Section 11.7 explicitly warns that implicit caching might happen that you would not expect:
> Implicit caching occurs when a memory element is made potentially cacheable, although the element may never have been accessed in the normal von Neumann sequence. Implicit caching occurs on the P6 and more recent processor families due to aggressive prefetching, branch prediction, and TLB miss handling.
Again to use the crypto analogy: An implementation of, say, RSA that uses timing-sensitive memcmp to compare signatures would follow the RSA specification. But everyone would agree that such software has a severe bug.
The fact that speculative reads don't check the permission bits is arguably a design bug, not an implementation bug, but I'd still call it a bug.
So, the root causes of shared, on-chip resources were identified by mid-1990's, demonstrated again by later work like Percivals, being mitigated from that year onward, and ignored by CPU vendors. When I asked in the past, hardware people told me they didn't care about cache security because their sales were strictly tied to customers' benchmarks of performance per dollar and watt. Customers didn't care. Suppliers didn't care. That simple.
There were in fact (tiny) segments where customers were buying processors with more robustness or predictability. Those that come to mind were some PowerPC designs from Freescale that aerospace liked, Leon3-FT SPARC that was GPL, some smartcard components, and especially Rockwell-Collins' AAMP7G [5]. Designed with EAL7 methods from 1992, it had mathematical proof of separation at level, triplicated registers for fault-tolerance, ECC memory, and MILSPEC heat tolerance. It's used in guards to separate Top Secret/SCI info from other stuff.
So, these are old attacks with mitigations of various costs that were ignored for profit maximization by the big companies and performance maximization by most consumers/businesses who didn't buy security in general. Both old and new techniques were effective at assessing leaks and mitigations, though. They could've been used at any time, were by CompSci, were by security-critical suppliers (esp Rockwell), and even more techniques exist now for analysis [6]. Intel et al will just patch up until next attack since they and the market haven't changed. ;)
[1] https://pastebin.com/uyNfvqcp
[2] https://www.google.ch/patents/US5574912
(Note: Using the patent filing since the 1992 work is paywalled. It's the same person filing what they discovered on VAX VMM project.)
[3] https://pdfs.semanticscholar.org/2209/42809262c17b6631c0f653...
[4] https://eprint.iacr.org/2005/280.pdf
[5] http://www.ccs.neu.edu/home/pete/acl206/slides/hardin.pdf
These companies continue to employ a lot of smart people and have exciting things going on in some components, but the corporate culture makes it hard to get a consistent positive execution.
Intel has been sitting pretty for the last 20 years because chip design is not something that lends itself to modern patterns of disruption. I'm sure it will happen some day, but for now, Intel has kept its position mostly due to the difficulty of the niche than any intrinsic competitive edge.
The point being that Intel itself is not uniquely terrible; they're routinely terrible. They're just positioned in a space that's much more hostile to less-terrible entrants, so capable people do something that is less hard.
1. Process knowledge and manufacturing capacity. You can buy from others only as much as they have manufacturing capacity. Only real threat to Intel comes from combined volume of GlobalFoundries, TSMC, Samsung and UMC. Apple, NVIDIA, AMD, ARM and Qualcomm can get past Intel only trough these companies.
2. Profit margins. Intel makes 60 percent profit margins, AMD struggles from decade to decade. That's not a coincident. It's the direct result of pricing decisions by Intel. Whenever AMD gets ahead Intel in uP technology, Intel has always the option of cutting profit margins and prevent AMD from gaining more market share.
- AMD has just had a great release with Ryzen, showing they can compete on a price/performance basis.
- Apple is moving core OS functionality on its newest desktops/laptops on to Apple designed ARM chips.
- Mobile platforms are getting bigger, especially with things like ChromeOS that could be (are being?) easily run on ARM based hardware.
- Open Power has come a long way and could be poised to take some of the server market for customers who want more control than they got with Intel.
I’m excited for this. Obviously the vulns are an issue that needs to be solved, but we could get some real competition in terms of manufacturers, and even in terms of architecture. The industry will take some time to readjust to compiling for/running on multiple architectures, which I think much of the industry hasn’t needed to deal with for a while. The result though will be a market where customers can choose an architecture that makes sense for their use case, and can choose from a range of good options.
(I realise other chips are vulnerable, not just Intel, but the publicity has been Intel focused and I don’t think the technicalities of it matter too much)
Last 5 years they were slacking off, because economically there is no reason to go over the usual 10-15% yearly performance bump. But actually they were accumulating aces up their sleeves. Again, no reason to show your hand, if you don't have to.
But the time has come. Right now Intel has 3 major problems: 1) Meltdown/Spectre situation 2) AMD is awoken from sleep with surprisingly good Ryzen lineup 3) Apple craves new powerful CPUs to satisfy unhappy MacBook Pro customers
Intel can fix all of this with one sweep. Just by releasing a brand new CPU that will surprise everyone. Of course with hardware Meltdown/Spectre fix. They were holding off, but it's time to drop all these hidden aces on the table. And I believe it's gonna happen. Not right now with Cannon Lake, but with the one after - Ice Lake on 10nm transistors, by the end of 2018. It's going to be even bigger than NVIDIA's GTX 1080 success.
Intel's process advantage is shrinking. They're struggling like everybody else because the physics is getting harder and harder. Apart from the fact that it would have been nice to get easy process shrinking forever, this is good news for almost everybody: it means competition for them is getting tougher.
What you and the ggp are basically saying is that Intel slowed down the improvement in their processors on purpose over the last several years. Why on earth would they do that?
Besides, all the evidence points to the contrary, what with them being unable to compete in the mobile space.
I'm not a big hardware person, but from what I've heard the speed they released 6 core processors after Ryzen makes it likely they were capable of producing 6 core (consumer) designs earlier.
Maybe quantum computing, neuromorphic chips, GPGPU and 3D NAND are where it's at for them in the future, and traditional CPUs will be more or less commoditized.
I don't think CPU capacity failing to double every 18 months is good news for anybody. I'd rather have a monopolistic Intel churning out 2x powerful chips every 2 years than a competitive market giving 5% performance bump per year.
It's gonna take 3 years minimum until even the easiest of those things is resolved and silicon hits the street, even if Intel begrudgingly admits they need to do this.
What Intel could do in the short term is reduce their prices drastically. They have the profit margins to afford it.
...
> And I believe it's gonna happen.
I don't notice anything in your post supporting those beliefs, aside from Intel having a motivation to make them true.
Intel has a big problem. They're probably going to have to replace a lot of CPUs.
But the commodity hardware is SOOOOOO much cheaper, you know?
Everything about Intel was always about "good enough" from since probably ... 1982? And "good enough" security in hardware was always "nobody cares". IBM, certainly, was screaming about the level of insecurity in commodity hardware and software forever. DEC similarly.
But commodity hardware is SOOOO much cheaper.
No one. And I mean NO ONE was every going to give up even 10% on performance or cost in order to be even slightly more secure. At any level of the stack. Intel, Microsoft, Google ... all are guilty of this up until probably this year. Anybody who suggested that would have gotten laughed at and/or fired.
The market spoke--and security became an afterthought.
Sadly, this is STILL true. While there is much gnashing of teeth about Intel, everybody's implementation of security is far worse. The thing that is biting Intel is that the monoculture means that it is a universal and scalable tool as opposed to at the software level where each individual company has to be compromised in a slightly different way.
It's only been since everybody is putting everything in the cloud that people now care about actual absolute values of security.
The problem is that all of the hardware solutions to these bugs cost RAM somewhere. And RAM is now the gating factor of cost and performance on most chips. RAM fell off the Moore's Law curve back about 32-22nm and isn't coming back. So, the performance hit to mitigate this is real.
No one was going to be the first to fix this EVEN IF THEY KNEW AND CARED. Everybody learned from painful business experience that the first guy gets all the arrows and the second guy gets all the profits. So, everybody was going to wait until they were the second guy--which was only going to happen when something bit everybody.
In any case, it is not clear to me that you need that much SRAM to greatly minimize the various side channels that Spectre uses; it's on the order of a portion of the L1D size, which is pretty minimal compared to how much SRAM there is on-chip already for the LLC.
I lost count of how many times this article co-mingles Meltdown (an Intel flaw) with Spectre (which affects all high-performance CPUs), and uses that confusion to imply that Spectre is an Intel problem instead of an industry problem. Then it quotes Intel making that distinction and implies that's disingenuous on the part of Intel.
Either the author doesn't understand Meltdown and Spectre or they're intentionally making misleading claims.
The combination of speculative execution, virtual memory, caching and user/supervisor privilege separation isn't ten years old.
These flaws upend something like 40 years of conventional wisdom.
// it is impossible to read past the end
// of the array in this code:
if (i >= 0 && i < array.size())
x = array[i];If the hardware is broken, all bets are off. Nothing in the conventional wisdom has been challenged, except perhaps the complacent assumption that "Intel hardware is unlikely to be broken in ways that invalidate our security".
The mitigations for variant 1 I've seen are either introducing a speculation barrier on array bounds checks, or faster masking tricks which convert speculated out-of-bounds values into safe values. It won't surprise me at all if conventional wisdom changes to include speculation effects, much like it has changed to include cache effects once memory got slower than the CPU.
So, everyone is like Intel is so bad, when they just happen to be the one that people are running the most untrused code on.
Let me quote IBM "This vulnerability doesn’t allow an external unauthorized party to gain access to a machine, but it could allow a party that has access to the system to access unauthorized data."
People on this board are upset, because after spending the last decade+, making promises about how secure "cloud computing" is, once again the naysayers were proven correct. This time the flaw is so fundamental that in order to fix it you have the OS vendors making changes the destroy system performance for most applications that are I/O or just syscall intensive. This likely won't be the last time either if history is to be believed.
But javascript you cry, again I'm going to say that you shouldn't be running random code from random people on the internet.
I'm not a big fan of intel, but in a way I applaud them for pushing back against what I view as the crazy extent people go to in order to allow native code execution from untrused sources. I would much prefer this change be isolated to hypervisors, and have chrome/ff/etc detune their JIT's a bit to keep people from running cache timing attacks.
So, I likely will be turning the kpti off on most of my machines the same way I run them with the iommu's disabled because I'm not running VM's with untrused code.
(BTW, once the dust settles, i'm guessing a pretty large number of other aggressive OoO processors are vulnerable as well (old Alpha/PA-RISC/SPARC/etc)).
Certainly, if there's proof, they would, under the circumstance, provide it. Without out that proof the sale of this many shares in that time window tells us all we need to know.
Off topic: How has Intel become so dominate, and not been pursued as a monopoly? Legal issues aside, how did so many big customer allow all their collective eggs to be in a single basket? It seems to me, at some point, some of the responsibility needs to be shared by other industry titans.
I remember reading about a "cyperweapon" the US government was using to cause North Korean missile tests to fail before launch. Could attacks like this be possible through meltdown?
Intel doesn't care. They will downplay it, misinform and misdirect till the cows go home. No point beating a dead horse. The only think Intel (and most other corporations) unerstand is to stop buying their products. This is the only language they 'speak'.
This is a falsehood. They MCU updates only cover IvyBridge and newer CPU's.
Check the dates on the MCU files.
>> Starting in the mid-2000s, Intel added a layer of security within its chips and began encouraging developers to store users’ most sensitive information in the walled-off area rather than in regular software memory.
Can anyone explain what the author is referring to as "walled-off" area? L1, L2 Cache?
What "layer of security" and "area" they are talking about?
Why they're deciding to credit Intel with protected memory models is beyond me, though. Maybe they thought they needed to give some credit to Intel for something to make the article seem more balanced.
Maybe they’re referring to the fact some of these bugs are present on other chipsets but that seems weird in an Intel article. Am I missing something?
But Spectre cannot break kernel memory. Its more of a "new class" of bug, similar to how "Buffer Overflows" don't describe a particular attack, but a methodology that hackers will use to exploit new bugs.
Spectre affects virtually every high-performance computer in the world. Smartphones, SPARC, PowerPC, Intel, AMD, ARM. All of these designs use out-of-order execution, and in theory, a rogue Javascript would be able to read the rest of process memory if a programmer isn't careful about how things work.
Meltdown took it one step further: and showed that code could read Kernel memory. That was an Intel-specific mistake.
EDIT: Having done a little reading, I can't find an instance of a notable (even marginally so) smartphone using an M-series chip.
It's all in ARM Ltd. report and whitepaper.
See http://infocenter.arm.com/help/index.jsp?topic=/com.arm.doc....
>>Instruction prefetch and branch prediction >> >>The Cortex-M4 processor:
>> prefetches instructions ahead of execution
>> speculatively prefetches from branch target addresses.It's not out-of-order execution, it's speculative execution (all forms of branch prediction) plus the ability to affect the cache state during speculative execution.
It allows you to attack any security domain you can "call out" on the local host, whether it is some other process, another security domain in the same process (e.g., a JIT running JavaScript) or the kernel.
The authors used JavaScript and cross-Hyperthread/process disclosures as their primary examples probably because at leas the former is especially devastating given the ease of running JavaScript remotely when someone visits a malicious website, and because it differentiates it from Meltdown (which is mostly only about reading kernel memory) - but executing either variant is even easier against kernel syscalls (no need to attack across the JIT layer).
The huge amount of work been done on kernels to mitigate Spectre is evidence of this.
Meltdown and Spectre have opened up new hacking threats, sparked class actions, and enraged longtime partners.
At this point in time, it is known that Intel isn't the only vendor producing hardware susceptible to Meltdown and Spectre, which is another of saying AMD is in the same boat. Given this fact, I'm struggling to understand why Intel is being continuously singled out.Meltdown and Spectre aren't the first, won't be the last. I personally feel that a more interesting discussion should take place: how to prepare/plan-for/deal-with similar issues further down the road. One particular thought that comes to mind is that this industry lacks an effective recall mechanism.
Is AMD affected by Meltdown ? Do you know something that we don't ?
A75 is meltdown. Qualcomm Snapdragon 845. Go ahead and skip the next generation of Android flagships. The whole line of them will be based on 845 garbage.
I believe the number of people who would have looked at Intel's design decisions related to meltdown and thought to themselves this was a credible security holeis very close to zero. I can say this with some confidence because I used to do processor performance modeling for AMD, security verification for Intel and now do research in hardware security verification. This is not a bug until you look at the exploit.
I am certain AMD/ARM/RISC-V or whatever other Intel competitor you can think of have similar security bugs in them. It's just a question of whether people will find them. In fact, knowing what I know about AMD and Intel, I'm willing to bet these other companies have more security holes, not less than Intel
Same as Takata's airbag design which used an unstable compound. They were not unaware of the chemical propties of their inflating device. Instead, they miscalculated the impact.
Every high performance processor in the world performs speculative cache updates. That did not cause meltdown, nor did it cause spectre.
In other words, as obvious as the vulnerability is in hindsight, the general design is "textbook computer architecture" -- at least it would not raise any eyebrows. In fact the alternative, eager checking, would be the less obvious design, because almost every other exception is handled at retirement, and it would require careful thought and some additional complexity: do you cancel the instruction? (Potential problem: dependent instructions are already scheduled and will expect your result next cycle, because scheduling is deeply pipelined. Now you need some auxiliary scheduling information to let you cancel those too, which will be very complex and ripe for bugs.) Do you zero out the result? (Potential timing issue: this requires a few more gate delays between the data-cache word-select MUX and the results bus, a pipeline stage that's probably already timing-critical.) Clearly, AMD and ARM made other design decisions, but I would bet this is due to incidental aspects of their pipelines rather than an explicit prediction of a Meltdown-like vulnerability.
Overall: the issue looks very obvious in hindsight, but it arises due to an unexpected interaction of common design principles of modern cores, rather than a negligent oversight or shortcut.
(I worked on processor design for a while; throwaway for obvious reasons)
Generating the exception can still obviously happen at retire. It's not even a problem to run an extra bunch of code with constant placeholder garbage data (if that would be too costly to interrupt it sooner), given you know all is gonna be cancelled by the eventual exception anyway. But you just don't fucking load the privileged data to begin with!!! Timing side channel attacks have been known for maybe a decade for heaven's sake. Even if Meltdown would not have been as simple as it is, as soon as you speculate on privileged data you are completely dead; because there will be tons of obvious side channels to retrieve said data. (measuring occupation of execution units through HT comes to mind, actually before meltdown was disclosed I thought this was the trick to leak the data that should not have been loaded, turns out it was even simpler...)
Maybe the field needs better "textbooks", or to return to the classics about speculative execution. I've yet to see anybody blaming Intel for Spectre. Some other designers might have made the same mistake as Intel for Meltdown, but that does not make it less a mistake.
Maybe AMD only had it correctly by some kind of strange accident, and it will probably be very hard to know. That would still be a silly mistake for Intel, even then.
Yup. I think the subtle distinction I'm trying to get across is this: in the mind of a computer architect, up until Meltdown, there was a powerful and useful simplifying principle available: speculation unwinding (due to e.g. an exception) will clear away the results of any instructions after the excepting point, so it doesn't matter what we do on the "wrong path" (the instructions that will be cleared). If you set a bit in the ROB entry for an instruction that will trigger a page fault at retire, you know that it doesn't matter what data is returned, because the load will never commit to architectural state; it will be flushed. You can design the logic as "don't care" at that point.
I'm not saying that this is the correct way of thinking now, in a post-Meltdown world. I'm simply saying that the blind spot can be understood from the point of view of that principle (which seemed reasonable to many people at the time). To the layman, "allow operation on privileged data" sounds careless and negligent. To a core architect, it's a (seemingly) correct design. Speculation unwind will reset your state anyway, so one might as well omit the (non-free) eager checking logic. It's a simpler, cleaner design, easier to verify, etc.
The blind spot was that side-channels make this a leaky abstraction, and we do have to care about what happens during speculation that will be cleared. That is extremely non-intuitive to most computer architects (or at least, to me, and I did research on microarchitecture in academia then at a large chip company).
If you like, just interpret my post as a report of the widespread mentality in the industry -- I'm not saying it's right, just that this is how it likely came about.
A company which had a major flaw in its products, secretly negotiated a silent phase out of its flawed products for its bigger customers creating a private settlement, while avoiding the public courts by declaring the error trivial. This company was too entrenched to fail, as in every big actor agreed upon that seeing the company go under and half of all legacy software rewritten to accommodate new hardware, was not in the interest of a "informed" public.
Liberal fanatics of course, had no such distributed hostage effect in there market-models, and where under the illusion that the Anorexia state they created, was still too much of a influence, while in truth the upholder of citizen rights was not even present anymore at the negotiation table.
This catastrophe would later on lead to a heretic movement among the fanatics, that viewed hardware dependence in legacy not as something ugly but inevitable, but something threatening to their deity, the one free market.
Please follow me into the next exhibition hall, where we will see the reintroduction of generational slavery by debt for the underclass. Please watch your steps, some of the tiles on the floor are in slight disrepair-
The idea of what is high performance has really changed.
Moore's law is still alive, it's just dwindling now.
I thought it was transistor count in IC's?
If we assume that for state of the art chips the area is more or less fixed (by speed of light, manufacturing constraints, thermal constraints etc) then there is no difference between "transistor density" and "transistor count" - but is that the case? Is there no margin for die growth (Disregarding yield/cost - considering only physical constraints)?
I wouldn't say that the security of the "modern OS" is a solved problem. The Meltdown/Spectre issues happen to be broad-based and heavily publicized, I don't think you can draw any more general conclusions from them.
In general, as long as human beings are designing and implementing information processing systems, there will be bugs. Once AI systems are building them, there will also be bugs, but no human being will be able to understand them.
Since when do datacenter servers run the internet ? The journalist seem to not understand the role of routers and network equipment to connect those servers. You could add AS to the mix but the datacenter servers are connected to the internet not running it. The internet would still work if we removed all those servers, it would work as intended even.
Also, the initial press release was worded such that all CPUs had the same problems. This is only true for Spectre, Meltdown is specific to Intel CPUs and some other CPUs (such as recent Apple CPUs), but it is definitely not a general bug.
Then there were benchmarks one or two days ago that were seemingly intentionally misleading. E.g., they included typical number crunching benchmarks, while Meltdown mitigation only concerns system calls and Spectre mitigations are still very limited in software. Rather than communicating this clearly, it was a big cover-up screaming 'see, nothing happened'.
The problem is that after the ME and now this debacle, very few technical people still trust Intel. There may be no alternatives for e.g. HPC, but if Intel doesn't clean up this mess, people might vote with their wallets as soon as they can.
And Meltdown... Well, that's just the FIRST platform specific vulnerability found using the Spectre strategy. There will be more, hell, there are probably more already, just not published yet.
No brand is safe, this is not an Intel bug, it's a bug in Computer Science itself. If anyone wants to profit from it, they'll need more than paid Bloomberg posts, they'll need to rethink how we build processors.