Intel Licenses ARM Technology to Boost Foundry Business
bloomberg.com
bloomberg.com
That is not just a glib remark. Intel builds leading edge fabs, and runs them very well. Then it simply looks at every design it could run, and ranks them by gross margin per wafer. When the fab is full, low margin projects either run on non-Intel fabs or alternatively the teams get redeployed.
And the kind of fab capacity Intel has is a strategic weapon. If Intel decides your corner of the market is going to be paved with silicon, prepare to get paved over.
My take on Intel licensing ARM is two-fold: 1) They have decided the gross margin per wafer of running some ARMs is looking pretty good, and 2) they perceive a benefit of using Intel's fab technology as a strategic weapon to own as much of the ARM market as they want.
A high-end (higher margin) ARM SoC like the A9x is in the 150-200mm² range. I'm guessing it's the same for qcomm/samsung/lg's offering. How would it be more profitable to go from a $30-$50 SoC to an $300-$900 one ? Even when accounting for processor node differences (negligible between Samsung and Intel at 14nm or TSMC at 16nm), I can't see an ARM chip bringing high-enough margin.
Unless they envision some magical tiny IoT chip that would be expensive ? Or that smartphone SoCs price will continue to go higher? What am I missing ? Are they preparing for the end of the x86 PC/Server market (hence the need to fill the fabs) ?
We have been on a 3/5 cycle for the last 10 or so years. (3 year laptop/5year desktop)
We are likely moving to a 5/8 cycle with our current hardware.
> their dominant computer is their smart phone/phablet, which is being refreshed every two-years,
I am seeing less and less people willing to replace their phones on the 18-24 month cycle. The primary reason this occurred originally was cell phone carriers used the "upgrade" to lock people in to an additional 24mo contract. As we see less and less contracts and more people paying for their phones out of pocket or in installments, the replacement window is increasing, I expect by the end of 2017 to see the replacement be 36-48 mos for Cell Phones. This is especially true for any phone priced over ~$350.
Most of this comes down to how long the OEM will continue to issue OS Updates for the hardware, if your phone is getting updates to the OS 36mos after you bought it there is likely very little reason to spend another $400 on a phone. We will likely see some Manufacturers/Carries taking the opportunity to screw over their customers by refusing updates after 24 months in order to extort them into a replacement. Hopefully the market responds to these companies by pushing them to bankruptcy but...
I think you have missed the most important issue: phones and tables are not getting to the point where they are useful without an upgrade. I had the first android phone when it came out, and it was just barely usable: I was happy to upgrade once it was paid off. The upgrade was usable, but it needed a little more power. My next phone got upgraded only when it broke, the upgrade is better but it isn't significant for my workflow.
Personally, I consider myself showing self restraint if I don't upgrade my phone on an annual basis, and don't think I've ever gone more than 2 years - but I realize I probably fall into a niche.
We do, I have probably ~2,000 systems under management, not a single one of the Apple.
We always buy the 3 year extended warranty from the manufactures, and most of the Major Vendors will also sell a 3 year extended hardware support after the extended warrany, giving us 6 years of full hardware support on the systems, after 6 years most of them have been full deprecated by accounting anyway so for year 7 and 8 (desktops only, laptops would have been replaced at that point) we likely just keep a few spares and replace as they fail.
The nice thing about Applecare, is pretty much any major city in the world there is an Apple authorized reseller that I can bring the gear in to be replaced. Downside, is there isn't the same type of "Onsite, 4 hour repair service" that companies like Dell offer.
I've got a laptop (2013 Macbook Air) that's up for renewal, but, honestly, it's fast enough that I don't really feel the desire to get a new one until Apple comes out with their new MacBook Pros. For the first time in a long time, I don't feel any urgency to upgrade...
Yes
>>That would be very unusual for a technology company, where being able to use Apple gear to get work done is a pre-requisite for a lot of engineers considering applying for a job.
Where did I say I work for a technology company? We do have alot of Engineers however, Mechanical and Electrical. None of them use Apple.
Further I assume by "engineer" you mean software developers, I know a TON of devs that will not touch apple with a 10 foot pole. They us Linux or nothing.
Of course I also avoid silicon valley like the black death so....
>Downside, is there isn't the same type of "Onsite, 4 hour repair service" that companies like Dell offer.
This is what we have. We use Dell and Lenovo Mainly
I don't think there is anyone using Linux as a desktop, (outside of a virtual machine) though we are pretty much 100% linux on the server side (outside of IT) - where we have several hundred servers, and about 800 employees.
But now people are talking about augmented/virtual reality. lots of silicon. phones or phone like devices might take a big part in that.LG's chip will probably be for that purpose excatly . What happens if Intel loses that market ? maybe TSMC wins.
My guess is that this is all too little too late for the Mobile Market ( if that is what they intended, ) Apple A11 (10nm) and A12(7nm, since it is an evolved 10nm much like the 20nm and 16nm) all this is pretty much set with TSMC.
Also, using their advantage this way by selling at a cost might be borderline illegal (with regards to monopolistic and anti-dumping laws).
I specifically mention Apple because not only could Intel work their Modem ( Infineon tech ) within the SoC, it is also the single most important customers on the market. Arguably the only customer. There is really only three major volume Mobile SoC competitor, Apple, Qualcomm and Mediatek. The last one being on the low to medium end market which doesn't fit well with Intel. Qualcomm is an direct competitor to Intel in terms of IP and other cases. Only Apple which fits the bill, and actually will pay for using latest Fab Tech, both Qualcomm and Mediatek has many 28nm design.
With Intel's help, now ARM chips will be even more competitive against x86 chips. So far at least Intel had an edge in manufacturing, and even then it couldn't keep up. What do you think it's going to happen now? Well, ARM chips are going to eat into its x86 business even faster.
I'm not saying that this is a "bad" strategy for Intel, just that it could lead it to become just a foundry like TSMC in the end. But it remains to be seen just how well this strategy works, too, as I'm sure Intel will have a lot of "caveats" for chip makers, which will only detract them from using Intel as a supplier of manufacturing process.
I think you're missing the important point that Intel can take away market share from TSMC using this strategy, thus succeeding regardless of whether x86 or ARM wins in the end.
They've increasingly broken this cycle in the past few years, with the final demise of tick-tock and the continued number of devices which opt for ARM.
It's another in a series of developments that are reshaping the processor market and slowly moving Intel from a dominant position to a supportive one.
There's still a large number of brilliant engineers and amazingly competent people there. I wonder when they're gonna start really fighting for their future.
I don't think fighting for their future necessarily involves keeping x86 as their only chip architecture.
I think that right now their "fights" and entries into these markets have mostly involved trying to keep the world from changing, not really fighting to survive. Less new and interesting strategies, more taking out old dusty x86 cores, polishing them up a bit and shoving them into the packages and market segments where they aren't dominant.
This is their fault at this point. It used to be that without x86 compat, you had an uphill battle. They don't have that high ground anymore. The fact that companies like yours were more locked to ARM than other options should be a huge, amazingly large warning sign for them.
I think what they should have done was the work the RISC-V team already started. I'm not sure they have the right architects to do it there though, frankly. The last time they tried to start a new architecture they really messed up on keeping it simple and providing clean hardware. They didn't recover confidence after Itanium and I think there was probably a lot of "Never Again" sentiment (probably rightfully so) both inside and outside the company.
But I think at some point it is going to become clear to them that they don't have any option but to try a move that desperate and risky again if they want to be a dominant player in the processor design business in the future.
Who knows if they'll actually get it right though. I probably would bet against it and would expect them to shift more and more towards a legacy designer (see also IBM) with substantial fab tech and capabilities that remain world class even if the processor design part loses most of the device markets.
How come this wasn't put in the contract?
Object code, in some cases, can be smaller. That's why modern RISC ISAs like RISC-V include a "compressed" instruction set option, to combat the object code size disparity with CISC ISAs.
I wonder how it works. If the opcodes can all be compressed, why not just make them smaller to begin with? I guess I always thought of opcodes as being enumerations of "commands".
The 16 bit instructions map 1:1 to the 32 bit instructions, so decoding is unchanged, The difference is a pre-decoding transmogrificaiton stage where some of the resulting fields in the transmogrified 32 bit instruction are fixed (e.g. 0) and some are filled in from the original 16 bit instruction fields.
Switching between 32 bit and "thumb" is fairly painless but the instructions cannot be intermingled. When you do a call, you have the opportunity to switch 32/thumb. When you return, the processor switches back as appropriate.
The resulting thumb code is more compact than full 32 bit instructions, but it isn't 2x more compact because you still need some 32 bit instructions (subroutines) and the 16 bit instructions are not as powerful as the full 32 bit instructions.
There is very little reason to use ARM code in ARMv7 and up.
Incidentally, ARMv8's 64-bit mode (AArch64) adds a whole new instruction set, called A64. It's fixed width 32-bit per instruction, and the only option for 64-bit code.
* Excluding some really obscure, mostly long deprecated ones
Any halfword where bit[15:13]=='111' && bit[12:11]!='00' is the leading half of a 32-bit Thumb instruction
See slide 31: https://riscv.org/wp-content/uploads/2016/07/Tue1130celio-fu...
Arguably code density can be higher on a CISC device. You'll note that ARM has implemented basically code compression in their ISA with thumb.
As for microcode on x86, most of the actual instructions running on the processor are pulled from L1 already translated. This used to be the trace cache, but after P4 that was dropped and now comes back as a u-op cache, meaning mostly that the instructions are stored in decoded format and so don't get run through translation again. (In P4's trace cache, branches were smashed and whole trace sequences up to three branches long were cached and then speculatively retrieved and executed.)
Interestingly, there's research that shows you can spend plenty of time and gates to optimize the instructions as they come out of translation and go into the trace or u-op cache without negatively affecting performance, which allows some optimization basically for free. No one would do this these days because it's extra power and heat, but it was seriously considered back in the gigahertz wars.
In addition, it should be noted that the translation between x86_64 and u-op is not super larger or challenging and is much much less challenging than most things we'd think about as a compiler pass.
Decoupling internal architecture from externally exposed ISA is not an unreasonable choice, but x86 is certainly no longer a highly desirable external interface these days either.
In a world where single thread performance isn't a significant differentiator, there's fewer reasons to use x86.
English is far from the best language, but everyone on this site speaks it. Gasoline is far from the best source of fuel for a vehicle, but every corner has a gas station. x86 is far from the best architecture, but basically every OS at leasts supports it in some way.
Uh, no. CISC is far more efficient than RISC. You can do much more with fewer bytes, increasing your cache efficiency. Also, the fastest instructions are just as quickly decoded as RISC, if not faster. And any complex instructions we can leave for microcode.
The fact that early microprocessors are CISC shows you how much they could do with just a few thousand transistors. RISC architectures never took off until they could pack millions of transistors in, which should hint at how inefficient they were. This inefficiency manifests itself in modern times through power consumption. DEC Alphas & Power CPUS were already power consumption beasts, which is why Apple had to switch from Power to x86 for Macs.
Right now there is nothing intrinsically better about RISC architectures. Like VLIW, RISC is a nice theoretical computer-engineering experiment, nothing more.
This was the case briefly when RISC came out (e.g. Alpha), but x86 fought back with improved fab technology and also threw more power at the problem. Lots of power.
What blindsided RISC was that the CPU clock speed is not as much a limitation as memory access speed[1]. The voracious appetite of RISC instruction fetching[2] through the memory subsystem slowed it down more than it could speed up the CPU clock. (Yeah, yeah, broad generalizations.)
Oh man, my point in a great graphic! https://www.cs.virginia.edu/stream/
[1] If you run the numbers, todays DDRn SDRAM latency is maybe half the latency of the original IBM-AT even though the advertised when streaming bandwidth is several orders of magnitude faster. In other words, if you are not streaming the memory accesses, your memories are not much more than 2x faster than the IBM-AT. Caching enables streaming. Branches break the stream. That is why both caches and branch prediction logic is huge.
My numbers are out of date, but when I ran the numbers for the DDR2(?) memory in my hardware, the worst case latency was 56 clock cycles (e.g. if you had to close a page, open a new page, CAS latency, etc.).
[2] RISC takes 1.5-2x more instructions than CISC in my experience. YMMV.
See slide 31: https://riscv.org/wp-content/uploads/2016/07/Tue1130celio-fu...
That gives less flexibility to the processor maker for individual optimizations. A CPU can optimize its microcode generation to use its hardware as efficiently as possible. Made a hardware change that can be used for more performance?
Update the microcode generator and ship the new CPU, all updated (complex) instructions build on that are automatically faster/more efficient/whatever you optimized.
If you compile to microcode-equivalent, your compilers don't know about the new detail, so they won't try to use it. So you need to ship a new CPU, and then update all compilers to know about your new microcode, and get software recompiled, and worst case shipped in both a version for your new and all old CPUs, ...
Giving the CPU abilities to change things about how it exactly executes code means the compiler has to know less about the specific CPU running the code later.
AFAIK this is one of the things that killed the Itanium line: they tried to rely on sufficiently smart compilers instead of hardware optimizations, and had a hard time to adapt the code running to innovations on the hardware level.
If you ever have a question about computer architecture, ask two questions: 1) where does the memory bandwidth go, and 2) where does the die area go. All else follows from this.
RISC only made sense when CPU clock speeds were at rough parity with central memory speed, on-chip I-caches were limited, pipelines were short, and compilers were only good at modest optimizations.
If you look at functionality per instruction byte, RISC is fairly low. In today's world, CPU clocks are hugely faster than memory speeds, on-chip I-Caches are huge, pipelines are very very deep, branch prediction hardware is very good, and compiler back-ends are much much smarter than the peak days of RISC. CISC wins because the amount of functionality that you can move into the I-Cache per clock is higher and the amount of functionality you can keep in a given amount of I-Cache die area is higher, and compilers are good enough to target specialized instructions efficiently.
Is what Intel has crammed into X86 impressive? Yes, very. And it was damn hard work. I can tell you that walking all the way out to byte 15 of an instruction to look at the MOD R/M byte to decide if you have to raise an illegal instruction exception is a painful long path to squeeze under the clock constraint. But breaking backwards compatibility is just not something customers will put up with. So logic and circuit designers get to ply their trade in it's most convoluted form with the X86.
Perhaps these days I should amend my question list and add a third: 3) where does the power go? This is where ARM has an advantage over X86. The equations that you have to resolve to issue an X86 instruction are simply more complex than for ARM, and all that bit-flipping consumes power.
ARM conditional execution is very clever.
also mostly dead in 64bit armv8
[1] interestingly, POWER8 is capable of converting unpredictable conditional jumps over a small number of instructions to conditional execution.
There's no reason to believe this to be an immutable law going forward for much longer. With more and more going into the browser & cloud, how much is really left on the client side to be backwards compatable with these days?
The first thing I do when I reformat is install chrome, login, and sync my Google Account. I have all my important documents on Google Drive, a full office suite in Google Docs, all my bookmarks, account logins, etc.
Personally for me the web is HTML + CSS, for everything else there is a native application, even though I have done my share of web application projects.
I also don't put my private data in the hands of strangers.
Also do know lots of people that think like I do.
For desktops/laptops and mobile phones, this may be true. For embedded device (where a lot of ARM chips end up) it is not an option, usually.
Likewise, on Windows, there is a huge amount of third-party software that is, for better or worse, tied to the CPU architecture. And for a fair amount of that software, vendors are either unwilling to port it to a new CPU or have stopped supporting the software altogether (or have gone out of business). Consider the sad fate of Windows RT. Windows on low-end ARM devices might have been sweet (as far as Windows goes), but if the only piece of software it runs is IE/Edge and Office, it is not terribly useful.
You should watch this video: https://www.youtube.com/watch?v=Ii_pEXKKYUg
The most important information therein is that the difference in code density is minuscule between x86, ARM and RISC-V. The differences between GCC versions is even bigger.
> The equations that you have to resolve to issue an X86 instruction are simply more complex than for ARM, and all that bit-flipping consumes power.
It's actually just 3-10% of power consumption [1] and I even heard it claimed that the instruction decoding on ARM is larger than x86
[1] https://www.usenix.org/system/files/conference/cooldc16/cool...
ARMv7 OK, is-it still true with ARMv8? It seems to be a much more 'reduced' ISA.
Admittedly some of AArch64 is much cleaner, it involves fewer new conditional instructions than you would expect for an ARM architecture; however it sure as hell ain't RISC.
No, RISC was a response to CPUs becoming faster relative to memory access. Smaller decode logic meant more die space for registers and cache.
I wouldn't go that far. In particular, the encoding of instructions in x86-64 makes zero sense except for backwards compatibility. Since all code is littered with REX prefixes, you have the same i-cache footprint as RISC architectures do without the benefits (die area for decoding, etc.).
If you fixed the instruction encoding, extended the three-address forms of instructions introduced in AVX2 to scalar integer operations, and finally threw away real mode and port mapped I/O, etc., I agree that x86-64 would be pretty nice. I'd be happy to get even one of those improvements, honestly :)
RISC and CISC are both philophies that rose out of the contstraints of the eras they originated in and both, in their pure forms, are totally obsolete.
The effort that goes into the design of a modern CPU micro architecture is so high that the number of instructions has a pretty small impact on the overall design effort. The descendants of the original RISC ISAs have kept adding new instructions and for good reason.
On the other hand part of the increase in micro architectural complexity is that chips are designed to decode and issue multiple instructions per clock cycle. So the ease of multiple decode where every instruction is aligned to 32 bit boundaries is actually a noticable advantage.
The size of the instruction stream is still an issue but on that front x86-64 and ARM-64 are more or less tied these days in terms of bytes of instruction per task so that's a wash in practice these days.
"The thing you have to remember is that this was before the iPhone was introduced and no one knew what the iPhone would do... At the end of the day, there was a chip that they were interested in that they wanted to pay a certain price for and not a nickel more and that price was below our forecasted cost. I couldn't see it. It wasn't one of these things you can make up on volume. And in hindsight, the forecasted cost was wrong and the volume was 100x what anyone thought."
It was the only moment I heard regret slip into Otellini's voice during the several hours of conversations I had with him. "The lesson I took away from that was, while we like to speak with data around here, so many times in my career I've ended up making decisions with my gut, and I should have followed my gut," he said. "My gut told me to say yes."
Source: http://www.theatlantic.com/technology/archive/2013/05/paul-o...
I saw an interesting draft of a white paper that was pointing out the challenge of maintaining a technical advantage when you have to fab your chips with someone else. Especially if that someone else is under the nominal influence of a nation state that is hostile to your best interests. It was arguing that either Global Foundries needed to be "aligned" with US interests and oversight, or Intel needed to be drafted as "America's chip baker." The consequence of not doing so would be to create a threat to national security where the government had no way to procure the volume and complexity of chips they would need from a source they could be 100% sure was not out to get them.
And then there was this IDF announcement.
It is amazing how hard it is to imagine building a chip company "from scratch" which includes fabrication facilities. And it is hard not to see how important such chips have become in our day to day lives.
Could other countries pursue open-source chip designs?
Everything from sourcing silicon to the wafers to the packaging is such an amazingly intricate supply chain that it is a huge undertaking to try to develop it.
That said, I'm interested how you see TPP influencing the evolution (or distribution) of ARM manufacturers. Or the cost for ARM chips for that matter.
That might have to do with Global Foundries which bought IBM's fab in East Fishkill [1] which was a Trusted DoD Foundry [2].
Global Foundries is owned by the Emirate of Abu Dhabi apparently.
[1] http://www.theregister.co.uk/2015/06/30/regulators_ok_ibm_gl...
[2] http://www.aerospace.org/wp-content/uploads/conferences/MRQW...
https://www.u-cursos.cl/ingenieria/2010/2/EL653/1/material_d...
I keep it bookmarked. Anyone wanting to learn more just has to Google the phrase in question. Only thing that's missing is shuttle runs.
That said, the economics of this are pretty uncertain. Intel's business model is to spend an enormous amount of money on process R and D and cutting-edge fabs so it can produce the most advanced chips, and charge premium prices that pay off these costs. It can charge such high prices because x86 dominated computing, and it has had little real competition in the x86 space.
The ARM model, in contrast is to produce large numbers of chips at low cost for markets with many competitors and intense price competition. If Intel charges typical ARM prices, it won't be making enough money to pay for its fabs and R and D costs. If it charges premium prices, it will be more expensive than everyone else and won't sell many chips, and again won't make much money. My guess is Intel will go the latter route and make a relatively small number of premium ARM chips for the highest-end, most expensive smartphones. Better than nothing, but hardly a great success.
In terms of what in 'interesting' it depends what you are interested in. Yes all the sales growth is in low end phones, but high end phones are still selling in the hundreds of millions and still commanding decent margins, if only for Apple at least. That market isn't going away.
There are also the vr and automotive marked needing high end chips. Arm is not just phones.
Apple's history with ARM is particularly interesting and not very well known IMO. It looks like this Intel's new move will be at least as fun to watch.
The hundreds of millions they got from selling it (at a loss, IIRC) kept the lights on long enough for them get on their feet again.
On the other side ARM is beginning to deliver increasingly powerful SOCs that are at least getting close to Intel's woeful laptop cpu offerings.
It was looking like Intel could be in trouble. I mean why spend $150-300 just for CPU when you can get a powerful SOC for $10-30 with CPU, GPU and Memory all in one? There was no way Intel could compete with that.
If ARM desktops started appearing with proper Linux and Windows support, or Apple decided to put the iPad Pro SOC in a Macbook it could be trouble. Fortunately for Intel ARM chooses not to focus on this market even though the potential for low cost desktops and laptops powered by ARM is huge, with great cost savings for consumers. Driver issues hold back independent efforts. Vulcan and the new generation of mobile GPUs offer a ray of hope.
Then out of the blue Softbank buys ARM. Now Intel is licensing ARM nevermind Intel business model does not allow the kind of low cost SOCs ARM is popular for.
I hope this is not the beginning of ARM SOCs becoming pricey so Intel is less threatend in its X86 business, that in the absence of competition has frankly become extremely expensive and uncompetitive.
Something to be said for low power systems to pass all that relatively simple web traffic (JSON -> prepared statement ... row -> JSON) between the outside world and the databases, which might still well benefit from higher power / performance x86 chips. (who knows, maybe 128 low power cores serves a database better than x86, also, I can't say -- although being charged per-core by Oracle for little cores would really suck)
I'm a software guy, so feel free to expand on why this is BS, or not, hardware guys/gals.
still looks like big news
http://www.zdnet.com/article/intel-we-have-arm-license-no-pl...
> "So the short answer is, ‘No, we have no intention of using our own license to build ARM processors,'" Otellini said.
> Instead Intel is making a bet that in the long run its silicon process technology and manufacturing capabilities will give it an edge over the many fabless chip companies that design ARM SOCs and rely on foundries to manufacture them.
So, if I'm designing mobile SOC and need 2Ghz CPU in my SOC design, I check with Intel or TSMC to see if they have HardIP of that CPU that I need.
For the time being Intel has killed any remaining products that could hope to compete with ARM, that's why there is no non-pro Surface 4 because Intel killed Braxton and any successor products.
Intel had XScale and squandered that opportunity because it saw ARM as a "conflict of interest" internally. Intel has also spent over $10 billion investing in trying to enter the mobile market with its own Atom chips, and so far that has only led to a silent exit from the market.
So yeah...let's see what happens in 5 years with Intel's "ARM business".
Beyond this, even for designs that don't need the most current process, it could allow Intel to use prior plants to produce ARM chips for other clients as well. It's become clear that the demand for ARM will not shift towards lower power x86, and that x86 is unlikely to meet the power draw of arm in the near term.
Another way of reading this is that making more use of a given tech should pay the initial cost faster thus making possible more research and improved processes.
Sadly another way of reading this is that Intel doesn't have any tech that would capitalize in better chips anymore.
Oh, and all fabs require any IP as intricate as a memory or standard cell library to go through their verification process, as cell characterization is crucial to getting good yield (and not interfering with any other designs in the case of multi project wafers).
[0] https://www.arm.com/products/buying-guide/licensing/processo... [1] https://www.altera.com/products/fpga/stratix-series/stratix-...
Probably the only way they can be sure of that is to have a contractual relationship of their own with both parties.
It probably would have been better strategically to buy them and in hindsight they never should have sold XScale to Marvell but Intel isn't know to strategize. They just sort of blindly charged forward until they hit a brick wall then pivot and charge in the next direction. You saw this with Itanium, Netburst, Atom, and CMOS sensors. It's what they do.
My guess is this is an experiment or a way to burn excess capacity?
Perhaps not everyone here remembers, but StrongARM (Xscale) was a thing back in the day.