Intel Kills “Tick-Tock”
fool.com
fool.com
Secondary take away: due to the increased time to move to a better manufacturing process, Intel will likely not have as much a competitive advantage anymore. Though the number of competitors will (and has already) decrease due to the high investments.
Would be interesting to see how long intel estimates the process cycles are going to be, i.e. how Moore's law will progress.
Weird statement in the article: "it (TSCMs 7nm tech) should be very similar in terms of transistor density to Intel's 10-nanometer technology". This makes no sense as it would be comparing apples and pears. Surely they are referring to TSCMs 10nm tech?!
At this point, process nodes don't map to a particularly well-defined set of feature sizes across vendors. So what I assume they're saying is that what Intel calls 10nm is actually pretty similar in terms of transistor size/density to what TSMC is calling 7nm. (No idea if that's actually true but it's what I believe is being said.)
From a technology point of view the problem with EUV is more the line edge roughness, line width roughness, defects and the reliability... Given, more light will make the job easier
I think it's all but confirmed at this point that it's not progressing. What we're witnessing here is it's death.
If it's not dead, it's certainly dying. Perpetual exponential growth is irrational given finite resources.
I'm mostly speculating, though.
Process metrics are somewhat of a marketing nature nowadays (and not as directly related to the physical properties of the fabrication process as it may seem--at least not without adjusting for the differences between vendor-specific definitions in use).
Compare TSMC's 10nm to Intel's 14nm:
"In the case of TSMC they follow the “Foundry” node progress whereas Intel follows more of an “IDM” node transition 40nm versus 45nm, 28nm versus 32nm and 20nm versus 22nm. At the 14nm node TSMC has also chosen to call their node 16nm where everyone else is calling it 14nm."
Source: https://www.semiwiki.com/forum/content/3884-who-will-lead-10...
"Although the nominal gap in process nodes between Intel and TSMC appears to be narrowing, TSMC is not likely to catch up in terms of actual Moore’s Law scaling any time soon. TSMC’s 16FF+ process delivers only 20nm scaling, so they are still a generation behind Intel’s 14nm in terms of actual die area. TSMC said that 10nm shrinks by 0.52x from 16nm, nearly identical to the 0.53x scaling that Intel achieved from 22nm to 14nm. So if they stay on schedule, in 2017 TSMC will be in production on a 10nm process that is equivalent to the 14nm technology that Intel began producing in 2Q15. At that rate, even though Intel has slipped 10nm to 2H17, they will remain at least a year ahead of TSMC."
These numbers refer to the length of the channel in the MOSFET or FinFET between the Source and Drain dopant regions. The channel is always the smallest feature the fabrication process can create, but how it's created is actually with lots of solid-state physics tricks (such as annealing to cause the dopants to spread out from where they were originally implanted, looking a lot like a physical Gaussian blur), as the original masks used to define the transistor are nowhere near that small.
Depending on how small they can get the mask details (mostly an issue of optics and mask creation, now bounded primarily by the frequency of light they expose to the mask material, though interesting holographic tricks can help if they've left the laboratory phase) you can define the minimum size of the transistors, themselves. But there are other considerations like heat dissipation and the melting points of the various metals you've used that may force the transistors to be larger than what you can theoretically build.
The channel length is still very important because it determines how many Coulombs of charge you need to switch the transistor on, and how long it takes to propagate a signal across the transistor.
It also correlates with how much leakage current the transistor has (smaller = higher) and therefore the standby power consumption, so you do continue to make an engineering trade-off to get lower, which is also why server hardware tends to get the smallest sizes first, as they least care about idle power consumption if they're being utilized correctly.
Everything I've said is about 4 years out of date, when I switched from EE back to Software Engineering for financial reasons (and also my wife is an EE so we didn't want to put all of our financial eggs in one basket), but the semiconductor industry moves so slow, and proof-of-concept fabrication techniques take so long to productionize, that I doubt this is very far off.
The only thing marketing has done here is fight between each other on whether or not channel size is the number to focus on, as it benefits their corporation or not.
That's certainly true for Google or Amazon but not for most real corporate IT departments. It's very common for small offices to need a local server for one reason or another even though the thing will be idle most of the time.
Moving the server into a larger datacenter would require a faster network link to that office which would be both more expensive and slower for local employees than the local server. So every little office gets a server which is mostly idle.
And then you need two of them for redundancy. And then you need four of them because one app touches credit cards and needs to be in the PCI CDE and the other doesn't etc. etc.
Compute cost matters when compute is a substantial part of your budget.
I can't speak for all corporate IT departments, but I'm at a very large corporation that certainly cares about idle power consumption. Over the years our servers have gotten more powerful, and more power hungry. It's not that hard to add more power, but adding new cooling isn't so easy. The less power we use overall, the less cooling we have to add, and the happier we are.
> These numbers refer to the length of the channel in the MOSFET or FinFET between the Source and Drain dopant regions. [...] The only thing marketing has done here is fight between each other on whether or not channel size is the number to focus on, as it benefits their corporation or not.
Right -- I believe this is where the variety (or marketing) creeps in.
Here's what I mean: The various "gate length" definitions (printed, physical, effective, etc.) have a direct (and significant) impact on the ambiguity of the meaning of "process nodes" (that's in addition to the half-pitch interpretation in the DRAM industry); cf. Table 1: ITRS 2013 Data for CMOS Technology “Nodes” in http://semiengineering.com/a-node-by-any-other-name/
For reference (for anyone else reading, I know you know this), "channel length" vs. "gate length" is another difference to take into account: http://vlsi-soc.blogspot.com/2015/12/channel-length-vs-gate-... (More formally -- "Sidebar: Gate Length (Lg) versus Channel Length (L) and Experimental Data versus Equations" in http://www-inst.eecs.berkeley.edu/~ee130/sp06/chp7full.pdf)
Given that, I think it's fair to say the result is that this is more of a marketing name than a directly (in terms of physical feature sizes) interpretable number; as in:
- "At the December meeting, for example, Chenming Hu, the coinventor of the FinFET, began by mapping out the near future. Soon, he said, we’ll start to see 14-nm and 16-nm chips emerge (the first, which are expected to come from Intel, are slated to go into production early next year). Then he added a caveat whose casual tone belied its startling implications: “Nobody knows anymore what 16 nm means or what 14 nm means.”"
- "The switch to FinFETs has made the situation even more complex. Bohr points out, for example, that Intel’s 22-nm chips, the current state of the art, have FinFET transistors with gates that are 35 nm long but fins that are just 8 nm wide."
http://spectrum.ieee.org/semiconductors/devices/the-status-o...
Related: http://spectrum.ieee.org/semiconductors/design/shrinking-pos...
So the "tweak-tweak" can become the major cycle soon on more levels.
Not to judge, they make things that are mere atoms in size...but the corollary to Moore's law should always have been EY's law: The cost of each major change in semiconductor production methods doubles
It's hard to tell with Intel.
Some phones have Intel processors but they're not popular
An 2011 i7-2600K desktop CPU is still competitive with current generation i7 CPUs for general-purpose computing – most improvements were in specialized instruction sets like AVX or TSX (which infamously was broken on the first generation shipping it). Lower-end i3 and i5 are even worse off, because they don't get most acceleration instruction sets in the first place.
For consumers, it makes absolutely no sense to upgrade their desktop PCs unless they break.
The way I see it, this is our fault as software developers. A 2016 PC would be much faster than a 2011 PC if we as software developers made good use of SIMD and GPUs. But we don't.
void blur(const Image &in, Image &blurred) {
Image tmp(in.width(), in.height());
for (int y = 0; y < in.height(); y++){
for (int x = 0; x < in.width(); x++){
tmp(x, y) = (in(x-1, y) + in(x, y) + in(x+1, y))/3;
}
}
for (int y = 0; y < in.height(); y++){
for (int x = 0; x < in.width(); x++){
blurred(x, y) = (tmp(x, y-1) + tmp(x, y) + tmp(x, y+1))/3;
}
}
}
The optimised-for-speed version (order of magnitude difference): void fast_blur(const Image &in, Image &blurred) {
m128i one_third = _mm_set1_epi16(21846);
#pragma omp parallel for
for (int yTile = 0; yTile < in.height(); yTile += 32) {
m128i a, b, c, sum, avg;
m128i tmp[(256/8)*(32+2)];
for (int xTile = 0; xTile < in.width(); xTile += 256) {
m128i *tmpPtr = tmp;
for (int y = -1; y < 32+1; y++) {
const uint16_t *inPtr = &(in(xTile, yTile+y));
for (int x = 0; x < 256; x += 8) {
a = _mm_loadu_si128(( m128i*)(inPtr-1));
b = _mm_loadu_si128(( m128i*)(inPtr+1));
c = _mm_load_si128(( m128i*)(inPtr));
sum = _mm_add_epi16(_mm_add_epi16(a, b), c);
avg = _mm_mulhi_epi16(sum, one_third);
_mm_store_si128(tmpPtr++, avg);
inPtr += 8;
}
}
tmpPtr = tmp;
for (int y = 0; y < 32; y++) {
m128i *outPtr = ( m128i *)(&(blurred(xTile, yTile+y)));
for (int x = 0; x < 256; x += 8) {
a = _mm_load_si128(tmpPtr+(2*256)/8);
b = _mm_load_si128(tmpPtr+256/8);
c = _mm_load_si128(tmpPtr++);
sum = _mm_add_epi16(_mm_add_epi16(a, b), c);
avg = _mm_mulhi_epi16(sum, one_third);
_mm_store_si128(outPtr++, avg);
}
}
}
}
}
I don't know about you, but that looks like an error prone maintenance disaster waiting to happen.And, just for comparison, Halide code that produces results as fast as the second code:
Func halide_blur(Func in) {
Func tmp, blurred;
Var x, y, xi, yi;
// The algorithm
tmp(x, y) = (in(x-1, y) + in(x, y) + in(x+1, y))/3;
blurred(x, y) = (tmp(x, y-1) + tmp(x, y) + tmp(x, y+1))/3;
// The schedule
blurred.tile(x, y, xi, yi, 256, 32).vectorize(xi, 8).parallel(y);
tmp.chunk(x).vectorize(x, 8);
return blurred;
}(this is kind of a weird coincidence; last time I replied to you I mentioned Halide[1] as well)
(defun language-choice (developer)
(if (> (developer-hipness developer) (developer-experience developer)) (lang-du-jour)
(if (developer-scared-of developer 'parenthesis)
(c-family-language)
(lisp-family-language))))3D rendering, video editing, and music production all need as many cycles as you can afford, and then some.
VR and all those AI/ML technologies waiting around the corner are going to be even more greedy.
At work, our CAD people use Autodesk Inventor heavily, and that thing will happily gobble up all the CPU cycles one can throw at it. (It it the one example I have first-hand experience with.)
What I meant was that for most users of desktop PCs in an office environment, a faster CPU is not going to make much of a difference in overall system performance. (I might be a little sore because at work, users will sometimes complain there computer is too slow and then demand a new one with an i7, and then I have to explain to them why that is not going to help, while a RAM upgrade and an SSD are going to make a big difference.)
But you are right, there are plenty of examples where there is no such thing as "fast enough". ;-)
Every time you buy a cellphone with an ARM chip, Amazon buys an Intel chip to process all the data, services, apps, and websites you access from your phone and all the data those services, apps, and websites collect from you.
Heavy equipment causes vibrations in the ground, which most people probably can't notice, but when you're literally printing at the atomic scale, microscopic vibrations probably make a difference.
Thunderstorms used to just shut down litho for hours/shifts
-Dust -Chemical over spray -Cow poop
there is no such thing as a perfect filter. What you learn over time (not me, the factory as a whole) is that wind direction matters, everything matters, and you start to pick apart causes and figure out why.
I guess I kind of get why. Might be a socket compatibility and cost issue, the allocation of die space to a GPU, but it would nice to see some movement. Also probably zero demand outside of software developers, but I have to wonder if it is kind of chicken and egg problem.
It's actually kind of funny if I recall on desktops more of the die is GPU than CPU.
Server CPUs don't even use the same sockets they are fundamentally different.
36-core workstation based on Xeons: http://www.mediaworkstations.net/i-x2.html
Makes your MBP look pretty cheesy when you could have something like that under your desk. Then again, for developing rails and angular apps, you don't actually need it, and the PCI-E SSD makes a bigger difference in practice (you can RAID 0 NVME SSDs in workstations too though).
I don't know. Low-end i3s have 4 virtual cores, i5s 4 physical cores – and most consumer software still makes little use of even two. I just downgraded a bunch of users from quadcore 55W i7-HQ to dualcore 15W i5-U CPUs (laptops) – they don't notice a performance difference, because Windows and Office and all the web app bullshit is still single threaded.
Office applications may be single-threaded, but if you need to open Word, Excel, and Powerpoint simultaneously they can each use a separate core. Similarly, you may need to run a web application at the same time as one or more Office applications. This sort of multiprocess parallelism still benefits from multiple cores.
As for "Windows ... is still single threaded", I'm not sure what you are talking about. The concept of a thread is part of the OS implementation.
If you need to run multiple applications simultaneously and they do sufficient amount of computations to generate noticeable load.
> Office applications may be single-threaded, but if you need to open Word, Excel, and Powerpoint simultaneously they can each use a separate core.
And idle. Unless you're actively editing a document, the applications are doing exactly nothing with their separate cores. And so far, our attempts to create four-armed employees able to write two concepts at the same time have hit some unexpected difficulties, so generally only one document is doing any CPU intensive work (and promptly hits 100% load for the single core it can use).
Excel would be an exception, if we did any heavy processing in it – which we don't. We have real databases for that. (And Postgres does scale nicely on its Xeon servers.)
> As for "Windows ... is still single threaded", I'm not sure what you are talking about.
Windows Defender / Security Essentials is singlethreaded, throttling all I/O to 50 MB/s effectively on an i5-6200U (and inducing crippling latency). Windows Explorer is singlethreaded (or has very unfortunate locking), so a single stalling I/O request (spinning up disk/CD drive; SMB over WAN) hangs up the device's entire UI. Windows Update is singlethreaded, resulting in hour-long churn while computing applicable updates. And so on. While you get some noticeable improvements from using a dual core, and a few more from using a dual core with hyperthreading, a full quad-core (with or without HT) is just a waste of silicon for our Windows clients.
On a typical desktop workload, at least the ones users at work put on their machines (i.e. running Windows, Office, a browser, maybe a PDF viewer and sometimes something like AutoCAD for P&ID) it is rare to fully utilize more than one or two CPU cores.
In fact, for performance on a typical office PC, the CPU hardly matters any more compared to I/O and (to some degree) RAM (and insufficient RAM again becomes I/O load when the swap-fest starts). When one of our users complains about performance, unless the machine still has a Core2, rather than replacing the CPU, we tend to upgrade RAM and/or replace the hard drive with an SSD.
And multithreaded JavaScript for webapps seems like a pipedream. Virtually no one uses Web Workers because they were designed to avoid sharing data between threads, making them inordinately cumbersome to use in any routine/real-world cases.
There is a lot of room for improvement across many fronts. But I am hopeful it will occur sooner or later because I suspect adding more cores will be easier than increasing the speed or each core for the foreseeable future.
For right now any settings other than one are not tested or supported. After electrolysis ships tuning the dom.ipc.processCount value based on the number of cores and and such is planned.
Would software developers change their approach to coding if 32-core chips started to show up in consumer machines? I'd like to think so, but then again, I figured everything would be multithreaded by now.
(Which is sad because I want 16 core CPUs, but may be a long time because that isn't so useful for normal people)
You can turn hyperthreading on if possible
There's probably a narrow set of workloads that would benefit from more slower cores than 4 faster ones
The biggest bottleneck tends to be linking steps and other serialized items. (And those are getting better too; gcc's current linker supports parallelization using -j, integrated with make.)
AMD tried exactly that with the Bulldozer architecture, which was a dismal failure. Desktop workloads are frequently bottlenecked on one or two cores.
8 core 16 thread, 2.1 GHz base clock, 2.8 turbo, latest tick, latest tock, supports up to 128 gig ram so double the i7, mini itx boards @ 800 dollars including processor. No GPU though. Nice thing is, it does ECC ram. Oh and, 10Gbe. Awesome product. Not for gamers but superb for devs IMO. Throw in a gtx980 for display/compute, 1TB ssd, plus said 128 gig ecc ddr4 @ 1200 dollars, and you can do top-of-range mac pro class power for a third of the price. IE around 3.5k dollars before monitor for a very serious little rig. And if you're creative there are some stunning gamer cases for the mini itx form factor.
This is going to be "build your own" btw but anybody who's used a philips screwdriver can do this. Here in UK I have the total mentioned above at just over 2200 GBP including PSU and a cute Corsair case (http://is.gd/0pfKhg) so we're talking 3300ish dollars leaving 200 left for a pro mech KB and mouse. For the ram I had Crucial supply me for 780 british pounds (1200 USD ish) for 128 GB of 2400Mhz DDR4 ECC DIMMs (4 x 32 - there are only 4 slots so don't do 8 x 16), which works with these boards. As I said total including RAM about 2200 pounds so 3300 dollars. Or if you don't need 128GB start off with 64, say, but stick with the 32GB DIMMs so you can upgrade later.
http://www.newegg.com/Product/ProductList.aspx?Submit=ENE&DE...
Asrock Rack made some nice boards with onboard IPMI, dual Intel gigabit NICs, and 8x SATA. Kind of pricey but still under $500 for the board back in 2013. Been a great board: http://www.asrockrack.com/general/productdetail.asp?Model=C2...
Here's their Xeon D stuff: http://www.asrockrack.com/general/products.asp#Server
Denverton is the successor to Avoton; 16 cores, 16 lanes of PCI-E v3, 10GbE, DDR4: http://www.fudzilla.com/news/processors/39156-intel-s-denver...
Haven't heard anything about it for months, though.
I will have to consider this. I run a Mini-ITX board in a Micro-ATX case anyway. Perfect form factor.
I remember back in the day when SMP was the mark of a Real Programmer with a Real Computer. You kids today and your multicore processors! Fie!!
This looks like bargaining to me.
"Whenever Moore's law is no longer applicable, its definition is changed to accommodate some new version of 'computers get faster'."
Moore's Law was dead a long long time ago.
AlphaGo requires huge amount of energy and CPU power to accomplish what human brain does in 20W, just using a smallish portion of its capabilities. There must be a plenty of undiscovered architectural improvements to computers that can still kick the Moore Law for a couple years.
[As an aside, I had thought that Moore's law was originally transistor density, and the hype machine spun it into performance. But like I said, Moore's law is doomed regardless, so whatever.]
... they're switching to Tic-Tac-Toe?
(Thank you, don't forget to tip your waitress) ;-)