How Intel Makes a Chip
bloomberg.com
bloomberg.com
This is something I've never heard of before. Anyone have some insight into this? Is it a relatively recent phenomenon?
The most common example I can think of is BCH coprocessing, which most modern application processors use for verifying data stored in embedded Flash.
https://en.wikipedia.org/wiki/Application-specific_integrate...
In this case Intel is building certain accelerators for some of its customers, and enabling it only for them before shipping it out. I don't think any of these cores contain trade secrets, rather specific functions they've requested which helps accelerate their stack.
Instead of making several versions, each with with more features, only one version is made that contains all possible features.
Whether they are turned on or off is what correlates with the "model" and the price.
Turn features off and sell as "basic" model. Turn features on and sell as "pro" model.
In this story, turning on features correlates with whether the customer has deep pockets and consistently buys in large quantities.
Say you're making an i7 and half the chips on the wafer have scratches on one or two of the cores. Instead of throwing out half the wafer (doubling the per CPU price) you just disable the damaged cores and sell them as i5s instead.
Imagine the potential conversation between Intel and the NSA:
NSA: Please backdoor your RNG for us. Intel: We're making xx billions of this chip. What happens if someone finds out and we have to recall the chip. Are you paying us the xx billions back then? NSA: ...
It is fairly unlikely the NSA could afford Intel.
Also, remember that it's easier to do if you disguise it as a coding flaw. It wasn't a backdoor: just one among many flaws accidentally hitting our systems. Wouldn't cost any market share.
NSA: Please backdoor all sorts of stuf Intel: ok
Security Researcher: these chips have backdoors
Intel: Maybe they do, what are you going to do about it?
Dan Luu wrote that Intel's cache allocation tech particularly helped Google, so they could run multiple workloads on one machine without totally trashing the larger caches at every context switch: http://danluu.com/intel-cat/
The Xeon D chips, which provide cheap low-clocked big cores and integrated NICs for small servers, were requested by Facebook: https://code.facebook.com/posts/1711485769063510/facebook-s-...
Someone who follows server chips (maybe AnandTech?) called Xeon D's one of Intel's coolest product lines to come out recently. They're also arguably a smart way for Intel to seal off the lower end of the market against approaches with tons of cheap ARM or AMD cores.
[1] zeptobars.com/en/read/open-microchip-asic-what-inside-II-msp430-pic-z80
A lot of Intel's chips are recycled binned Xeons. They've almost completely stopped making desktop processors.
The I7s you buy are usually failed Xeon chips. To my knowledge Intel currently only fabs laptop and server chips.
There were only a few xeons that were made into desktop processors and they ended up as i3's there were a few of those that supported ECC.
The 'E' edition CPU's can potentially be xeons but then they aren't failed ones but highly cherry picked ones as they boot and overclock very high.
Powering all that useless (to us) real estate must have profound consequences on our energy budget
Basically binning for different market segments.
Wait, what? Isn't that UV-free light? Ultraviolet light is used in the mask exposure step, so using normal light in the room would basically remove all of the photoresist before photolithography.
I guess they've automated the process to the point where the wafers are never exposed to light even when moving between steps.
By the way the people at Intel are working with ultra UV, and every material is opaque to them. That is the reason for the use of reflection surfaces instead of lenses for new machines.
Extreme UV is mentioned in the article, but they'll use that for the 5 nm process. I don't think ultra UV exists.
(As an aside, since the air quality of the room was mentioned...god forbid you ever broke a wafer in the foup. Entire lot in that box is ruined, and cleaning the foup itself and the equipment front ends is a giant PITA)
Source: was engineer at Varian Semiconductor / Applied Materials, who make all those giant tools that Intel and the other fabs use. Intel was definitely one of our largest customers and our most technically demanding.
By the way, how do tool manufacturers make money? Is it basically sales + maintenance + repairs? Do you have any idea what kind of profit margins they operate on?
As a whole the industry was very cyclical though. We would employ a ton of people in the factories in the years when everyone was moving to new process nodes, but in the years in between there could be lots of manufacturing layoffs / shifts to part time. On the engineering side work was more constant since we would be busy working on the next generation of tools. When I left we were working on the transition from 300mm wafers to 450mm (12 in. to 18 in.)
As far as the profit margins went, I forgot our divisions exact numbers, but AMAT as a whole had roughly ~40% gross profit margins from what I remember. Being on the supplier side of the millions/billion dollar fab builds is a good business :-)
I was just looking at the list of top semiconductor equipment manufacturers [1], and it seems that all of them are either American or Japanese companies!
Thanks for taking the time to answer man :D
[1]: https://en.wikipedia.org/wiki/Semiconductor_equipment_sales_...
To misquote George Burns, too bad that all the people who know how to build fabs are busy hanging out on HN instead...
See the problem? Someone has to build the fab in the first place so you can use it.
There's something to be said for tight integration between fab and design when you're trying to go in to uncharted territory.
If you don't need to be at the edge of process technology, you're completely right of course.
Regarding the edge of technology - Intel doesn't beat others to a new node by making both chips and fabs, it would beat others even if it was only building fabs. And it's not like TSMC is that much behind Intel. Grandparent who mentioned 28nm could mention 16nm except that there the masks will costs much more but it's still millions, not billions.
Periodically, the wafer is washed using a form of water
so pure it isn’t found in nature. It’s so pure it’s
lethal. If you drank enough of it, it would pull
essential minerals out of your cells and kill you.
I guess this bit of silliness is our Western version of Korean fan death. It's a commonly-repeated myth that humans need minerals in their water, or that distilled water has meaningful biological effects that "normal" water doesn't. But... Bohr’s solution, unveiled in 2007, was to coat parts of
the transistor with hafnium, a silvery metal not found
in nature
... somebody had to make that up completely at random.Who did that, and why did they do it? What goes through a journalist's (or an editor's) head in the process of putting a statement like that in print?
The article is actually pretty interesting... but how much of it am I supposed to believe?
http://chemistry.about.com/od/waterchemistry/ss/Distilled-Ve...
I've had this argument on here before and am not interested in revisiting it. Suffice it to say that you will be able to find plenty of sources for various old wives' tales along the lines of "Don't drink distilled/DI water, it'll kill you/rot your teeth/give you an itchy rash." None of them will include respected medical texts, legitimate peer-reviewed journals, or even blog posts written by people who remember what their fifth grade Health textbook had to say about the operation of the human kidney.
It's worth objecting to this kind of mythology because some people may assume that the converse implication in the Bloomberg article is also true, and that drinking arbitrarily-large amounts of "normal" water is harmless. The truth is that you will die if you force yourself to drink too much DI or distilled water, and you will die just as quickly if you drink the same amount of tap water.
I'm not arguing it will kill you, even in more than a sip amounts, and my "don't make a habit of it", was more on the lines of "it's likely not great for your teeth and mouth", since its a bit more reactive, not "OMG THE SPECIAL WATER WILL KILL YOU. ALSO WHAT DID MY 5TH GRADE TEXTBOOK SAY?" Again, sorry if I wasn't clear in my first post!
Quick edit: I don't know if you edited your response or I just missed it the last paragraph (probably the latter) but I whole-heartedly agree on the point of addressing this kind of thing! Totally wasn't trying to argue for accuracy of that statement, just being very nit-picky about it being DI vs distilled...for really no reason what-so-ever, haha.
Just my best guess from a single processor architecture class so definitely not positive that's the answer.
Correct. The chip has to be small enough that the clock can propagate everywhere within the chip within a single cycle, or problems will occur.
> If a chip is physically bigger, it takes longer to move bits inside of it.
Yes, so either you would need to delay for some cycles to ensure that the information has propagated (which will basically nullify the performance gains from cranking up your clock) or clock parts of the chip differently, but you're pretty much always limited by the slowest part of the chip (which is why every modern chip has a cache, because otherwise it would stall waiting for data).
Problems like what? These chips are already chopped up into different clock domains, and it's easy to install some PLLs so that perfectly-synchronized clock signals can blanket a chip even if it's inches across.
Moving data around is also not a big deal. The Xeons in the article already have multi-nanosecond ring busses running around between cores[1]. They don't slow the chip down because the design simply lets long-distance data transfers take multiple cycles. L3 and I/O don't have to be blazingly fast in terms of latency.
[1] http://images.anandtech.com/doci/8423/HaswellEP_DieConfig.pn...
The chip won't work.
> These chips are already chopped up into different clock domains, and it's easy to install some PLLs so that perfectly-synchronized clock signals can blanket a chip even if it's inches across.
Sorry, I didn't explain myself well enough. Of course chips have different clock domains, but these also come at a cost. The more synchronization you need to do between domains, the less die space you have for computationally useful stuff.
> L3 and I/O don't have to be blazingly fast in terms of latency.
I would argue differently, the impact of latency is highly dependent on the type of computation you're doing. If you're doing something with a lot of data (say, encoding video) then you need to be moving data as quickly as possible between the processor and memory. Any additional latency in cache or I/O will cause the performance to suffer.
Ideally, you want the latency of L3 and I/O to be as low as possible.
And encoding video is nearly the platonic ideal of not caring about memory latency. You could easily make memory requests ten thousand cycles before you need the results. You just need throughput.
That calculus (so-called Dennard scaling) has now broken down, which is why Intel has abandoned their tick-tock development model in favor of a model that doesn't solely rely on process shrinks in order to achieve better performance.
Because the amount of power leakage (heat) is proportional to the size of the transistors. So you cannot improve the efficiency of the chip without making the transistors smaller or lowering the frequency.
> Obviously there would be downsides in cost of materials and power consumption,
Huge costs. Data centers aren't only worried about the power consumption of chips, because for every watt that's generated as heat, they have to use >1W to cool that (because cooling systems aren't 100% efficient themselves).
As you may know from other branches of science, the resistance of a conductor rises with heat, meaning that electrons running through the conductor are more likely to hit a vibrating atom and dissipate as heat. This is why so-called "super conductors" are usually super cooled. Very little atom movement = very low chance of an electron hitting a moving atom. Silicon is a semi-conductor, but the principal is the same. If the temperature of the chip rises, so does the heat generated. This is why world-record setting overclocks are done using liquid nitrogen to cool the chip.
To prevent things from getting too toasty, data centers would have to reduce the density of the servers, which means they would need larger buildings and more land to house the same number of servers.
> but wouldn't that be offset by avoiding the need for nano-manufacturing advances in every single round?
In short, no.
Small correction: leakage actually increases as transistors shrink. This is why high-k dialectric and fin-fets were such important developments. They pushed back the point at which leakage power overtakes switching power as the dominant source of waste. Even with these technologies, we have to do a lot of design work to reduce leakage. I can't even guess how many power domains are on modern cpus -- certainly dozens if not hundreds. Most of those domains can be switched off to eliminate leakage in those domains altogether.
I'm nearly certain you meant switching power reduces as transistors shrink, which has been true so far. Things are getting weird with these new processes, and a lot of things that we've held as fact are looking less and less reliable.
Yup. Thanks for the correction.
http://synisma.neocities.org/perf_scale_cheatsheet.pdfAfaik temperature dependence of materials is a more complicated relation than this. The graph of their coefficient does not even need to be monotone and can change depending on what properties dominate.
E.g. a semiconductor can have NTC because charge concentration increases with the temperature, however when it reaches saturation this effect diminishes and it will behave similar to a regular PTC conductor.
https://en.wikipedia.org/wiki/Temperature_coefficient#Negati...
Not at all. A typical cooling system might use 1 watt to move 4 watts of heat outside.
And that's not really how superconductors work. That's how normal conductors work.
Can you name any super conductors which function at room temperature? I'm not aware of any. [1]
[1] https://en.wikipedia.org/wiki/High-temperature_superconducti...
All this feeds into a competitive market between chip manufacturers. If you are a memory manufacturer and can find the right balance that makes the product 5% cheaper to manufacture, they can make a ton of money and gain marketshare. There are a few companies like Apple, Intel that can differentiate on brand name but the rest are commodity products that compete on cost and features.
Five wafers stacked on top of each other would take 5 times as much capacity and materials to produce and cost about 5 times as much. The interconnect would be very difficult. Lastly, while you could cool the top wafer, the bottom layers would have to go through a lot of silicon to remove heat. Five times the number of layers on each wafer would also be prohibitively expensive, you don't just cut it thicker, you vapor deposit each layer under a mask and often also need to etch off parts of previous layers or ensure that higher layers are still planar. That gets difficult when you have a few steps of logic layering and some metal layers for interconnect, and would be much more expensive (and hot) with many layers.
This isn't so much of a problem with eg. NAND flash chips which are low power and the address and data pins can just be shared with a couple wafer select wires to separate them, but processors are neither low power nor trivially paralleled.
But true, heat is a huge problem.
Imagine you have a 4in by 4in square. If your chips are 1in^2 you can fit sixteen chips but if your chip is 4in^2, you can only fit four. Now imagine you have a thick scratch going diagonally right down the middle. With the bigger chips, you might lose all four to that single error, wiping out your yield. With the smaller chips however, you'll only lose part of your wafer.
Specialized processors like those made for mainframes or RAD hardened ones can be much bigger since the set up costs will vastly outnumber the fabrication cost anyway. Companies like IBM that aren't in as cost competitive a market as Intel make big chips all of the time.
Intel has an advantage here: they use 12 inch wafers while other fabs use 10 inch wafers. This improves their relative yield.
Overall these are very very strong factors which push towards shrinking chip size at every opportunity.
Considering Intel's recent exit from the smartphone SoC business and concentration on the data center, I have a suspicion why Bloomberg was "given the most extensive tour of the factory since President Obama visited in 2011".
Anyway, HN's dupe detector isn't all that performent either.
Well, looks like they have really lowered the bar, I seem to remember it being twice as much power, for half the price, every 18 months.