Ask HN: Why don't transistors in microchips fail?
Wouldn't a single transistor failing mean the whole chip stops working? Or are there protections built-in so only performance is lost over time?
Wouldn't a single transistor failing mean the whole chip stops working? Or are there protections built-in so only performance is lost over time?
The typical issue at sea level is from neutrons hitting silicon atoms. If a neutron hits the neucleus in some area of the microprocessor circuitry, it suddenly recoils, basically causing an ionizing trail of several microns in length. Given transistors are now measured in 10s of nanometers, the ionizing path can cross many nodes in the circuit and create some sort of state change. Best case it happens in a single bit of a memory that has error correction and you never notice it. Worst case it causes latchup (power to ground short) in your processor and your CPU overheats and fries. Generally you would just notice it as a sudden error that causes the system to lock up, you'd reboot and it would come back up and be fine, leaving you with a vague thought of, "That was weird".
So, if we had more hydrogen (either free or compounds) in the air this would not be the case, right?
The column of air on top of your head is equivalent (in terms of mass) to a column of water 10 meters tall, with the same base section area. But the composition is quite different, of course - the only major component they have in common is oxygen.
Hydrogen is pretty good at moderating neutrons down to thermal energies (eV range, ie room temperature) via elastic scattering, but gasses don't really have enough density to do a very good job. If you really want to protect something from neutrons you just coat it with boron. A mm coating of the stuff will keep out pretty much any common source of neutrons.
It's very surprising to see how efficient boron is. I thought neutron shields (paraffin, water) are supposed to be very thick. Maybe boron does the job via a different mechanism?
Boron's probability to capture a neutron is astronomically high, that's why you can get away with so little. Environmental sources of neutrons are actually pretty rare normally and most neutrons you do see will be pretty low energy and won't have a huge amount of penetrative power. A thin layer of boron will pretty much stop them. Pyrex (like the stuff baking dishes are made of, which is borosilicate glass - glass with boron added) is actually commonly used as control material in nuclear reactors.
http://www.google.com/patents/US7309866
Is Intel using any such things in their commodity chips?
Most notably, gamma rays killing columns was the suspected cause of dead columns while filming Superman a few years back. As a result, cameras were shipped by ship rather than air to minimize this happening ( a bit of a reactionary tale to this particular incident and not generally what happens).
Unfortunately, most other things are also of the same order of magnitude of ~20g/cm^2 - with gamma rays the single most important thing is just "how much mass is in the way". Which is exactly what you don't want.
There is a lot of disagreement on bitflips from ionizing radiation. They are unequivocally real, and unequivocally very rare. Even when they do happen, a large portion of the chip is dark a lot of the time, and a lot of the live data in the chip is simply thrown away and never used. (Think prefetching) Some bits, if flipped, will break something but will not corrupt the disk and the machine will be able to recover.
Nobody really knows for certain exactly how big of a problem they are and how often they happen- it's all statistics, and it depends on things like where on the globe your computer is, what your building is made of, and what phase of the solar cycle we are in. It even depends on workload. Anybody who claims to know for certain...
Also, FWIW that experiment will include people subject to bit errors in DRAM, not just in the CPU- and I would even guess that bit errors are more common in DRAM than SRAM given their electrical characteristics (a tiny floating capacitor vs. two inverters driving eachother)
Another slightly older paper with similar data: http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.115...
Caveats: most computers don't have ECC, and I don't remember if Blue Waters was completely installed when I visited.
Oh, and get off my lawn.
Indeed, that had to be, what, in the early to late 70s? I remember the Cray X-MP in the early 80s supported up to 16MB and frequently came with 4 to start.
The prediction was that chips below 32nm wouldn't work reliably and the only option would be to use these exotic architectures.
Well, here we are at 14nm and everything seems to be going okay.
As others mentioned, most of these problems are caught when testing the chips. Most of the transistors on a chip are actually used for caching or RAM, and in those cases the chips have built in methods for disabling the portions of memory that are non-functional. I don't recall any instances of CPUs/firmware doing this dynamically, but I wouldn't be surprised if there are. A lot of chips have some self diagnostics.
Most ASICs also have extra transistors sprinkled around so they can bypass and fix errors in the manufacturing process. Making chips is like printing money where some percentage of your money is defective. It pays to try and fix them after printing.
Also, as someone who has ordered lots of parts there are many cases where you put a part into production and then find an abnormally high failure rate. I once did a few months of high temperature and vibration testing on our boards to try and discover these sorts of issues, and then you spend a bunch of time convincing the manufacturer that their parts are not meeting spec.
Fun times... thanks for the trip down memory lane.
But even more to the point, in modern electronics the ratio of problems caused by either software or assembly level issues is very high compared to ASIC level hw problems. High volume ASICs are designed with relatively large margins to ensure good product yield. Further the cost of investigating and root causing an issue to the transistor level is very high, that unless you have measurable trend defect rate, any actual random failure not related to design would probably be attributed to something else first.
That is not to say that there are no HW design level issues with ASICs, but generally when they are discovered you would try and change the low level software to either make the problem not happen ever, or make it very very infrequent. You might also simply screen (test) the parts so you don't ship any that exhibit the undesired behavior.
So it's not that such an event that the OP mentions can never happen, it's just that it's so infrequent compared to all the other types of problems that can happen in modern electronics that unless you have a very specific reason to believe it's being caused by hardware you would never be able to differentiate it from a random software bug.
Here's a quick presentation I found on laser repairs: http://www.ee.ncu.edu.tw/~jfli/memtest/lecture/ch07.pdf
The last time I worked with some hardware folks speccing a system-on-a-chip, they were modeling device lifetime versus clock speed.
"Hey software guys, if we reduce the clock rate by ten percent we get another three years out of the chip." Or somesuch, due to electromigration and other things, largely made worse by heat.
Since it was a gaming console, we wound up at some kind of compromise that involved guessing what the Competition would also be doing with their clock rate.
Companies that make console hardware want you to be happy; they're not going to ship you a hunk of hardware that generates a Warranty Expired Interrupt at one-year-and-one-day because they want you to keep buying games and services. Having to buy a new console is a hassle. Maybe you'll buy the competitor's console instead, who knows?
On the other hand, console margins (and yes, generally they have margins these days and are not sold at a loss) are razor thin. There are knife-fights in meetings over three cent changes to components because at production scale those pennies rapidly turn into millions of dollars. You don't make a console with a reliability of 20 years because it would cost way too much and be obsolete long before the failure curve started to inflect.
So you optimize the product lifetime for user expectation of value, how long you think the technology will remain relevant, what the market will bear, and a bunch of other things (chip sourcing, cost of manufacture, architectural headroom and so on). This involves hard-won experience, spreadsheets, testing, figuring out how vendors are lying to you, plane trips to godforsaken industrial parks, and fist-fights in hallways. It's awesome :-)
Are there companies selling stuff that will break the moment the warranty is over, or earlier? Sure; they are betting that actually getting warranty service is so inconvenient that you won't bother. On the other hand, having watched (from the outside!) the execution of a product-recall-scale warranty, I was impressed with the company's professionalism, and it didn't strike me as a company that didn't care about repeat customers. It sucks on both sides for something like this to happen, but there's very little cyncism involved.
I can provide a different and less cynical interpretation for your example too. Companies that sell stuff that only lasts until the warranty is over are providing a valuable service for customers who don't want to pay as much as they would have to for a higher quality product which would outlast its warranty.
Also, with regard to the gp specific point about the discussion being in regard to a gaming console: they want the product to last as long as possible. Each additional function unit in existence counts toward their installed base and increases the attractiveness for third party developers.
The practice still makes no real long-term sense though. What do you do after `warranty_period` expires and no one wants to buy your products anymore?
Or the flip side, if you want to go that way. Higher reliability allows you to offer a longer warranty, attracting either higher prices or more customers.
This is part of how Japan made themselves mainstream car sellers in the US. They calculated that the cars had to be more reliable, because in the early days without a lot of infrastructure, recalls would be very expensive. So they made the cars reliable. Allowing them to offer longer warranties.
I just bought a two year old Toyota, after driving a Honda that I bought new for 25 years.
For something like a car or a Microwave this is not the case, but this is why the profit margin on a Microwave is closer to 50% than 5%
For example, take a look at the financials [1] for Black & Decker, a manufacturer of tools. For the quarter ending 4/4/2015, they had gross revenue of 2.6 billion. Less cost of goods sold, they have a gross profit of almost 1 billion, or about 61% gross margin. But then observe all of the numerous expenses and taxes they have, which causes their final Net Income (aka net profit) to drop to 162 million, or a net margin of about 9%.
9% can be considered an outstanding rate of return in this industry. If you could come up with a way to build a manufacturing company with such a return, you could have your own IPO.
[1] https://www.google.com/finance?q=NYSE%3ASWK&fstype=ii&ei=-Am...
OK, so what is the average workload of a laptop? What is the real-world worst case workload of a laptop? Don't forget to consider that different software operates different pieces of the chip! Ok, now that you have all that information, start modeling the odds of a weak[1] chip ending up in the hands of a worst case user...
The rabbit hole is very deep.
[1] chips have a failure distribution, like hard drives
I don't think there's any power management or throttling due to heat active at that prompt.
Luckily, the thing isn't plugged in at such times. Now, the laptop battery is dead after an hour or so.
But the failure rate after initial burn-in is phenomenally low. They're solid state devices, after all, and the only moving parts are electrons.
For example Nvidia runs it's GTX 980 and GTX 970 production at the same time. The only difference is GTX970's can have up to 2 of their compute units non-functional.
This is very commonly done in the industry. If you remember Phenom Dual, Tri, and Quadcores. Which were the same chip, just it was expected that only 5% of produced chips would be fully featured quad cores, the rest would be sold as other core counts. This was done with 27xx series i5, which were 34xx core i7's but with hyper threading disabled due to issues with yields on dye shrinks.
If a single transistor fails, normally the whole thing dies.
For large, expensive parts or parts in which a single common defect could easily blow the whole yield (for example DRAM, especially when embedded, CPUs with lots of cache or cores, and so on), regions (or rows and columns of memory) are generally fused off so that if one specific region fails qualification, it can be disabled without discarding the whole chip. This is the source of most 3-core CPUs, as well as the difference between most models in a single CPU family (they're often binned off based on how much of their L2 cache actually works).
However, once parts are manufactured and qualified, they're pretty much done. Some hardware has BISR (Built In Self Repair) but as far as I know it's not particularly common outside of DRAM.
No one has yet figured out a way to have a shorted polysilicon feature un-short itself in situ. :)
As years go by, the chip starts slowly degrading and some of the high performance chips start to get higher temperatures, worse power consumption, needs higher voltages, etc. The power management software counters this by keeping the clocks lower and the voltages higher, causing performance degradation over time to avoid catastrophic failure.
When the same chips are used in products with higher reliability requirements, they are clocked down and more conservative power management software is utilized.
disclaimer: not my area of expertise, I work on something completely different than power management.
That said, it would really interesting to see just how much the CPU actually degrades over time. I guess it's around few percent.
I'm not sure that bodes well for smart watches selling at 4+ figures.
As for the whole market segment of "this watch will pass through generations", I guess the honest thing to say is that we just don't have that kind of experience with integrated circuits yet... besides, does this type of traditional watch never need repairs? They must have failures as well.
I'm pretty sure the smart watch makers don't expect them to last for too many years, definitely not decades. After all, they want to be selling you a smart-er watch in just a few years.
This consumerism drives the whole thing, if your new watch was to last decades it would be designed in a whole different manner. And it's not only the chips, you won't be able to get a compatible display, battery, PCB or case or anything to replace a broken/worn out one in just a few years.
This sad state of consumerism is why I do woodworking to balance my mind. The pinewood dovetail box I built last week will still be there when I'm dead.
After the 2 year mark the chip became unstable and over the period of the next 6 months the clock speed it would reliably maintain was 3.2Ghz. On that progression it would be down below its default rated speed of 2.6Ghz in presumably another 6 months or perhaps outright failed.
The current estimate is that Intel targets about 15 years for a CPUs life at the clock speeds they ship, overclocking can vastly decrease that.
SSDs are solid state (duh) as well, and yet they degrade over time.
So, simplicity and hard work by fab designers is 90+% of it. There's whole fields and processes dedicated to the rest.
Past that question, Ive still plenty to learn on hardware and will take your word for it about errata sheets. Sounds right given the things described in them.
Yes but that is either a manafacturing defect (if persistant) or a transient error, or running out of speced tolerances. Or simply details of reality.
Whereas a logical design flaw, is more the actual design/implemenation, is fundementally wrong more akin to a bug.
Faults don't always manifest themselves as a binary pass/fail result; as chip temperatures increase, transistors that have faults will "misfire" more often. As long as this temperature is high enough, these lower-grade chips can be sold as lower-end processors that never in practice reach these temperatures.
Am not aware of any redundancy units in current microprocessor offerings but it would not surprise me; Intel did something of this nature with their 80386 line but it was more of a labeling thing ("16 BIT S/W ONLY").
Solid state drives, on the other hand, are built around this protection; when a block fails after so many read/write cycles, the logic "TRIM"s that portion of the virtual disk, diminishing its capacity but keeping the rest of the device going.
Sure there are. That's why Intel sells chips with a thousand different cache sizes, for example. Bad bit in the cache? Just turn that block off. Likewise for whole cores in some of the bigger chips, I believe.
http://www.tomsguide.com/forum/id-2142984/amd-phenom-710-4th...
AMD used to software-disable the 4th core on the Phenoms, but then they switched to disabling them by hardware, so that rendered any software means useless.
AMD only disabled the cores to save on manufacturing - most of the time, these cores were not working properly. But as the process got better, people got 4 or more cores for the price of 2-3 :-D
Good luck if it turns out the disabled SPU is actually bad!
As geometries fall, the effects of "wear" at the atomic level will go up.
You test by scanning in a bit pattern, issuing a single clock andscanning out the result.
Smart software generates the minimal set of test vectors that tests every wire and gate between the flops.
Chip testers are expensive (millionish) so minimising tester time minimises chip cost - we make special testign logic for things like srams
[1] http://www.caltech.edu/news/creating-indestructible-self-hea...
This seems to be a nice overview of aging effects: http://spectrum.ieee.org/semiconductors/processors/transisto....
http://www.anandtech.com/show/4142/intel-discovers-bug-in-6s...
Yes, generally speaking it would be. Depending on where it is inside the chip.
> Wouldn't a single transistor failing mean the whole chip stops working? Or are there protections built-in so only performance is lost over time?
Not necessarily. It might be somewhere that never or rarely gets used, in which case the failure won't make the chip stop working. It might mean that you start seeing wrong values on a particular cache line, or that your branch prediction gets worse (if it's in the branch predictor) or that your floating point math doesn't work quite right anymore.
But most of the failures are either manufacturing errors meaning that the chip NEVER works right, or they're "infant mortality" meaning that the chip dies very soon after it's packaged up and tested. So if you test long enough, you can prevent this kind of problem from making it to customers.
Once the chip is verified to work at all, and it makes it through the infant mortality period, the lifetime is actually quite good. There are a few reasons:
1. there are no moving parts so traditional fatigue doesn't play a role
2. all "parts" (transisotrs) are encased in multiple layers of silicon dioxide so that you can lay the metal layers down
3. the whole silicon die is encased yet again in another package which protects the die from the atmosphere
4. even if it was exposed to the atmosphere, and the raw silicon oxidized, it would make silicon dioxide, which is a protective insulator
5. there is a degradation curve for the transistors, but the manufacturers generally don't push up against the limits too hard because it's fairly easy and cheap to underclock and the customer doesn't really know what they're missing
6. since most people don't stress their computers too egregiously this merely slows down the slide down the degradation curve as it's largely governed by temperature, and temperature is generated by a) higher voltage required for higher clock speed and b) more utilization of the CPU
Once you add all these up you're left with a system that's very, very robust. The failure rates are serious but only measured over decades. If you tried to keep a thousand modern CPUs running very hot for decades you'd be sorely disappointed in the failure rate. But for the few years that people use a computer and the relative low load that they place on them (as personal computers) they never have a big enough sample space to see failures. Hard drives and RAM fail far sooner, at least until SSDs start to mature.
https://en.wikipedia.org/wiki/Electromigration
At the scale of your house wiring the effect is not so noticeable but for integrated circuits it is definitely a factor.
As for your house wiring, if it is really 70 years old you might want to worry about the insulation, not the copper.
The persons actually question is roughly why do transistors last so long compared to other types of mechanisms. No one in the comments made an attempt to answer that, at all.
That's why our boxen have power-on self tests.