Toshiba and WD NAND Production Hit by Power Outage: 6 Exabytes Lost
anandtech.com
anandtech.com
>Toshiba Memory and Western Digital on Friday disclosed that an unexpected power outage in the Yokkaichi province in Japan on June 15 affected the manufacturing facilities that are jointly operated. //
Surely that's not the reason, it would have to be "and local [backup] power failed, and the failovers for that failed too"??
Toshiba manufacture generators too, it's not like they'd need to go far to get backup power designed for them.
There must be more to this? (Which explains why people are assuming it's suspicious, I guess; and this site is making 35% of global NAND output).
FWIW, I hadn't realised that it takes ~2months to process a wafer in to a chip.
In the US in certain industries, you have to do quarterly disaster recovery testing.
I am surprised they don’t do something like that here. The losses would certainly warrant it.
One application of this kind of product is chip fabs because they are so sensitive to power disruptions.
Whether Toshiba/WD had this type of system and if so, why it didn't prevent loss of product was not mentioned in the linked article. I have heard that there is a glut in chips for SSDs so a reason to cut production can't be ruled out. However it seems like Toshiba/WD would pay the price for this outage while their competitors would reap the benefits (unless the competition agreed to somehow share the cost.)
No it's not? Isn't that basically capitalism working as intended (in this rare case)?
One would think. Though, the tech industry is much like any other industry, and I imagine a conversation like this:
Engineer: We need to install backup generators in case the grid goes down.
Middle manager: Can you do it without stopping production?
Engineer: No.
Middle manager: Screw it. It's the next guy's problem.
I can't help but feel very skeptical about the timing of this event, given the history of price-fixing in the industry.
As a result of this and SSD, Thailand's HDD industry is mostly gone now. (among other losses techies don't hear about.)
>Five fabs and an R&D center, outage was after the batteries also ran out.
For perspective, the batteries at GF's leading fab can run the 1/3 of the systems for only a few minutes. That's the scale we're dealing with.
I think before we do all sort of conspiracy theory, we need to look into reason for why was there an outage in Yokkaichi.
Judging from the scale, the "Generator" would have to be a power plant? I.e It is not feasible to have generators to operate at this Scale?
Getting an exact figure on how much utility power they use is proving difficult, but let’s shoot on the very high side and say it’s 100MW. It’s fairly easy these days to buy generators that put out 10MW of power and are either diesel or natural gas powered. Price wildly varies based on a number of factors, but even on the very high end that would cost $50M for ten such generators.
The facility itself was in the multiple billions range to build, so the added cost would be a rounding error. The environmental hazards alone due to losing containment, let alone how much the outage costs in lost business, seems pretty logical to me then that the generators existed.
So the question really is, was it incompetence (unexpected failure of backup systems) or malice (good excuse to justify constraining supply)? We will likely never know.
Outside that, have you not seen a multi-building data center complex? The power demands aren’t that different.
Lots of sites use them for backup power and simultaneously use them as regular power plants selling power back to the grid.
These values are experimentally tightened to get the highest possible accuracy to the desired effect and improve the number of working chips that leave the factory. If the power cuts out you don't know what conditions the wafer experienced while the system was winding down completely uncontrolled and your processes haven't been designed for the wafer going through the ramp up twice.
The reason why it's lost so much output is because modern semiconductor processes have hundreds of steps and (I believe) a lead time in the months, so the amount of material that's in flight at any one instant has to be huge to get any reasonable throughput.
I can imagine that if the factory is entirely automated and a full-restart has never been attempted. Every single machine will probably be in some bad state with unknown chemicals settled into unknown pipes in the machine, requiring custom flush processes to be designed, and in some cases machines might have to be replaced, which in a human-free clean room isn't easy...
Similarly, a drought that doesn't last that long (relative to the life of a big tree) can nevertheless kill that tree even though said tree has been growing for centuries.
In other words, there need be no correlation between how little time it takes to ruin something that takes a very long time to make.
When you lost power, you are not sure if the chips stayed with acids for too long or too short, or coated with unwanted amount of materials, the uncertainty kills the yield rate, which can be already low since memory chips require repeated stacking nowadays.
Similar to https://i.stack.imgur.com/yTQqw.jpg
https://www.reddit.com/r/DataHoarder/comments/c6mt9l/a_13_mi...
One analogy is to think about if you had a batch script that you were working on that touches a lot of files (1000s). Now imagine power was cut and the batch script was interupted because the computer turned off, but that computer hasnt been turned off in a long time (say it was a server).
First, you have to turn the server back on after a power outage. Was there corrupted files in it? You have to now get that server in a known working state, and if you have kept it on for years....then you may be in a world of hurt.
NOw you got your server up and running. You have the option of going through each of the 1000s of files your script was working on....but that will take time. Does it make sense to start from scratch? You will have to through our all the files you were working on, but at least you can start that script again. You could attempt to salvage every file, but that will also take time too.
These soufflés take two months to bake. The baking has to be done in such a precise fashion that even fractions of a degree in variance results in the entire batch being ruined.
Worse still, it takes a long time to bring the oven up to temperature and stabilise it on the precise temperature. You can't just scrap the batch that's ruined and start production again.
These soufflés take two months to bake but you need soufflés every day for sales. What does this mean?
You always have 2 months of soufflés in various states of production at all times.
Now you lose power and all these as very fragile soufflés in production are lost because of the power failure. Furthermore, it will take you two months to get the first soufflés off the restarted production line.
Giant UPSes are not an option in the industry because fabs eat oodles of electricity, and it is cheaper to loose a megabuck once a year than build a stabilisation/ups plant
But we all know that markets are driven by emotion: losing 3.5% of your raw materials in a market that is projected to grow 45% will cause big fluctuations. But that is just my opinion.
According to your data, the total quaterly production is 41 exabytes for SSD, which would mean losing about 37% of the total SSD production this quarter.
That being said, it is the first time I read about the scale of storage production worldwide. It makes you wonder what does the humanity store in those hundreds of exabytes per year. Probably many duplicated data or unused bytes.
That makes it a poor choice for serving anything but the rarest of YouTube videos.
Let’s as a hypothetical assume that Apple iPhones average 128GB of storage in 2019 (they go up to 512gb after all now). Let’s also assume Apple sells 50M iPhones in 2019. Doing the math, assuming my wild estimates are right, that gives us about 6.4 exabytes of storage usage in 2019 for just iPhones alone.
Android however ships something like 1.25B devices per year. The storage average is way lower I’d assume, but that’s still easily in the tens of exabytes per year most likely.
I can understand how NAND prices don't look suspect if you're not terribly familiar with the history and low-level factors of the industry, but if you really look into it, it's kind of ridiculous. The same companies price-fixed DRAM chips and got busted. Then they price-fixed LCD panels and got busted. Then they price-fixed DRAM chips again and got busted again. They were being investigated for price-fixing NAND chips, but the South Korean president shut down the investigation. Shortly before being ousted for rampant corruption (and then being bailed out of prison by Samsung). Anything that is present in such a gigantic variety of devices should cost almost nothing. That's just economics. It becomes commoditized. The materials involved and their rarity become the primary drivers of price. Comparing price per terabyte of storage between mechanical drives and NAND-based storage is the most telling to me personally. The technology that goes into modern high density mechanical hard drives is utter madness. They should be, by all accounts, astronomically expensive. They use helium, of which there is a global shortage. They coat the platters with rubidium and other rare materials. They include neodymium magnets. They include high precision mechanical motors that spin platters fast enough that the surface tension against the air becomes a significant factor (leading to the use of helium) and still maintain enough precision to be able to seek to a very precise spot in nanoseconds. Also, you've got 'hybrid' drives that include both the mechanical and NAND storage... which incurs almost no premium over the pure mechanical solution. Now they're beginning to produce drives with integrated lasers for heat-assisted magnetic recording. And these are still many times cheaper on a $/TB basis compared to.... just a dumb parallel array of NAND gates that don't require anything rare?
I was mostly thinking of serial flash chips there, which use a very different interface from plain NAND; they have their own little controller built into them. They are used in relatively big numbers for firmware etc. Even if they had the same interface, there is still a huge gap between a 128 MBit chip and the densities you find in PC storage, where we now have 512 GBit chips.
Samsung & Co. might make them too, but you mostly see other semicons badged on them.
Sure there will be some clever parties that will make some money anticipating this. But that's the same reason why the price of the gas at the pump that was already in the tank jumps up because of a shortage somewhere else. The whole stock is instantly valued at a different price.
Giga / Tera / Peta / Exa
6 Millions Terabytes of solid state memory.. quite a mass.I hope the postmortem will be public !
So do I.
It turns out that backup power fails more often than one would hope.
Generators fail to come online, batteries not performing as expected despite recent maintenance, switching gear failing, or the switching gear's safety mechanisms preventing a successful switch etc.
Source: I work for a smallish ISP, and have heard lots of stories from the ISP community, and am always eager to read about outages when there's a public postmortem.
Based on comments on the site, it appears that even a very short power disruption can mess up semiconductor manufacturing.
If that is true, then a backup power system that involved detecting an outage and starting up generators might be too slow.
If based on generators, they'd either need to have the generators always running, or have a second redundant system based on batteries that can immediately take over during the time it takes to start the first redundant system.
Or they could run their stuff off batteries all the time, with the batteries charged from the grid. They will still need something that can very quickly switch to the grid in the case of their own battery powered inverters failing.
All of these are going to add complexity and cost that may drive up the effective cost of electricity enough that if may be cheaper in the long to simply go with the grid, if they are in a place with a reliable enough grid.
Anyone know how reliable the grid is at their location?
> If that is true, then a backup power system that involved detecting an outage and starting up generators might be too slow.
I expect them to have anticipated this. I expect them to have applied a system that would have worked for them, had it worked as designed.
I would suspect the failure to have happened somewhere behind the redundant power supply (but then again why aren't the individual production steps not independently redundantly powered?).
Given the massive consequences of quite a short disruption maybe they need to figure out how to weather disruptions more robustly?
Fabs are engineered to have redundant power, but what's interesting is that the same thing happened to Samsung last year: https://www.anandtech.com/show/12535/power-outage-at-samsung...
That's my point though. If power outages hurt these fabs so severely why aren't their power supply systems more robust?
I know it's easy for me to say but I'm having a hard time wrapping my head around it.
Say in another engineering space where an hour of power outage means roughly an hour of downtime then you'd maybe not care so much.
But if, as you link here, a 30 minute power outage can "destroy 3.5% of the global NAND supply for March" wouldn't they make sure they have 0 minutes of power outage – heck, that's nearly national security levels of threat – wouldn't the South Korean government install two (or three) sets of power lines from different parts of the grid. Or a local power source (diesel generators and a small coal power plant.) Expensive? Sure. But so is 3.5% of global NAND supply for a month?
http://up2v.nl/2017/06/02/datacenter-complete-power-failures...
there are a lot of individual failure cases. DC operators learn from each failure, but there are a lot of ways things can go wrong. Fabs can be upwards of 50MW, which puts them in the range of a good-sized datacenter, so the challenges probably end up similar. (I'm saying the last part carefully - I'm much more familiar with datacenter power design than fab power design!)
And of course, it's just an estimate. Maybe the real damage will be different.
I have absolutely no idea what I'm talking about, by the way ;-) - I'm just positing a hopefully plausible explanation why the ratio of wastage might not impact the amount of "lost" (in the sense of reduced from baseline) production, even if it help reduce the amount "lost" (in the sense of unrecoverable expended resources) production.
I doubt non-experts can do much better than believe their own projections - assuming nobody with a real background here comes up with a solid reason why not.
If you mean "they should not lose so much product when equipment loses power", that's just not possible. Modern semiconductor manufacturing involves hundreds of steps where the wafers need to soak in a chemical bath for a very specific time, and missing deadlines by a few seconds causes the entire wafer to fail.
The question is very much: "why did their UPS fail?".
Either that mechanism failed somehow or there's more to this outage (maybe it damaged the production equipment as a result of spikes corresponding to the outage or repair).
I wonder if this is standat hi-tech factory process reliability.
I remember the RAM in particular - the chips had nothing etched on them, or a single 5char line. Even the firmware was unbranded, with strange timings, too (probably loosened because they would not work at standard specs).
I'm guessing it was Chinese companies buying up "bad" or excess stock and reselling it.
Haven't seen this in a while, either they tightened up regulations or the margins are too low to make a profit nowadays.
I also don’t think if I needed tons of storage like this I’d want to acquire it over a 20+ month period.
Obviously I think I’m kidding but the thought is interesting.
Wow.
Maybe they switched to using electron internally?!