SSD drives’ 40k-hour “death bug” continues to catch enterprises unawares
thestack.technology
thestack.technology
The vast majority of the people responsible for the servers on the billion-dollar company I work for have never physically touched a server, a rack, or an Ethernet cable.
There's a whole generation of "point and click" admins quickly gaining prominence. I suspect cloud services have contributed to this lack of knowledge, because if all you know is cloud, you don't need to know things like this.
All of these little companies starting with brand new hardware as well. Far from all but a few of those 'garage startups' I knew of did limp along on used hardware for some time. It happened again a bit after the dotcom bubble burst (especially for office equipment) and again around 2008.
But for drives in particular, people are pretty superstitious about giving them up.
This feels honestly like my whole career. Just get to a point of competence and then get an offer you can't refuse.
It can help seeing people struggle with tech debt or bad decisions you had nothing to do with, if you can put yourself into the shoes of the person who did it. The wise man learns from the mistakes of others. But it does mean you have to do a little more sleuthing than the typical postmortem that just blames everything on an engineer who quit six months ago. That's facile, and neglects the contributions of people who still work here and/or think the same way.
> if all you know is cloud, you don't need to know things like this.
That's part of the cloud value proposition, isn't it? Leave the physical (and other low-level) stuff to people who do know that stuff, and can also reap economies of scale, while you focus on your own parts of the problem. Do you not approve of such specialization?
When AWS fails, it fails spectacularly. It also experiences outages fairly often, and you as the developer often have nothing you can do about it, but twiddle your thumbs and hope that your application was amenable to "high-availability architecture" (which is more expensive...). Finally, this particular failure mode brought down HN's servers, rented from a reputable hosting provider, and arguably this is something that a hosting provider should know.
It's also really hard to fix things when they're surrounded by a black box operated by someone who doesn't care. I kind of enjoy doing it, but it would be a lot more productive if I was able to diagnose network link problems by checking the error counters on all the routers/switches rather than by sending pings that hash over the numerous lacp alternate routes between two servers to figure out how to get a traceroute that shows the issue so the network owner can look at counters and fix the underlying issue.
This is a specific bug, like an overflow that is consistently triggered after about 4 years of use.
https://www.dell.com/support/home/en-us/drivers/driversdetai...
https://support.hpe.com/hpesc/public/docDisplay?docLocale=en...
No one registers their hardware at the vendor because all they do is spam you with offers then. Or they change their system and despite you registering you'll still get no notice.
They weren't end-of-life when they were sold this way. That's just the PR department trying to make their failure seem less severe than it was.
Response was epic.
Everything you say is true.
But we don’t care, we rent the servers.
As in, no one believed him at the time?
And what does Netflix have a need for SSDs for? Their load seems entirely predictable / serveable / stable with spinning disks and reasonably (but not lightning) fast access patterns? I would have guessed the network latency is always dominant.
I dunno what talk the above comment is referring to but most likely the young hero was saying something somewhat clueless, not being kept down by the tyranny of older idiots or whatever is being insinuated. Netflix's content is pre baked. They have to do like a dozen versions of every episode to get the best experience across the variety of hardware, but it's not livestreaming or such. Where they use SSDs it's not an erase/update heavy workload, at least for the CDN part (Netflix does a ton of other stuff including basically running entire movie studios out of the cloud, not referring to that). Density and power also matter a ton as they want their POPs to be as easy as possible to convince people to run at peering points.
I have no idea what that guy is talking about or why Netflix would care about SSD failures more than a normal company that just uses mirrors/raidz for everything.
I haven't seen any drives that really stood out as terrible in their reports in the last couple of years for 12TB+ spinning disks, although the HGST product do seem to be the most reliable.
That said, there's no real standard blanket vendor recommendation so long as the "Enterprise" criteria is met.
Well yeah that’s probably why you see “SanDisk” and think of consumer SD cards.
So the code was never tested even once?
Engineering processes are systems with their own bugs, and the people/systems that implement them can’t deliver perfectly either.
There are always going to be things that slip through, and some of them are bound to look embarrassingly sloppy.
I'm not talking automated testes or anything. Just once, on your own machine.
I agree that it is a bit unfair to pick an example after the fact. But it does not inspire confidence.
I'd love to see an entire thread on this topic alone.