Update on Samsung SSD Reliability
pugetsystems.com
pugetsystems.com
https://www.reddit.com/r/buildapc/comments/x82mwe/samsung_ss... https://www.reddit.com/r/DataHoarder/comments/x8arle/psa_sam...
How I came across this: Ran into this last week(!) on a 6-month old drive -- but I'm not in China....hmm. Not just one bad batch? Interestingly, it's non deterministic - the data is backed up but trying ddrescue, it occasionally succeeds at reading a few kilobytes from the 5 MB of several runs of 512-16384 bytes that can't be read or written. Curious to see what happens with a firmware update and secure erase.
The scamming company in question: https://zh.lobcomgroup.com/
tl;dr: All 3 of my Samsung M.2 NVMe SSDs have failed in less than 3 years. 100% failure rate.
My first SSD was a 1TB Samsung 970 EVO. It failed after 2 years and 8 months. It was replaced under warranty with a 1TB 970 EVO Plus.
That replacement has now also failed after 1 year and 9 months.
I bought a 2nd 1TB 970 EVO Plus in May 2019. It has now also failed (2 years and 7 months).
Both are expected to be replaced under warranty.
The 2 970 EVO Plus SSDs clearly had hardware errors (that were not accurately reflected in SMART data) that caused everything from system hangs, game crashes to file corruption on OTHER drives. I couldn't believe it at first but after 5 days of testing and trial and error, I had it confirmed. As soon as I removed those SSDs, my PC was completely stable again.
In the meantime, I have bought a Kingston KC3000 1TB drive as I no longer trust Samsung M.2 NVMe SSDs. On the other hand, I have a Samsung EVO 850 SATA drive which has been rock-solid.
There may be multiple, different issues with Samsung parts at play here. The 900 series issues seem to have been addressed with a f/w update; the 870 EVO issues were - allegedly - caused by bad NAND and the devices needed to be replaced.
ofc part of the problem here is the lack of public acknowledgement / information from Samsung on these issues.
Could also just be sheer chance, of course.
Anecdata is the worst, I'm sorry to hear about this happening to you. It's surely frustrating and upsetting.
Ah this is a fantastic and true hacker mindset :)
Willing to tamper with fairly expensive equipment just for the heck of it.
I'd point the finger at your PSU or motherboard. That's way too many failures for it to be the SSDs.
Samsung couldn't stay in business if that was a normal failure rate.
Interesting. Maybe my M2 (WD 570) is the cause for the hangs in my system. Thank you very much!
<https://thomas.glanzmann.de/samsung/>
The pattern is always the same: I have them configured in a raid 1. Once a month debian does a raid check. During the raid check Debian reads all data from both devices. I get uncorrectable read errors. I no longer use Samsung SSDs and replaced them them with SSDSC2KB076T8, Micron SSDs and KC3000 Kingston NVMes. No failures since then. In 2021 I told a friend of mine about the issue. He also had a 870 EVO, issued a dd if=/dev/sdX of=/dev/null bs=8M and guess what, he got uncorrectable read errors. Due to running them in RAID 1 I caught the issue early and I had no data loss or downtime because the Linux software raid compensated for the bad hardware. However I replaced them in a hurry because I no longer trust Samsung SSDs. As you can see from the smart log they're barely used. Less than 4 months in service and 10 TB written. I also got uncorrectable read errors when evacuating data from the devices.
The warranty is from the date of sale of the original unit. Replacing one doesn't reset the warranty end date.
https://github.com/torvalds/linux/blob/69f2c9346313ba3d3dfa4...
Can be a bit eye opening to look down that and see equipment you're using listed. ;)
https://github.com/torvalds/linux/blob/69f2c9346313ba3d3dfa4...
Not quite as easy to read and understand as the ata driver code though.
With the occasional further device specific workarounds in other parts of the code.
eg for specific Toshiba, LiteON, and Kioxia devices:
https://github.com/torvalds/linux/blob/69f2c9346313ba3d3dfa4...
While this seems to be special handling for Samsung X5 SSD external drives, and also Samsung 970 Evo Plus drives:
https://github.com/torvalds/linux/blob/69f2c9346313ba3d3dfa4...
Is it even possible to buy SLC drives any more? For the past 5+ years the only outlet I've been able to find that even advertise SLC is https://www.delkin.com/, and you need to speak to sales to even get a price. I just assumed they and any other similar suppliers bought giant lots of chips at the tail end of SLC production and jack up the price on every new order as their supply dwindles. Or maybe they cobble together drives from the tiny SLC chips used for cache on modern SSDs?
Yes, small ones for industrial use. They're extremely expensive, however.
Looking at the raw NAND flash prices, SLC seems to still be around $4.60USD/GB, or roughly the same as it was over a decade ago, while MLC is already <$1USD/GB despite only a doubling in capacity. TLC and QLC seems to be down in the $0.10USD/GB. You can still buy raw SLC NAND flash in the smaller capacities of few GBs; this one is only 512MB, at the price mentioned above:
https://www.newark.com/micron/mt29f4g08abaeawp-it-e/flash-me...
If the pricing was sane, SLC drives would be only 4x more expensive as QLC ones for the same capacity, but that's not what we're seeing today.
SLC needs almost no FTL. 100K endurance. Very low raw error rate that can be handled with basic ECC.
I've had an OCZ SSD in the past that also became read-only, in the sense that all changes after shutting down the computer were gone.
If you rebooted everything was fine, but as soon as the computer was shut down and the SSD also ran out of power, it converted to the state before turning on the computer. That was so bizarre. Once I had a hefty Windows upgrade installed and it was gone after the reboot - as if I had never installed it. It also took me a while to realize it because first of course you start to doubt your own memory, you don't realize that the sector- mapping on the SSD has suddenly become read-only
OCZ eventually solved this via an RMA and eventually OCZ also went upside down. The Time Warp bug it was called I think.
Real ultimate write lifespan on 3-level-cell and QLC consumer grade SSDs varies wildly for things of the same capacity and similar price.
Such as this series of tests from 7 years ago: https://techreport.com/review/27909/the-ssd-endurance-experi...
It looks like the bar charts and other data in that URL are now broken, which is sad, because I recall reading it when it was first published and it shows some amazing differences between the drives that died first, and the ones that died last.
another similar: https://www.guru3d.com/news-story/endurance-test-of-samsung-...
The official line is that endurance should not matter for most people. For example the Samsung 990 Pro 4TB is rated for 2400TB TBW - which, if the drive has a service of 5 years, is 1.3TB of data written per day. The average user will need < 1% of that.
Where that falls down though of course is cases like this. The point of a review is to show when real-world performance doesn't match the marketing. Tech reviewers seem to be blindly trusting the marketing on this one. They're really dropping the ball.
They should just become read-only, but it seems that in the vast majority, the controller just shuts off and bricks the drive.
>from 7 years ago
It's from 7 years back for good reason. They stopped doing those tests when it became impractical as endurance increased. The drives are now good enough that you can't wear them out fast enough to make sense in a review setting
...unless fundamentally broken like these
Funny how that worked out.
Even assuming a conservative 300MB per second, there's 86400 seconds in one day. That's 25920000 MB per day. Or close to 26TB per day. The samsung 960 Pro 2TB is rated by its manufacturer for a total 1200TB of write endurance lifespan.
Or at least leave it running for a couple of weeks and then see what the SMART-reported remaining write lifespan data reports it to be, versus the brand new out of box baseline.
Right. So around a month and a half. In a world where hardware news drops simultaneous by multiple outlets literally within minutes of news embargoes being lifted.
That's a lot of time investment to get result that are boring AF ("we tested them. they work"). Have very little real life consumer relevance. And the manufacturer that sent you the review sample definitely doesn't want to see (focus on edge case negative).
>Or at least leave it running for a couple of weeks and then see what the SMART-reported remaining write lifespan data reports it to be
Yeah that would make a bit more sense. Run various units down to 95%. That said the resulting story would still have watch paint dry appeal only
Any outlet interested in journalism rather purely PR can purchase a retail sample on release and publish an endurance report at a later date, however long it takes.
From my experience people buy storage whenever they have a need for something faster or larger, unlike CPUs and GPUs which have peak interest around their release dates. Storage is evergreen in that sense.
> That's a lot of time investment to get result that are boring AF ("we tested them. they work").
What's boring about that? This is extremely valuable information for any perspective buyer. Either way you gain reputation for being a trustworthy outlet that people can rely on for accurate information.
> Have very little real life consumer relevance.
I disagree and I think most consumers would be extremely interested in durability of their storage devices, especially since most of them rely on it in the absence of backups.
> And the manufacturer that sent you the review sample definitely doesn't want to see (focus on edge case negative).
That's hardly relevant. Informing potential customers about extremely serious flaws in the product is quite literally their job - at least if they wish to have any semblance of integrity, trustworthiness, and respect.
Many choose to sell out and simply echo the approved selling points they receive directly from the company, but not every outlet does this and it shouldn't be held up as something that tech journalists should aspire (or be allowed) to do.
[citation needed]
Since the firmware update last week Robocopy has not frozen the drive at all this week.
One of my earlier forays into switching to SSDs, I installed Intel... I think it was 525, 535, something like that, 2.5 inch SATA drives in several different machines. Every one has failed by now with this similar mode of (in)operation. On my desktop where I had one, it would simply bluescreen, but then come back fine until eventually reading certain parts of the disk would just always cause it to hang and it had to be replaced. Failed SSDs like this are interesting because Windows (and to a lesser extent Linux) really aren't prepared for the disk to just hang, so trying to recover anything off them can be a challenge.
Just recently I found out the last one I had around, in a little headless desktop server, was the cause of my problems with it where it would partially hang after a couple days of uptime. Having finally gotten around to having it hooked up to a display, I was treated to a sea of red dmesg errors from the disk.
I think ultimately part of the problem was new power-saving features Intel had tried to add for these disks, which would cause them to write to themselves a large amount and just eat through their useful lifetime much faster than you'd assume.
In almost every case, I replaced these with, of course... Samsungs. Though I believe I've been lucky enough not to choose any of their bad ones.
Given that it's two gen4 drives in a laptop being subjected to a moderately heavy sustained workload, I'd also suspect a thermal problem or maybe even power delivery. Those two slots are probably being fed off the same 3.3V regulator.
Since the firmware bug appears to have caused catastrophic write amplification, what may seem to the user to be only a modest and reasonable workload may be causing the drive that is the backup destination to be running at full tilt doing a ton of writes to the flash and causing the drive to hit its peak power consumption and heat output.
It is strange though that after the firmware update there have been zero freezes.
Link to specific firmware version please?
@echo off
pause
robocopy "C:\Users\o\Desktop\2023" "D:\2023" /e /mir /np /v /tee /r:0 /w:0 /log+:"C:\Users\o\Desktop\log_robocopy.txt"
pause
@echo on
I'm morbidly curious how much it reports lifespan remaining for its internal write-wear-leveling system.
SMART data is as follows (both since 2022-11-25)
Primary drive C: 5.1 TBW Model Name, Samsung SSD 980 PRO 2TB Serial Number, S***** Drive Type, NVMe Result,Byte End,Byte Start,Description,Raw Data,Status ,0,0,Critical Warning,0,OK ,2,1,Temperature (K),320,OK ,3,3,Available Spare,100,OK ,4,4,Available Spare Threshold,10,OK ,5,5,Percentage Used,0,OK ,47,32,Data Units Read,6465577,OK ,63,48,Data Units Written,10998930,OK ,79,64,Host Read Commands,150273501,OK ,95,80,Host Write Commands,157439035,OK ,111,96,Controller Busy Time,1083,OK ,127,112,Power Cycles,199,OK ,143,128,Power On Hours,571,OK ,159,144,Unsafe Shutdowns,12,OK ,175,160,Media Errors,0,OK ,191,176,Number of Error Information Log Entries,0,OK ,195,192,Warning Composite Temperature Time,0,OK ,199,196,Critical Composite Temperature Time,0,OK ,201,200,Temperature Sensor 1,320,OK ,203,202,Temperature Sensor 2,328,OK ,205,204,Temperature Sensor 3,0,OK ,207,206,Temperature Sensor 4,0,OK ,209,208,Temperature Sensor 5,0,OK ,211,210,Temperature Sensor 6,0,OK ,213,212,Temperature Sensor 7,0,OK ,215,214,Temperature Sensor 8,0,OK
Secondary drive D: 2.9 TBW Model Name, Samsung SSD 980 PRO 2TB Serial Number, S***** Drive Type, NVMe Result,Byte End,Byte Start,Description,Raw Data,Status ,0,0,Critical Warning,0,OK ,2,1,Temperature (K),320,OK ,3,3,Available Spare,100,OK ,4,4,Available Spare Threshold,10,OK ,5,5,Percentage Used,0,OK ,47,32,Data Units Read,4919136,OK ,63,48,Data Units Written,6128916,OK ,79,64,Host Read Commands,164977799,OK ,95,80,Host Write Commands,94324034,OK ,111,96,Controller Busy Time,78,OK ,127,112,Power Cycles,199,OK ,143,128,Power On Hours,538,OK ,159,144,Unsafe Shutdowns,20,OK ,175,160,Media Errors,0,OK ,191,176,Number of Error Information Log Entries,0,OK ,195,192,Warning Composite Temperature Time,56,OK ,199,196,Critical Composite Temperature Time,0,OK ,201,200,Temperature Sensor 1,320,OK ,203,202,Temperature Sensor 2,323,OK ,205,204,Temperature Sensor 3,0,OK ,207,206,Temperature Sensor 4,0,OK ,209,208,Temperature Sensor 5,0,OK ,211,210,Temperature Sensor 6,0,OK ,213,212,Temperature Sensor 7,0,OK ,215,214,Temperature Sensor 8,0,OK
But Samsung's Magician also listed my Seagate ST2000DM008-2FR102 2TB spinny disk. It found a SMART error. I ran a performance test and looked at SMART again, and the "Hardware ECC Recovered" value went from 80 to 81, with a threshold of 64. My other software labels this as "good". Nevertheless, this drive is now being replaced by a 4TB WD Blue. Thanks, article. Saved me some future troubles!
I don't trust Magician to report correctly on other vendor's storage
Aside, any idea why it thinks my drive is 208% used?
The high endurance SSDs appear to be only available in u.2\u.3\hhhl and god-help-me EDSFF form factors
Any suggestions? Micron's 7450 isn't readily available
> Life Expectancy 1.6 million hours Mean Time Between Failures (MTBF)
> Lifetime Endurance4 10 Drive Writes per Day (DWPD)
There was one update, but I don't believe it's m.2? [2]
0: https://www.newegg.com/intel-optane-ssd-905p-series-380gb/p/...
1: https://ssd.userbenchmark.com/ (sort by "Avg Bench" and you'll see these old Optanes still in the top 10)
2: https://www.intel.com/content/www/us/en/products/docs/memory...
Ha, you can't blame the consumer.
Great tech but expensive and locked to intel. Nobody was going back to the blue evil just because they had a really random reads and writes for the enterprise market.
A database server box, or even a CI build box, is a whole different business.
https://www.atpinc.com/blog/over-provisioning-ssd-benefits-e...
i would not be shocked to find tlc has a >1.5x impact on dwpd.
Been using in my threadrupper workstation with a lot of vms which are put to sleep every day with around .25tb written and read each time the vms are started. keep in mind these are 22110 form factor
I rely on the physical store i buy from where i live. that's the reason i only buy either physically or from amazon germnay (their support had been rock solid in the last 10 years i had been using them).
Any firmware issue? (https://forums.servethehome.com/index.php?threads/pm9a3-firm...)
It says to update firmware, but how can you do that from Linux? The instructions are all about some Windows program. Thanks!
EDIT: I'm quite happy with the warning from this article, fixed a potential future problem!
I'm on Arch and apparently I installed the update at some point in the past.
$ fwupdmgr get-updates
Devices with no available firmware updates:
• SSD 980 PRO 2TB
It would be good to put some pressure on Samsung to use the Linux Vendor Firmware Service. I just opened a support ticket about it.fwupd is at least manually adding a warning about the affected firmware. https://github.com/fwupd/fwupd/pull/5481
curl -O https://semiconductor.samsung.com/resources/software-resources/Samsung_SSD_980_PRO_5B2QGXA7.iso
mkdir /mnt/iso
sudo mount -v -o loop ./Samsung_SSD_980_PRO_5B2QGXA7.iso /mnt/iso/
mkdir /tmp/fwupdate
cd /tmp/fwupdate
gzip -dc /mnt/iso/initrd | cpio -idv --no-absolute-filenames
cd root/fumagician/
sudo ./fumagicianhttps://blog.quindorian.org/2021/05/firmware-update-samsung-...
And it seems to have worked. After extracting this updater tool and running it, smartctl kept showing the old firmware version (3B2QGXA7), but after reboot it now shows the new version (5B2QGXA7).
I took the risk of running this while the OS (Archlinux) was running with the disk mounted (this is the OS install disk), and at first sight this didn't cause issues. But still do it at your own risk!!
├─SSD 980 PRO 2TB:
│ Device ID: 03281da317dccd2b18de2bd1cc70a782df40ed7e
│ Summary: NVM Express solid state drive
│ Current version: 5B2QGXA7
My home is on non-redundant stripe of two 980 Pro which both had the bad firmware, so I was obviously motivated, but not panicked as it's replicated hourly to spinning rust (and I have offsite backups). I treat Flash memory as dynamic ram with only slightly better retention.After installing unzip, the firmware updated successfully.
isoinfo -R -i xxx.iso -x /initrd | gzip -dc | cpio -idv --no-absolute-filenames "root/fumagician*"
if you don't want to go through the mounting and extracting everything. bsdtar xOf xxx.iso initrd | gzip -dc | cpio -idv --no-absolute-filenames "root/fumagician*"
(And of course then running `root/fumagician/fumagician` in either case.)https://blog.quindorian.org/2021/05/firmware-update-samsung-...
I've done it some months ago, so I don't remember if it was that one exactly.
I'm now relieved to know that 5B2QGXA7 is still the current one.
nvme fw-log /dev/nvme0
nvme id-ctrl /dev/nvme0 -H | grep Firmware
nvme fw-download -f firmware.ebin /dev/nvme0
nvme fw-commit /dev/nvme0 -s 2 -a 3
nvme fw-log /dev/nvme0
In an unlikely event, may need to change the slot (-s)Extracting the actual firmware update from the files Samsung gives you might be an issue though.
and crucially, I always make sure they're 2 drives from different manufacturers, so that a bug of this nature should never be able to take down both drives in a pool simultaneously.
I think of this as the "if you're going to go to the trouble of wearing a belt and suspenders, make sure to buy them from separate brands" principle.
0: https://search.nixos.org/options?channel=22.11&show=boot.loa...
Also, their submerged in mineral oil aquarium computer was really cool, back in the day.
My 980 2TB crossed the river styx over the holiday break. Failure mode exactly as described. Nice Christmas present for me. Took 3 weeks to get the warranty replacement from Samsung.
Time to install some bloatware and see about updating their firmwares, I guess...
https://www.techpowerup.com/forums/threads/samsung-870-evo-b...
Fortunately they failed one by one, so I was barely was able to recover my RAID array by pulling out one drive at a time, powering the computer off, and waiting for the RMA replacement to arrive.
But imagine my shock to see one drive fail... only to replace it with an RMA... and then days later, seeing the next drive fail... and the next!
This issue is all over the Internet, yet Samsung would not acknowledge it was a known issue. They also refused my RMA because the corner of one of the plastic port guides was chipped - we're talking about a miniscule chip, barely visible with the human eye. So despite the fact this drive was obviously defective, Samsung won't replace it. So I'm down £350 (£225 for the original drive, and £125 for the Crucial I had to buy to replace it), through no fault of my own.
I'm not buying Samsung SSDs again - problems are one thing, but how you deal with them is paramount.
equipping a storage system with disks all of the same make, model, and vintage is invoking the statistics gods to strike failure all at once (or close enough that you won't be able to keep up with the rate of failure and time to rebuild)
personal experience: attempting to rescue a failing 192 disk system containing disks all of the same make and model. wearisome.
At least Samsung was fairly speedy with the RMAs and it was basically no-questions-asked... because I imagine they're getting tons of these mailed back to them.
> seems to primarily affect drives produced in January/February 2021
That is interesting, I wonder if it's another one of those cases where the supply chain shortages forced them into respinning the boards with some slightly out of spec parts. It's certainly been a major problem for anyone making PCBs.
The replacement drive they sent me has behaved itself so far, touch wood
What is the funny thing? The Samsung 980 died before the wear test even start.
Windows 10 writes 100KB/s constantly.
That should be illegal.
Also, if you are worried about overwriting the same files over and over, it also doesn't matter. Block device addresses are not physical addresses, controller maps them to wear the drive evenly.
It would have to be very misbehaving software or deliberate sabotage.
Normally, SSDs and operating systems both use aggressive caching to combine writes. That's the only way a drive can turn in extremely high random write benchmark numbers. Consumer SSDs do this caching even though they do not have power loss protection capacitors to ensure that data cached in volatile SRAM will be flushed to the flash in an emergency. But it wouldn't be smart for the caching to wait forever for more writes to combine with a sub-page write, which is why I'd be concerned that a slow and steady trickle of write activity may be able to cause serious real write amplification.
It's now common for consumer SSDs to have less DRAM than the normal 1GB per 1TB ratio, but they run their FTL with the same 4kB granularity and just don't have the full lookup table in RAM. There are at least a handful of special-purpose enterprise drives that use a larger sector size in their FTL, such as the 32kB used by WD's SN340: https://www.anandtech.com/show/14723
I've dug down and found random things doing dumb stuff in the past. Verbose logging turned on by default for some services, for example.
What are the benefits of it for you, compared to OS-level full disk encryption?
Such as?
I will say I love my Puget, its performance has been killer other than these lockups. And I’ve heard only good things about Puget support. I should have reported this months ago, but it’s just now that I’m doing some critical work on Windows that it’s affecting me.
One bricked itself in to read only mode after a few months.
The other has been losing 1% health each week or so. I caught it losing 2% in just two days recently.
These drives are older than the 990 model mentioned in the article but I have my suspicions anyway they’re dud drives.
Nothing lost except time - they can be swapped under warranty. But I used to buy Intel exclusively before swapping to Samsung when Intel started selling rebranded drives.
I guess the search for a reliable vendor starts again…
I’m not paying great attention to all the SMART counters day by day.
It’s for a dev workload so like … compiling code and stuff? I have the exact same workload on my desktop PC and its Samsung drive health is 99% after … years.
I think SKHynix might be the last as they supply a lot of OEM over the years. But their consumer base is small so we don't have a huge sample size like Samsungs.
However this seems like a systematic problem at Samsung.
What really stinks is that I've been recommending Samsung EVOs and PROs to relatives and friends for some years.
If I want to remain honest with myself I have to contact every one of them and have to run the SMART tests.
So now there are NO reputable SSD manufacturers left. Only reliable SSDs are MLCs from mid 201Xs especially Intel and Samsung.
The samsung-made drives lasted about 5-6 years. Everything seemed fine, and then one day you'd get a spinning pizza of death, power down, power it back on...and your SSD was...completely gone. Doesn't even enumerate on the PCIe bus. It's just gone.
Screw the SSD chipset manufacturers for not making sure that their controllers can at least a)still show up on the bus b)be read-only in some sort of recovery mode.
A recovery mode seems absolutely possible and I don’t care if it’s 100kbit over SPI. It’s better than losing everything.
And then last year with around 5500-7500 quite light hours of runtime (primarily reads, ~0.08 DWPD, well under official rating of 0.3 DWPD) drives started failing. These were definitely real failures, first indication came from regular automated ZFS scrubs and reporting increasing checksum errors and ATA errors. It was for so many drives and I'd always considered Samsung SSDs relatively reliable (even for consumer ones) that at first I thought it was a SATA controller failure, and our rep agreed and warranties back the server. They were great, gold plated support contracts pay off once in a while, and motherboard replacement and thorough testing later back in service. More drive problems. SMART short tests said everything was healthy, first longs did too. But then drives exceeded error limits and started getting faulted, and at last SMART long tests started failing. Digging in showed worrisome stats. So began swapping out and warrantying drives (cheers to the stress test to TrueNAS, in the end zero downtime or need to restore from backups). In the end, THIRTEEN (13) out of 24 failed. Brutal >50% dead drive rate. I talked to some others around and they'd seen <1 year rates also at 30-60%. Big :\. Rep also indicated they were hearing more about Samsung failures.
Anyway, gave me a talking point going forward to really, really press management on "it's worth paying for drives from 3-4x brands and maybe splurging for higher rated vs consumer too", but also does made me wonder if there is something going on, or was (pandemic related?), at Samsung's storage division. It's definitely pure anecdote but still, I spread those drive purchases out reasonably hard, and they had radically different serial numbers. Same with other folks I know at other businesses using various Samsung drives, everyone has been going to real effort following decent practices to prevent buying drives all from a single lot. Even 10% failure rate for consumer drives I could have seen, but 54%? And not a bathtub curve all frontloaded in the first month or two but after 7-11 months? That feels high? Samsung did replace them all no questions asked, they paid for shipping too. I don't have any global insight into how this all looks and it could be just plain bad luck for all of us in the region, but still.
What's going on at Samsung is the same thing that's always been going on at Samsung: they use their own flash and their own controllers and they have their own sets of problems as a result. It was a selling point in the early days of Sandforce and other turds (leading to things like the OCZ Vertex series) but now the commodity market has caught up and Samsung doing their own thing is kind of a negative. Like yeah it's fine as long as they don't screw up, but they're screwing up a lot more than they used to.
I don't see any direct correlation between the various failures that have occurred over the years. 840 Evo had something wrong with the flash NANDs that caused them to lose charge over time (leading to data loss) so they put out a new firmware that would just continuously write the flash in a tape-loop sort of deal (lol) to avoid the data ever aging out. I don't count that as a controller flaw, that's a flash flaw that was fixed with a firmware patch.
870 Evo, 970 Evo, 970 Evo Plus, and 980 have all been accused of having problems over the years, in addition to the 980 Pro and now 990 Pro, but there's actually a pretty good variety of controller models as well as different flash types (from 64-layer to 236-layer) there. It's hard to know how much firmware they all share though... or whether it was more flash problems in certain batches, or what. But overall Samsung certainly has had a lot of failures in recent years and I think it all really comes back to the fact that they're using their own controllers and their own flash, while everyone else is pretty much commodity at this point... meaning they get their own bugs too.
https://docs.google.com/spreadsheets/d/1B27_j9NDPU3cNlj2HKcr...
https://wiki.archlinux.org/title/Solid_state_drive#Update_un...
Jikes. A lot of people were hoping that the high wear is just a reporting artifact
I lost almost an entire computer lab of Dells thanks to the goddamn Sandforce firmware. One that the company acknowledged but refused to lift a finger to fix. Luckily it is possible to fix these yourself despite the vendor hostility towards the repair. Look how easy it is: https://computerlounge.it/how-to-unbrick-sandforce-ssd/
This drive costs $100, and will last 10 years or until 100TB has been written to it, as long as you keep it within the specified temperature/humidity/power conditions.
If it fails to do that, we will return $1000 to you.
Open-source SSD firmware would provide more transparency on performance and reliability.
This seems fantastic. Are you saying you could review the firmware source and know that the 980 Pro would lose ~1% of its endurance per week?
And even still, you could construct the controller so that it was burning e-fuses to indicate lifespan and the fuses could be readable through JTAG, short of complete controller death or lightning-strike level surges (which you can legitimately argue as being abuse and not warrantyable) you could make it offline-readable from an external device.
https://en.wikipedia.org/wiki/EFuse
The problems here are primarily economic/social, not technical. Companies don't want to hold warranty liability on their books for 10+ years, but they also don't really want to accept returns for defective products or other things either, and we make them do it anyway.
The EU is already pushing warranties to a minimum of two years for exactly this reason. Could it be 5 years, or 10 years? Sure, why not.
Companies will scream in the short term, of course. It's cheaper for them to push out crap that'll die and be in the trash in 3 years. Engineering products for longer lifespans would be a shift in engineering/design mindset. It probably would also push minimum device costs upwards at least a little bit, but, that's not a bad thing either - the slogan is "reduce, reuse, recycle", in that order, and "reduce" there means simply buy less or buy things that last longer. A shift away from planned obsolescence isn't the worst thing culturally, we don't want to encourage design-for-disposability.
Especially as Moore's Law slows, hardware is relevant for longer and longer periods of time. For example, a lot of people are finding that their GPUs are dying before they're actually irrelevant as hardware. It's not just NVIDIA who had bumpgate, a ton of hardware from that era failed over time due to faulty solder and probably could have been fixed with an hour of a tech's work.
Even worse, they're often dropped from support. There's really nothing wrong with a R9 290X as a GPU, but AMD won't support it with software anymore, despite the fact that it basically works anyway and it's pretty much purely a software lockout (which third parties have hacked and bypassed), because they want you to buy the new one. Wouldn't it be nice if GPUs were just expected to work for 10 years from purchase and that was covered by warranty and software support?
There are an increasing number of people who do hang onto hardware for 5-10 years because the relevant lifespan is getting longer and longer, and we should encourage that and require companies to support those consumption patterns. Just like not gluing together phones to make the battery irreplaceable, we really should be making sure electronics bumpouts don't fail in 3-5 years and that companies don't dump-and-run on the software.
Routers are another one where the software support is just egregious, too. How many rando Linksys or TP-Link or whatever actually get an update when a bunch of new vulnerabilities in WPA or whatever are discovered? Not that many, and "just install OpenWRT" is not a society-level answer especially when companies are locking down hardware.
I miss that too. Thrift stores suck now, they're pulling all the good clothes out and selling them to upcyclers and pulling all the cool electronics and cameras and other stuff and selling them on ShopGoodwill and ebay.
And ShopGoodwill is pretty absurd, almost everything is sold as-is and uninspected, and prices are just as high as ebay if not sometimes higher.
The days of wandering through a goodwill and finding some neat stuff at a bargain price are gone now, unfortunately.
I probably still have a few screenshots of these forms somewhere.
Seems clear the idea is to make sure that companies err well on the side of lifespan rather than designing something that fails a month after the warranty expires. Because if they're cutting it close, a decent number of units are going to fall under the warranty line and they'll be liable.
Even if a company is required to stand behind the product, a lot of consumers won't pursue it if it's not perceived to be worth the trouble. Do you care about the 120GB drive you bought in 2012? Not really. Do you care if you can get 10x the original ($1/gb) purchase price for it? Sure, $1200 is worth my trouble.
As they say - "A times B times C, if that's less than X, the cost of a recall, we don't do one".
I'm not OP and am not gonna die on this hill as a point of policy, but if 9/10 consumers just shrug their shoulders and accept that their 8yr old drive has failed and throw it in the garbage, that's still a bad thing at a society-wide level where you want people to be using hardware for longer and longer periods of time. Especially as moore's law tapers down even further and hardware becomes relevant for longer and longer periods of time - a R9 290X is still a pretty nice piece of hardware!
Michigan used to do something very similar with checkout price scanners - if the price coded in the system was more than advertised, you got 10 times the difference up to a limit. And the point was to get retailers to pay fucking attention because a 50 cent pricing error on a can of chili could cost them 5 bucks. Punitive damages, with citizens who spot the violations receiving the bounty.
https://www.canr.msu.edu/news/michigan_changed_item_pricing_...
Failing that, maybe a bookmaker.