How to estimate an SSD’s working life
eclecticlight.co
eclecticlight.co
"sudo smartctl -l ssd /dev/sdaWhatever", for anyone interested.
apt install smartmontoolsSSDs often have a data retention spec that basically defines how long you have until your bits start flipping, and it usually falls off a cliff w.r.t. temperature, which can make SSDs non-ideal for offline backups.
However, I've read that reading from the SSD periodically allows it to detect these errors. Some say that even keeping it just powered on is enough.
My question is, do SSDs run some sort of internal scrub while they're powered on? I don't think so, based off of some power consumption tests I've done.
Also, if they do detect an ECC error, will they actually re-write the block in question, or just correct it and return a successful IO, while leaving the compromised data still on the media?
I don’t mind having to have the SSD replaced by Apple if the cost is reasonable, just as I do with batteries, but would be good to know what to expect.
While I don't like it being that way, I can definitely see the benefits of non-upgradeable RAM. The performance of the SOC on my "lowly", entry-level M1 Air is out of this world.
But an SSD that's glued on to the motherboard has 0 benefits that I can think of and basically only serves to give any computer a hard coded expiration date. And thinness is not an excuse. There are computers as thin as the Air that have removable storage drives.
That's really hyperbolic.
The 13" M1 MBP is 0.7mm thicker and 30g heavier than the equivalent Intel of not so old, 2015-2019 when the dimensions dropped to the lower bound.
I was surprised to realize they'd made the newer machines significantly thicker and heavier. Both machines have their pros and cons, neither is perfect or terrible.
Thank you, buran777!
Can't really think of many (any?) laptops that have this mix of thin, light, high performance, and long battery life.
Sadly, I've tried PC laptops (Surface tablet, Surface laptops and Asus laptops) and the performance (speed, thermals and battery life) are still nowhere near the Apple M1 hardware. I wish there was better price/performance from apple, but until the PC world has a strong contender, I don't see that happening.
At least Intel has strong competition from AMD now :)
Some things are worth more than saving 10 grams.
And the only viable mitigation strategies involve giving Apple more money.
Spec'ing a larger onboard SSD to spread out the writes for hopefully longer endurance should be effective.
And, perhaps, opting for more RAM to reduce VM writes to disk? I'm not sure if that's effective. Perhaps the resulting sleep file will be larger as well, resulting on more writes overall. I'm sure somebody can chime in on that?
At least the newer models have SD card slots; can use SD cards for semidurable storage. They are obviously unsuitable for some things, but fine for others.
And why would the laptops suddenly ship with a variety of modules if they're replaceable? You can still ship it with the same modules in every laptop and get those benefits. And if someone upgrades it, that's still an improvement over no upgrade path, so this makes no sense to me.
Yes, compare the specs of DDR4 and DDR5 to LPDDR4X and LPDDR5X. The latter are significantly higher performance.
This is also the reason that Dell recently introduced CAMM memory modules -- it is an attempt to address the packaging bottleneck that is limiting the speed of DIMMs currently.
The amount of time that it takes a signal to go down a wire has been relevant for DRAM for a while now. If you look at the traces on the board of any relatively modern computer, you'll see some that take circuitous routes, for the purpose of having the signals arrive at the CPU at the same time. You can see this even on relatively low performance devices like a Raspberry Pi. https://www.cnx-software.com/wp-content/uploads/2019/06/Rasp...
Having all the signals arrive at the same time isn't the same thing as having the signals arrive soon.
I'm not sure what "evidence" I could provide that would convince you since we don't have high latency Apple chips to benchmark against. However, there's a reason VRAM is soldered onto GPUs.
That higher frequency helps phones save on pin count by using a narrow memory bus, and allows laptops to have lots of memory bandwidth to feed the integrated graphics when using typical laptop/desktop bus widths.
It's because, since the T2 chip and going on with Apple Silicon, they're not SSDs in the NVMe sense. They're an Apple-specific technology derived from their Anobit acquisition, that only look like an NVMe device to the upper layers.
It's not even clear that this has a performance benefit on typical workloads. At release the M1 was non-trivially faster than contemporaneous PC mobiles, but its competition was also using TSMC 7nm and DDR4.
Now we've seen Zen3+ mobiles on 6nm with DDR5 and upgradeable memory and they're about the same speed for nearly everything despite the M1 being on 5nm, which basically proves they didn't need to solder the RAM.
Whatever ram and SSD config you buy when new is how it will be forever.
I have several old (5-15yrs) laptops still providing useful service. None of them is on the original HDD/SSD.
On the one had it absolutely makes life a lot easier and all the software I need runs great on it. On the other I would by buying into a device I simply cannot upgrade or maintain myself. This makes me extremely uncomfortable.
I dont want to run Windows as a daily driver since it is really jarring for my personal workflow etc. Yet Linux lacks quite a few of the essential pieces of software I need outside of development. E.g. (Krisp.AI, Reincubate Camo etc)
Rock <--> Hard Place
I'm not agreeing with Apple soldering the SSD to the logic board. But they do seem to be significantly more reliable than hard drives.
I've never had any kind of problem with internal SSDs.
https://appleinsider.com/articles/21/06/04/apple-resolves-m1...
But every time I get a little bit paranoid about the idea of "what if it suddenly dies and I loose everything?" and then I start thinking about backuos, at which point I'm like "nah scree this".
Are there any recommendations for particular manufacturers to look for? (Seagate, Samsung, Toshiba etc)
Also, how much is the average lifetime of a consumer-grade SSD these days? I always think they'll die in 5 years, but that's totally out of my head and not from experience.
I've got a 240 GB Intel SSD with 49254 hours on it. That's 5.6 years worth of constant work. Still kicking. Not doing anything heavy though. The only reason it's still doing anything is because it's in a firewall. It does run Prometheus though, so it's not completely idle.
Also, SSD only for music collection seems expensive. I’d get a 2TB (or whatever size seems reasonable) spinner, and put the rest of the funds towards backups (like Backblaze).
WD Blue 1TB Desktop Hard Disk Drive - 5400 RPM SATA 6Gb/s 64MB Cache 3.5 Inch - WD10EZRZ - $50
I can't say what $20 is 'substantially cheaper'. Sure, if you need dozens of TBs...
And currently you have a big chance of getting an SMR HDD, which is... quite a gamble.
I've had great luck with surplus HGST drives, they tend to be more reliable than Seagate.
I wouldn’t recommend relying on a single drive for long-term storage of anything that’s worth more than ~$1000 - neither with SSDs nor HDDs. Freak accidents can happen with any component. You PSU can go rogue and fry your SSD, etc.
The best option is most likely a single local drive with continuous mirroring to a cloud service. The drawback is the ongoing cost of the cloud, and the possibility of partial data loss, because cloud mirroring won’t be instantenous when saving bulk data.
A local RAID1 array is cost-efficient, but doesn’t save you from black-swan events like floods or house fires.
What numbers we have show that failure rates are low, but honestly I would give it another decade or two before using SSD for persistent data storage (of important data without backups to rotating disks). I don't think we're there yet.
Alternatively you could get a NAS, they are great.
Life experience has taught me not to trust detachable drives at all
The primary issue is not the 10 to the minus X failure rate they quote, but the much more likely chance that you will lose access to the data for some reason. For example, the account is hacked, someone closes the account, or the account just deleted/restricted by the provider.
But I’d say it’s unlikely that your local copy and the cloud backup would get destroyed at the same time.
Also, when I said “cloud”, I meant a proper cloud service with an SLA like S3 Glacier, not Google Drive which gets wiped if your Google account is disabled for uploading a YouTube video with background music.
I would look into Blackblaze personal backup for your use case. Just make it a habbit to download and test your backups.
This is not a hypothetical; all drives do fail at some point - it's just a matter of when.
The only way to be safe are external backups, not even RAID (which in unlucky cases, can suffer from serial failure).
If one really doesn't want to spend any money, an option is to periodically backup to an external drive, although cloud backups are relatively cheap and automated, nowadays. But also this is subject to localized mass-failure, ie. burglaries.
Can you tell the difference in "what if {HDD,SDD} suddenly dies and I loose everything?"?
> Also, how much is the average lifetime of a consumer-grade SSD these days?
Years.
Actually with all that TLC/QLC/xLC you can get even less than some old-time ones, but they are cheap and you can get much more than you need, so you will have plenty of resource to spare, ie if you have 250GBs of music - you can buy 1TB and have x4 'over-provision', if you have 1TB - buy 2TB or more.
And as others said, music collection is basically WORM so you probably would have the problem of finding an ancient USB3 controller to attach your ancient USB->SATA/NVME external drive for your new shiny holodeck with USB23423 in 2042 than it to die on you.
> Also, how much is the average lifetime of a consumer-grade SSD these days? I always think they'll die in 5 years, but that's totally out of my head and not from experience.
Some have warranty that long but the fact you wont reach TBW doesn't mean it won't lose data.
JESD218 spec only guarantees year of retention unpowered (for enterprise, 3 months for customer) soooo good fucking luck. Many manufacturers don't even say the guaranteed retention in datasheet.
Also, once you finish setup your backups make a calendar entry to test restores at least every year (preferably more often).
That retention spec is for a drive at the end of its write endurance. A drive that's less worn-out will have longer retention, more or less by definition (reduced retention is the main problem caused by the wear of writing lots of data). Also, IIRC it's one year for consumer and only 3 months for enterprise (albeit at a higher storage temperature).
It's bundled on the Ubuntu installer, though it taints the kernel.
You need a mirror or raid pool. Set the checksum to sha256 before copying your music over if you want the utmost protection.
Buy 2. Sync the devices regularly. Better to have 2 512GB drives and deal with syncing them once a while and half the space than 1 1TB drive and it's all gone due to any issue.
Never completely depend on a single storage device.
If your data is that important you'll want 3 - 2 - 1. 3 copies of data: two in different media, one offsite(offsite includes cloud these days)
Anything can always randomly suddenly fail. Low probability, but reasonably possible. So if it contains anything collectible, you'd want mirroring (ZFS is pure awesome) for availability and separate backups just in case.
While I know that SSDs will die eventually, this has yet to happen to any of my private drives.
Some of my drives are over 10 years old and still work something unheard of when it comes to rotating rust.
And those are regularly used too some with ridiculous uptimes measured in years.
The drives I replaced where changed for capacity reasons and not wear.
Our company runs a bunch of servers with SSDs and even there they hold up pretty well. The last error was caused by a faulty controller some months ago.
Since then I don't buy off-brand drives. I usually use samsung drives and have had no problems.
I stick to Samsung for PCIE4 and TLC+Marvel for anything else.
I've only had an old OCZ Agility 3 64gb die, but it had a cruel life of being a cctv storage drive and still lasted about 8 years.
Off-brand? Did you know Kingston produces >60% of world's DRAM and >20% of SSDs?
I was fairly sure Kingston just slap their brand any old rubbish and sell it.
Anecdotally I've had more Samsung drives fail on me than any other brand. Everything from OEM drives to retail EVO drives.
Happens all the time when you buy storage in batches. We learned a lesson, go through 3-4 suppliers and split your orders over a three month period. Buying from one place in bulk is just guaranteeing data loss in your future.
I've dealt with deployed systems that had (in many different clusters) a total of upwards of 100K HDDs and also with 10K SSDs and for extended periods.
I saw tons of drive failures of many types but never even once did I see two or more drives of the same batch fail soon one after the other.
Individual sellers probably use the same shipping method for every order and they may order in large batches which are transported at the same time. Although SSDs are a bit more resilient, hard temperature spikes in a single batch won't be detected after shipping. In the past, we would do this for hard drives because although parked heads are safe, high G-loads could actually unpark a head.
The hard drives were for bulk storage too. The SSDs have been system drives.
That's false.
I have a dozen original disks from 1988-1994 PCs that are still working to this day, and much more from the late 90's-early 00's.
Before you say it, of course I keep backups of everything. That's why I still use decade old hardware without much worry.
We do have some that sit at 1% life left and refuse to die tho
External SSD's seem to be significantly less reliable. My theory is manufacturers are only designing them for short burst of data transfer. I often transfer 100's of gigabytes at a time and these external drives get extremely hot when doing so.
I've only had one fail several years in, and I'm pretty sure it was environmentals (smoke) that did it in.
Not at all unheard of! HDDs tend to fail young, or last a very long time.
The oldest working hard drive I have is in my SPARCstation10, so they're about 30 years old now.
In my ZFS home server nearly all the drives are from its original build in 2010, so over 12 years old now. I had to replace one drive early in the life of that system (year 2-ish, forgot exactly) and the rest are going strong.
After installing, on the first boot KDE immediately informed me my drive was moments from death. I really appreciated that, as I had no idea. I didn't know a thumb drive even had SMART capabilities. I had been using that drive in Windows for random things and it certainly never told me.
Checking SMART it said something to the tune of data loss expected within 24 hours. Yikes.
My first instinct would be to do the backup (of course), but a close second would be "yeah, that's probably not so precise as to make me be quite that lucky today..."
smartctl --all /dev/disk0
Percentage Used: 2%
Data Units Written: 97,255,284 [49.7 TB]
This is after 1 year of full time development.My drive is 1TB (with 16GB RAM) so it should be 2.5PB Total Byte Written Lifespan.
Percentage Used: 2%
Data Units Read: 1,118,827,429 [572 TB]
Data Units Written: 432,707,145 [221 TB]
Power Cycles: 484
Power On Hours: 1,106
The power-on hours, during which the read-write activity happened, correspond to only 46 days, because the rest of the time the SSD was presumably powered-down by the OS, due to inactivity. Percentage Used: 14%
Data Units Read: 1,101,421,696 [563 TB]
Data Units Written: 902,660,422 [462 TB]
Power Cycles: 238
Power On Hours: 2,489
Assuming the %age used is even close to accurate, I'll be upgrading to a new device long before the disk craps out.https://support.hpe.com/hpesc/public/docDisplay?docLocale=en...
This took down HN:
They're so cheap I don't mind.
Does it fail to read? Does it go read-only? Do OSes actually keep track of the stat and pop up a warning "Only 5% life left!"?
Obviously without the manufacturer's internal debug tools it's really hard to properly root cause issues.
Second failure was complete disconnect from the OS (USB drive) no warning. Once it cooled down i tried again, device appeared and reported no disk installed
https://techreport.com/review/27909/the-ssd-endurance-experi...
Eg: BackBlaze: 2022 Drive Stats Mid-year Review
> As of June 30, 2022, there were 2,558 SSDs in our storage servers. This compares to 2,200 SSDs we reported in our 2021 SSD report. We’ll start by presenting and discussing the quarterly data from each of the last two quarters (Q1 2022 and Q2 2022).
https://www.backblaze.com/blog/ssd-drive-stats-mid-2022-revi...
And the Winner Is… At this point we can reasonably claim that SSDs are more reliable than HDDs, at least when used as boot drives in our environment. This supports the anecdotal stories and educated guesses made by our readers over the past year or so. Well done.
We’ll continue to collect and present the SSD data on a regular basis to confirm these findings and see what’s next. It is highly certain that the failure rate of SSDs will eventually start to rise. It is also possible that at some point the SSDs could hit the wall, perhaps when they start to reach their media wearout limits. To that point, over the coming months we’ll take a look at the SMART stats for our SSDs and see how they relate to drive failure. We also have some anecdotal information of our own that we’ll try to confirm on how far past the media wearout limits you can push an SSD. Stay tuned.
* `/root`, System/OS: fastest SSD with plenty of spare space
* `/home`: 1+ TB for data, games, media, etc...
* swap, `/tmp`, `/var`: separate cheap SSD (or even HDD) for frequent and unimportant writesEven if you have a RAID 1 setup (two SSDs operating redundantly), there are plenty of points of failure, many of which are more likely than a modern SSD failing:
* Theft
* Physical damage
* Water damage
* Data corruption
* rm -rf
You need backups in any case, even if SSDs were 100% reliable.
Preface .There was a test about memory endurance of the SSD, where they were actually written to death and precise number of data was measured (not estimated). That was in the age of Sata SSDs, so not super recent, but memory cells were actually more durable in those times, due to bigger tech process and lower number of bits per cell. In the end 256 Gb drives failed after writing 1-3 petabytes of data.
So when I'm buying 1 Tb modern mid range drive, I expect about 4 Pb of writes on it, but I halve that number because of the worse tech process and more levels per cell, and additionally halve again due to by personal build (SFF PC) operating at higher temperatures, not good for the lifespan.
None of my drives have reached even 10% of 1 Pb yet, so I don't care too much about memory lifespan.
What I want to know is there any general purpose utility that is capable of testing a variety of brands of these devices and doing a decent assessment of their state/reliability?
It seems that drive manufacturers don't provide much help here as I've not seen any such utilites for them. I raised this matter with a SanDisk rep at a trade show several years back and he said he didn't know of any.
Same problem applies to SD cards (SDHC, SDXC, etc.), so any info on them would also be useful.
There used to be a great site before that was a data center publishing physical hdd stats, not looked for it for a long time, but I presume they would have ssd stats these days too.
I also reckon much of the problem comes from the fact that information about them is proprietary—simply manufacturerers just don't tell us much about them as to do so may reveal trade secrets. Several decades ago I was involved in work where we had to have as large storage as possible irrespective of cost, back then a 1GB SanDisk was worth somewhere between $1k and $2k (we replaced single-sep time lapse remote monitoring film cameras with TV and needed large storage).
We approached SanDisk (about the only manufacturer with such large drives at the time) and they were very reticent about telling us anything worthwhile—even though we were a large international organization and had clout. We needed to know the reliability so we had to investigate it ourselves, whilst we made some progress it was never fully satisfactory.
Anecdotal info we learned from various sources was that manufacturerers had ways of testing them by altering the threshold voltage—the point where the gate potential would switch from 0 to 1. At a critical point one could check how many gates failed to switch and this voltage altered over time/with use. Monitoring this could provide useful info such as knowing when to retire a device before it failed.
How accurate this info is I don't know but it seems to make sense. If true, we users should be demanding of manufacturerers utilities that are capable of doing such testing. Trouble is, manufacturerers continue to maintain this secrecy.
PS: several days ago I put a brand new SanDisk 128GB in one of my PVRs and it's really hot to touch even when it's on standby (not recording TV). This isn't the first time I've noticed how hot they get. This isn't the PVR's fault as I've several different brands and the thumb drive gets very hot in each one including my PC. One wonders what this elevated temperature does to the reliability/service life.
Unfortunately, most plastic packaging precludes this. Solder connections would also be a problem.
https://hardware.slashdot.org/story/12/12/02/2222235/self-he...
It may be worthwhile heating it up in the oven and see what happens (I can control the temp pretty accurately). Take the case off first etc. If it falls apart or the solder melts nothing's lost over and above what's happened already.
I'd recommend instead (or at least first) desoldering the component memory chips and trying to read the data off of them directly with a microcontoller (bit-banging whatever protocol the drive uses internally). It's more (and slower and fiddlier) work, but also more likely to get at least some data back.
The only reason for suggesting that course of action was out of curiosity, as I recall from old data would sometimes come to 'life' after we erased 2716s, 2732s etc. and then exposed them to excessive heat (it was never intended as a means of unerasing them after we'd exposed them to UV light.
The real issue remains and that's that manufacturers aren't prepared to tell us anything about them and I reckon that's a significant problem.
In that case feel free to bake away with impunity. The warning was really only relevant in the context of possibly needing to recover the data.
While pretty much every SSD controller by other vendors does it by itself because bringing it to OS level is a waste of CPU cycles.
They usually fail early or very late. Assuming good cooling.
It would seem that the vendor is incentivized to not report the most accurate data so their drive doesn't come across as bad and/or to avoid warranty claims.