What every programmer should know about solid-state drives (2014)
codecapsule.com
codecapsule.com
How can they get that if they stuff enough fake reviews, plus the legion of consumers who would have no idea that the drive was the issue and not "viruses".
No one is going to be selling a million laptops with a drive from RandoDriveManuGoodBrand off Amazon with no track record and no validation.
Anyone buying the non-name brand types of drives knows they're getting (at best) something that might only work a little while before exploding.
The name brands like Samsung, et. al. work hard to make their firmware not grenade something (and the drives overall to be AT LEAST as reliable as their competitors) BECAUSE they want the name to mean something. It is what drives customers their way, most of the time.
If they get a reputation as a company selling junk (cough Deskstar/Deathstar) that costs them billions over many years.
I remember an adjacent team to mine that had to store several gigs of data which changed often, but only a small percent changed at any one time. They needed to recover quickly from a crash so they wrote it to disk. But they wrote the entire data set out to disk after every update, instead of keeping it in e.g. rocksdb or even sqlite. Their entire fleet burnt through their SSDs at about the same rate, so machines were dying in rapid succession, ouch. Write amplification is a real problem, but SSDs great performance often masks it being an issue until down the road.
You can burn out a modern consumer drive in 2 days if you want to. Write perf ~6 gb/s, mtbf 700 tb written on a 1 tb drive. The tlc/qlc cells have very poor endurance imo.
[1]: https://du.nkel.dev/blog/2021-05-05_proxmox_influxdb/#config...
https://www.pcgamer.com/storage-study-finds-ssds-might-not-b...
"Are SSDs Really More Reliable Than Hard Drives?"
https://www.backblaze.com/blog/are-ssds-really-more-reliable...
Very confusing and might be incorrect. What are planes. And are pages made out of blocks or vice-versa? If blocks are grouped in pages, with erasing it sounds very different.. Only whole blocks, which sounds like blocks are bigger than pages.
Planes reflect the physical structure of the storage chips: there's multiple layers that share a common vertical bus.
Plane > Block > Page, that is to say Blocks are always made up of multiple pages (commonly 128 or 256 as the quote mentions). Pages are the unit of read and write, while blocks are the unit of erasure. The FTL tries to hide this page write vs block erase mismatch as best it can, but as the original article points out you may need to be aware of what it's doing in very high performance systems.
A drive with 8 dies each having 512Gbit capacity divided into four planes per die will perform almost as well as one with 16 dies of 256Gbit divided into two planes, other things being equal (eg. number and speed of the channels between the SSD controller and the NAND, page and block sizes and access times, all of which are subject to change at the same time a generational change increases die capacity and number of planes).
How do I tell my SSD to write stuff to specific pages? You can't really tell the SSD to do anything except read, write, or trim LBAs.
Does NVMe support this with its queues?
> 27. Over-provisioning is useful for wear leveling and performance
I thought most if not all SSDs were already overprovisioned. Does additional overprovisioning help?
> To ensure that logical writes are truly aligned to the physical memory, you must align the partition to the NAND-flash page size of the drive.
I think this is false. This assumes there is a one-to-one mapping of LBA to SSD PBA which you don't know. LBA 2048 could go to any PBA on any page/block/flash line in the unit and as things are written and rewritten, any correspondence that might happen due to sequential assignment of PBAs->LBAs would gradually diminish, IF you knew for sure that was happening in the first place. Because you wouldn't really know what the SSD is doing without reverse engineering or seeing the source code of firmware, unless there's things going on in NVMe land that are new and I don't yet know.
I think a big extra helping of overprovisioning is one of the major differences between consumer and enterprise SSDs.
https://www.anandtech.com/show/11436/nvme-13-specification-p...
https://www.anandtech.com/show/14543/nvme-14-specification-p...
https://www.anandtech.com/show/16702/nvme-20-specification-r...
https://www.anandtech.com/show/15959/nvme-zoned-namespaces-e...
All of the current HW-level performance hacks could actually get in the way if your software already enforces things like single writer, chunky writes and/or append-only log structures.
Give me a drive that only writes in 1 linear direction (until its full) and has a big red button to clean the entire thing all at once (which would clearly require some offline processing time & multiple disks for a realistic system).
https://nvmexpress.org/new-nvmetm-specification-defines-zone...
> In December 2012, Taiwanese engineers from Macronix revealed their intention to announce at the 2012 IEEE International Electron Devices Meeting that they had figured out how to improve NAND flash storage read/write cycles from 10,000 to 100 million cycles using a "self-healing" process that used a flash chip with "onboard heaters that could anneal small groups of memory cells."
So can I apply this myself by placing an SSD drive in an oven?
What every programmer should know about solid-state drives - https://news.ycombinator.com/item?id=9049630 - Feb 2015 (31 comments)
Maybe it's useful if you want to make something like a more performant version of grep? (aka ripgrep?)
There is probably a small but non-zero number of these on here.
If you're programming at enterprise scale, this sort of stuff is the responsibility of architect-level programmers and senior systems engineers.
Even most linux sysadmins know all about block alignment (well, if they predate most of the various tools figuring out block size/alignment stuff for you.) It's nothing new - RAID arrays work best when properly aligned, for example.
This is why we can't have nice things.
It’s not like reading 10 bullet points on the subject is “diving deep” and making huge time investment.
It’s just getting the minimal context, so later on at least some keywords are known.
Not sure how many computer related topics you know/want (“The more you know, the more you know you don't know”), but for me, 50 topics on programming seems sufficiently high at frankly a very low effort/commitment.
True, but you're using so many abstractions that the rule can't feasibly be "read a short summary of every abstraction you're using." There are just too many. At some point you have to choose a threshold where the likelihood of an abstraction leakage is sufficiently low. When you're debugging a CSS selector you will almost certainly never need to know about even the existence of, say, Fermi–Dirac statistics.
Rule - no. Goal - yes.
Some topics are more stable and valuable then others, so prioritisation helps.
“How utf8 generally works” vs “implementation details of js-node-utf-related-library-X.”
But it’s rarely because some developer didn’t understand page caches, and usually because it obviously didn’t revive enough QA or UX input.
Choice of db schema impacts physical layout on ssd. E.g. Different tables are more likely to be on different ssd pages resulting in random writes.
Databases are insanely complex, but not magic.
Makes sense to me. At Google we were told to stop thinking about all this stuff, that the storage hardware and software people were responsible for hiding things like wearout from application developers. This article is really "things you should know if you plan to directly access an NVMe device" but there is a huge class of programmers who are better off not knowing.
and as a result Chrome slams SSD by writing cached Youtube videos to disk .... except Youtube never reuses cached video data (not even when rewinding more than couple minutes to already watched spot in same video), it explicitly generates hashed requests with custom URL parameters googlevideo.com/videoplayback?expire (~6hour shelf life) &range &sig &lsig. Heavy YT viewing results in wearing out your SSD by tens of gigabytes per day for no particular reason. This is just one small example of side effects from such brilliant decisions.
1-13) General background info that informs the rest.
14-25) Important for any programmer that does enough file IO that they need to optimize it.
26-29) Important for any system admin to ensure they aren't inadvertently limiting the performance of their hardware.
"In a time of SSD, multi-core/processor, two terabyte memory and Optane App Direct Mode machines, there is no reason not to build from BCNF data. Time to do what Dr. Codd demonstrated. Technology has finally caught up with the maths."
https://drcoddwasright.blogspot.com (skip the distractions)
https://github.com/bradfa/flashbench
Plus, alternatively, there's FlashBench:
https://github.com/JonghyeokPark/FlashBench
These might be found useful for determining the underlying structure.
Enterprise class 3D TLC NAND is relatively close to enterprise class non 3D MLC NAND, the gap is bigger for consumer drives.
But I think as of 2022 only Apple still sells consumer desktops/laptops with entirely TLC NAND. Everyone else is racing to the bottom for their consumer stuff.
Basically I want what every programmer should know about storage but in the style of dreppers original article.
Why not? Blog posts aren't nearly as valuable.
Also since we’re talking about hardware, I imagine a lot of people with necessary domain knowledge can’t share what they’ve learned done because of IP restrictions.
Max throughput is around 6gbps with a fairly high latency. DDR5 has speeds of 52gbps, lower latency, AND your CPU will almost undoubtedly have a cache on it to increase that speed further.
This is all assuming you are putting your mem device on a pci-express bus.
In the consumer market, a number of performance NVMe drives will hit over 5GB/sec, which would be 40 Gbps.
The latency isn't anywhere near as good as even quite-old RAM, but modern SSDs are considerably less than an order magnitude off in transfer speed from even current, common ram (DDR4) and "only" about a hundred times higher in latency than RAM.
That's pretty stunning from mass storage. So is well over 500,000 IOPS.
Backup your stuff
Backup your stuff! I happen to also back up to an SSD these days, because the difference between minutes and hours is hard to argue with.
†edit: history of shipping with an SSD standard, that is.
If the backups are incremental it shouldn’t take hours.
Yes. With an SSD the enemy is electron leakage. Minute quantities of electrons trying to escape an unnatural state and return to equilibrium. (yes, I just anthropomorphized electrons.) Magnets however are more stable by nature. (yes there is nothing natural about hard-drive storage. SMR doubly so!)
Anecdote/anecdata: I have been able to retrieve full drives worth of data off of drives that have sat in a cardboard box for 10 years. I also have trouble accessing data on 1-year old USB flash drives.
But multiple copies in multiple formats cannot hurt, and the most important stuff should have multiple live copies.
Even hard disks should be powered on occasionally to test backups.
Those big backup HDDs use shingled storage, so they're not any good as general purpose hard drives, but they're excellent for strictly sequential writes, such as a full disk backup to a single file.
-
Oh, what my routine is? Uh. I `cp -a ~ /mnt/backup/date` a couple of times a month.
... Testing backups?