Unbalanced reads from SSDs in software RAID mirrors in Linux
utcc.utoronto.ca
utcc.utoronto.ca
http://www.rkeene.org/viewer/tmp/linux-2.6.35.4-2rrrr.diff.h...
Thus balancing reads is good for SSD endurance.
Ergo, non-balanced reading is better for endurance of the array, as it reduces the probability of contemporaneous multi-disk failure. Like disk 2 failing under the read load encountered during a rebuild.
But, it's possible that this solution is not particularly useful depending on what mixes of equipment and workloads one expects to operate.
You'd need some kind of cache below the md driver to cache "metadata relevant to the SSD innards". Since there isn't one therefore there can't be a penalty due to caching to switch between SSDs in a mirror. That answers your question.
But since you want an insight, here's one for you. There is no "metadata relevant to the SSD innards" available. The md drivers do not have this information. I wish they did. When I was still playing with the md code I'd love to know the block and page sizes of the drives. But that information is two bytes. Not really something that would need a cache.
I'm trying to figure out how, exactly, you've read my attempt at excessive clarity; it's not clear to me yet what you misread, since it seems like you're trying to correct some misunderstanding, but it hasn't worked because you haven't said anything that isn't already blindingly obvious.
There is no "metadata relevant to the SSD innards" available. The md drivers do not have this information
This is not an insight.
The cache I'm interested in is tied up in how the SSD presents as a block device, but isn't implemented as a block device, as block devices are normally understood (wear levelling and remapping, hidden parallelism / striping, etc.). Yes, the FS cache can't help with this. Yes, the RAID implementation can't help with this. Of course. It's internal to the SSD. It's innards. I don't see what information you added.
Now that I know what you're talking about, let's try to answer your question. There are SSD controllers that don't use a RAM cache on the drive. It's not necessary for performance and doesn't do as much as on a HDD if present. The main benefit of a drive cache while reading is for read ahead. This is not needed for an SSD as it doesn't have to wait for the sectors to show up under the head.
The only thing that will affect things is reading file system blocks from the same SSD page/block. The md code already does this. If the next block requested is after the last blocks read it will use the same SSD.
If the block is from somewhere else then it doesn't matter if it comes from this SSD or that SSD. Access time will be the same.
Agree and this would seem to make sense.
It is like having two rolls of paper towels in the kitchen. If you use both of them you are only extending the time till empty. If you use one of them only you can then replace that when it runs out running on the 2nd in the time it takes to replace.
Not that it is really critical, read-disturb exists but it is on a far lower scale compared to erase endurance.
Then again, not sure if author was really thinking about these sorts of things, who knows.
When I last looked at this, devices could suffer (say) 3K writes, and once written a cell would have to be re-written after 3K reads in order to maintain proper voltage levels.
It's much, much harder to keep track of logical block location, especially if you need to handle surprise removal of power.
Is there a good reason at this point for Flash controllers to bother doing this part at all, instead of:
1. maintaining the per-block erase count metrics internally;
2. exposing the raw physical blocks over ATA;
3. adding an ATA protocol extension that returns the erase-count for a block whenever it's read/written?
Under such a system, the OS would maintain the logical block map (probably as a "layer" in logical volume management, mapping a wear-prone PV to a wearless, but gradually shrinking LV.) It could even put the logical-block-map metadata on a different disk, or pool the logical block maps for a bunch of disks together for a RAID-set and then RAID-mirror the pool, or whatever else. Basically you could play around with such maps the way you can play around with Linux's thin-pool metadata volumes.
1. How many bits of error did we detect in the read? Maybe it's time to relocate the data before it dies completely?
2. How long does the write takes? Maybe it is time to retire this block because it is near dead for unknown reasons?
3. How many times was this block read? Maybe it is time to relocate it before we risk read-disturb issues?
4. Try to distribute data evenly across the flash dies?
5. Ensure we have enough erased blocks to write into to avoid blocking new writes.
6. Can I do some background operation? I really need to garbage collect and rescue some nearly lost data.
What you suggest (and I entertained for a while too) is to return to the pre-IDE (RLL/MFM) days where the OS would do all the hard work and get all the flexibility. But the OS vendors do not want that complexity, with the more advanced technologies involved it is hard work and more importantly, it varies between vendors and for flash it can even need tuning for different flash generations or even batches. The OSes don't want that complexity.
On the other side, the different SSD vendors want to use the lowest flash chips and get maximum revenue. The algorithms that they implement are their trade secrets and the thing that makes them able to charge more for the better devices. If they were reduced to a flash dealer they would find themselves getting less money.
I do think that having a nearly raw flash interface and having a better processor above with possibly better algorithms would yield better results (especially if you integrate it deeper into the user application) but the complexity gets that much higher as well for the entire app and when (not if) there will be another technology the entire app will need to be rewritten which sucks.
I added round-robin code and experimented with splitting large sequental requests between idle disks.