yes | openssl enc -aes-256-cbc -out /dev/test-drive
and then decrypt and grep. It's almost as fast as badblocks.
yes | openssl enc -aes-256-cbc -out /dev/test-drive
and then decrypt and grep. It's almost as fast as badblocks.
You assume incorrectly. There have been attempts, like SandForce drives in the early days of SSDs but there are a bunch of reasons why you don't want to do opportunistic general-purpose compression at the lowest possible level of the storage chain.
The problem with compression* is that instead of having one drive capacity number (say, 500 GB), you have two: the drive's physical capacity and it's "logical" capacity. Which is fine and dandy, except that the logical capacity varies wildly with what kind of data is stored on the drive. Highly patterned data (text and most executables) compresses very well. Data which is already compressed (most images, videos, archives, etc) does not. So how do you report to the user how much space they have left on the drive? Even if you know the existing contents and current compression ratio, you don't know what they're going to put on the drive in the _future_. Your best guess would be, "250 GB or maybe about a terabyte, lol i dunno".
There's also the fact that application-specific compression algorithms tend to do FAR better than general-purpose compression algorithms, which is why almost all of our media storage formats default to using them. JPEG, HEVC, and so on. Plus you get the benefit of having the thing compressed all the time (even over the I/O channel or network) instead of just on disk.
Compressing data which is _already_ compressed often results in additional overhead. So unless the drive is testing the data to see how well it compresses before writing it to disk (which would murder performance), your 500 GB drive could actually end up being a 450 GB drive.
Further, always-on compression would result in a substantial performance penalty unless special silicon is crafted to handle it. Storage is already an industry with razor-thin margins, companies are not going to add to the BOM cost for a feature that could ultimately make people buy _fewer_ drives.
In the case of deduplication, there's no OS-level standard for it which makes current implementations far less efficient than they could be.
That said, data compression is very popular in the enterprise storage space, but it is typically done at the pool or volume level (large groups of disks) rather than per-disk. These arrays usually combine compression and deduplication with other strategies like thin provisioning to optimize storage to an almost absurd level. It typically requires trained storage engineers to manage them.
* _I lump deduplication into compression because dedeuplication is actually just one kind of compression strategy, even though lots of things treat it like a separate feature._
SMR maybe, I don't know if it's really worth it.
For a hard drive dealing with 4KB sectors, that's a lot of trouble for almost no benefit.
If you write random bytes then read them back, you have nothing to compare them to.
1. Test N blocks of the disk, where N is small enough that you can keep a 128-bit hash of each block of random test data in memory for the duration of the test. You can use the hashes to verify that read data is correct.
2. Then test another Nb/8 blocks using the N blocks from #1 to store 128-bit hashes of the Nb/8 test blocks.
3. Then test N(1+b/8)/8 blocks, using the blocks from #1 and #2 to hold the hashes.
4. Then test N(1+b/8 + b/8^2)/8 blocks, using the blocks from #1 and #2 and #3 to hold the hashes.
...
(Maybe replace the 8's in that with 7's, and in each block of hashes include a hash of the 7 hashes, so you can check that hash blocks are reading back fine?)
tee /dev/the/disk </dev/urandom | sha256sum
and
sha256sum /dev/the/disk
should also work fine.
SSD blocks can be written 10k, 100k, maybe a million times. If one block is rewritten every time the most frequently changed directory is changed, that block might be exhaustible. But starting out with writing each block once won't make a difference.
It's not 2007 anymore.
SSDs also have wear leveling; you cannot wear out a particular block of flash by writing to the same LBA a bunch of times.
Challenge accepted, I tried to wear a device out (my setup had direct bus access to the flash part in question, no remapping layer). Went for months without an error and I eventually repurposed that corner of my desk.
30 years later, they've pushed the technology a little :-)
You won't get 10K out of an SLC cell (ten years ago those would suffer around 3,000 erase cycles, and it has to be much worse now). I don't know what QLC endurance is (16 voltage levels in a cell, run away), but I wouldn't be surprised if it was a few hundred cycles.
10k P/E cycles is the right ballpark for SLC 3D NAND these days, and it's never been anywhere as low as 3k—that's more like what good TLC 3D NAND gets. QLC is in the 500-1000 cycle range.
I'm surprised QLC gets as high as 1000, frankly. It's like the cells are on a first-name basis with the electrons they hold.
Indistinguishable from magic, man.
[Cue future archeologists: "Remember when our grandparents could only get a few gigabytes on a quark? How primitive, how sooo last-femtocentury."]
You nerd-sniped me. I couldn't resist calculating it. A femtocentury is about 3.16 microseconds.
Kioxia (formerly Toshiba) is sampling their version of the same basic idea, and SSDs using that memory should be coming out during the next few months.
Intel has of course replaced SLC NAND with their 3D XPoint memory, which is about twice the price but can sustain better write speeds because it doesn't have flash memory's slow block erase operations.