Overwriting Hard Drive Data: The Great Wiping Controversy (2008) [pdf]
vidarholen.net
vidarholen.net
dd if=/dev/zero of=/dev/sdX
See also [1] for the non-scientific claim backed-by-bounty.> It is unlikely that an individual write will be a digital +1.00000 (1). Rather - there is a set range, a normative confidence interval that the bit will be in [15]. What this means is that there is generally a 95% likelihood that the +1 will exist in the range of (0.95, 1.05) there is then a 99% likelihood that it will exist in the range (0.90, 1.10) for instance. This leaves a negligible probability (1 bit in every 100,000 billion or so) that the actual potential will be less than 60% of the full +1 value. This error is the non-recoverable error rating for a drive using a single pass wipe
It all comes down to -- what is your threat model? Dishonest craigslist scroungers or a well-armed state? The latter just may be able to recover things you'd be astounded of, but if it's easier for them to do it with Van Eck phreaking or a subpoena to the sites you mirror with, they'll just do that.
Note that AFAICT this paper does not discuss NAND or phase-change memory or anything else. Today in 2017 the levels of indirection between you and your actual real NAND cell that you erased and programmed six months ago are significant. You might actually want to take a hammer to your SSD.
[1] http://www.hostjury.com/blog/view/195/the-great-zero-challen...
100,000 billion bits is 100 terabit, or 12.5 terabyte. So, there will be, statistically, a single zero bit on a 12 terabyte disk.
Worse, AFAIK, modern disks use error-correcting codes to get the insane bit densities we find normal nowadays. That means you likely will get a single bit failure in a written zero that then gets corrected back to a zero.
There is a much larger risk, though: when physical sectors fail, modern disks will map spare logical sectors in their place. Specialized tools likely will be able to recover data from sectors mapped as 'bad'.
The solution is to use full-disk encryption with the passphrase not present anywhere on that hardware. In that case, if someone gets to the hardware but powers it off when they do it, you can be sure that the data is safe.
For new drives I started using full-disk encryption. And the hammer.
Also keep in mind that with SSDs that have significant smarts for compression, trim, deduplication, benchmark cheating, minimizing write amplification, and related. Writing a pseudorandom sequence is more likely to actually overwrite almost 100% of the disk. Sure bad sectors that are remapped won't be, but the rest should.
Even better if you just use full disk encryption and throw away the key. Also the "secure erase" functionality is a good start, but if you aren't the trusting type it's best to wipe then overwrite.
Not sure why some have downvoted me. It's easy to verify.
time dd if=/dev/zero of=file1.bin bs=1M count=600
600+0 records in
600+0 records out
629145600 bytes (629 MB) copied, 1.57948 s, 398 MB/s
real 0m1.651s
user 0m0.004s
sys 0m0.688s
time dd if=/dev/urandom of=file2.bin bs=1M count=600
600+0 records in
600+0 records out
629145600 bytes (629 MB) copied, 37.7828 s, 16.7 MB/s
real 0m37.786s
user 0m0.004s
sys 0m37.784sI do wonder what machine you used that's 20 times slower than a 4 generation old i5.
openssl enc -aes-256-cbc -in /dev/zero -out /dev/sdX
This post appears to have the reason:
https://superuser.com/questions/359599/why-is-my-dev-random-...
(This raises the question of why using openssl (per other comment) makes it faster.)
An all zeroes pattern would look nothing like all zeroes on the media. Modern read/write channels use a scrambler, RLL, and LDPC, which scrambles the data, limits the length of long magnets (since the heads don't like DC signals,) and then parity bits. From a signal processing perspective, writing random data will not look much different from all zeros.
One thing to keep in mind, however; if you tell the drive to write a large range of LBAs to the same pattern, depending on the firmware, it might only write one physical sector and point every LBA in that range to that physical sector. So now your data only has the first sector overwritten. If someone hacks the firmware in a way that allows reading of physical sectors, they can get your data.
That said I've typically followed any overwrites, random or zero-based, with use of ATA secure erase commands, which might be equally bad.
The authors dismiss the security value of wiping a hard disk, based on their thesis that weakly-deleted data cannot be recovered without a priori knowledge of that data's content.
They argue the requirement of a priori knowledge of the data to recover negates the security risk of said recovery; this--they argue--reduces the threat model to more of an academic exercise.
What the authors totally neglect, however, is the security risk of confirmation: the risk that an attacker might confirm that the target hard disk did, indeed, store certain data, where the content of that data is known a priori.
Example: Say I have obtained a trove of private incriminating documents associated with some anonymous person, X. I suspect, but don't know, that X is my target, Bob. I would like to prove that Bob is X, and X is Bob, so that I can definitively pin X's crimes on Bob. Say X uses some electronic signature to authenticate his original work as his own. If Bob is X, I should expect Bob's hard disks contain a statistically aberrant abundance of copies of X's signature.
Thus, to pin X's crimes on Bob, if Bob is indeed X, it is sufficient to recover data from Bob's hard disk--data of which I have complete knowledge a priori--namely, X's digital signature.
While I take no issue with the facts, I find the author's conclusions reckless. It seems in their haste to "bust the myth," they extend their result beyond its valid range of application. What could have been a useful clarification on the low risk of _unknown_ data recovery has become a wild and dangerous generalization, 'debunking' best practices.