Long story, short, buy 'industrial' SD cards if you care about them not getting corrupted.
Long story, short, buy 'industrial' SD cards if you care about them not getting corrupted.
Ended up mounting /var on a tmpfs - ensuring practically no writes to the card - and fetching the device configuration from a server in the network at boot. PLenty of work, with zero (or negative) profit for the company, but at least I learned a thing or two doing it.
We wrote simulation tools for typical access patterns and ran them heavily for testing out new cards, both for performance and failure rate. Luckily there was enough margin on the devices that we (R&D) could select the more expensive industrial graded cards by good manufacturers, but I did get to see some really shitty cards that purchasing preferred (because they had gotten a great price on them).
However in our case the corruption was due to a buggy FAT-driver for the obscure RTOS we used. In the end though I learned a lot about FAT which was fun (<sarcasm>and is a highly marketable skill these days</sarcasm>).
Turns out the card itself was shit and replaying the trace would just kill arbitrary cards.
Note that all these techniques only show you the read and write calls made to the library/OS and - importantly - not what actually happens to the card. To see that, the next level down is to instrument the card driver to track the actual I/O operations (i.e. so you see what the card is really being asked to do, sans all the caching and buffering).
Note that's not the end of the story; there's what the hardware controller decides to do and when the hardware actually reads/writes the flash array. That's the level where the quality of the firmware in the controller(s) matters.
For the first 4 files with identical prefix use the ~N scheme as in LongFile.txt -> LONGFI~1.TXT For the following files it was: 1) "LONG" + hex_str(hash(long_file_name)) + ".TXT" 2) if there's a name collision repeat step 1.
Compound that with a the fact that hash() was implemented something like:
int h = 0;
for(int i = 0; i < strlen(long_file_name); ++i)
{
h = (h + long_file_name[i]) ^ 0x42424242;
}
It's pretty easy to see where this falls apart...Image the system, and confirm the checksums on first initial boot, then never re-write any of them.
The rpi2 seemed to be the worst for fs corruption, I've never had a pi1 fail on me. Jury is still out on the pi3.
https://www.cactus-tech.com/resources/blog/details/slc-pslc-...
The prices seem to be trending up for SwissBit.
I'll just say that all consumer level cards pretty much don't really care about your data, even the good brands.
For how to use what you've shared if we trust you, can you throw us a bone more than:
>buy 'industrial' SD cards
For example can you name a specific card? Or give us enough clues that we can do so?
I for one take your advice very seriously but I don't know how to use what you've just shared. I've seen ATM-style Pi-based kiosks with corrupted SD cards that wouldn't boot. It looked expensive.
But in general 'industrial' is the keyword to get the good shit from manufacturers who'll treat you like an adult.
https://www.digikey.com/products/en/memory-cards-modules/mem...
Digikey[1] and Arrow[2] generally sell the 4GB for $15 and the 8GB for $25. They make larger versions too, if you need them.
The 'A' stands for aMLC, they're using normal MLC flash (which most consumer SD cards no longer use, but is generally much more reliable than TLC) but they use it in 1-bit per cell mode like SLC. They make traditional SLC cards as well, but the price skyrockets.
The aMLC cards have very good endurance ratings, but they're still cheaper than SLC cards. The firmware and controller are designed to prevent sudden power loss issues, which is apparently the root cause of a lot of SD card corruption on the Pi.
They're also supposed to have lifetime (i.e. SMART) monitoring, but it's a vendor specific command set rather than something smartmontools can read. ATP has a tool for it that probably only runs on Windows.
I've been using those aMLC cards in a bunch of Pi3 and Pi Zero W devices for months, I've never seen them become corrupted or fail to boot even once, despite being pretty hard on them, compiling stuff, yanking the power, etc.
For comparison, a Samsung Ultra+ card became corrupted after a single power loss. The device was running Windows 10 IoT Core at the time, it never booted again and had to be re-flashed.
[1] https://www.digikey.com/product-detail/en/atp-electronics-in...
[2] https://www.arrow.com/en/products/af4gud3a-waaxx/atp-electro...
The cheap SD cards just aren't designed for anything except being used in consumer devices with batteries, where sudden power loss is rare and losing data isn't going to cause a plane to crash or result in someone not receiving a dose of insulin.
So when they suddenly lose power, they aren't always capable of ensuring that whatever task they were carrying out at the time is actually completed and did not accidentally destroy data.
And apparently the consumer SD card controllers are really there to manage and remap parts of the flash that were defective before ever leaving the factory.
It's probably cheaper to build over-provisioned cards with a simple controller that can deal with manufacturing defects in the field, than to do QA on 200 million thumbnail sized NAND die every month and still try to profit while selling them for a fraction of a penny each.
EDIT: and in no cases I saw did the cards I was testing give me truly 'corrupt' data. Just either error codes, or stale data, or occasionally data from another sector entirely. They've got metric shittonnes of ECC internally (to make up for the crappy NANDs), and will do a better job than you can at detecting errors.
(I also want a filesystem that does this so that you have room for a proper authenticated encryption mode for your full-disk encryption - if your apparent block size is the same as your physical disk block size, either you have no room for an authentication tag and you're using a pretty fragile scheme for making your ciphertext tamper-resistant, or you kill performance because you need to read the authentication tag from another block. Current disk encryption software tends to choose the former.)
In my experience, flash corruption of the type found in SD cards are completely blank (00) or erased (FF) blocks, not single-bit errors. Remember that SD already has a layer of error correction to handle those from the raw flash.
XFS has metadata checksums, enabled by default when using xfsprogs 3.2.3+ but data is a much bigger footprint so you can still get hit with silent data corruption. And ext4 any day now is going to start to default to metadata checksums as well.
For those file systems, you can use dm-integrity or dm-verity. https://gitlab.com/cryptsetup/cryptsetup/wikis/DMIntegrity