Par2cmdline is a PAR 2.0 compatible file verification and repair tool
github.com
github.com
Maybe I used it differently: I'd have a collection of dozens or hundreds of files, and if some small portion of those files were missing or corrupt, the recovery would draw on the overall parity files.
With this tool, I might be misreading, but I think if a single file went missing, the recovery files would only be for that single file, so recovery would be impossible.
Did you try manually creating parity for that file (using par2cmdline)?
Edit: filed a bug report
My backups consist of compressed and encrypted archive files, which are then processed with PAR, to add redundancy that will allow the recovery of any corrupted files from the archive.
Always remember that tools like PAR are of limited system if the filesystem becomes corrupt.
Also worth mentioning:
Is that because of the time they were written in, or is that more suitable than other codes?
Would using e.g. LDPC codes work just as well?
The total possible number of blocks I believe is 16k for par2, or some similar number. The specific number of blocks for each file can be configured, and the tradeoff is that the more blocks you want per file, the slower the generation will be.
Then it generates some parity data. Using that parity data it is possible (and guaranteed, if the parity data itself is not corrupt) to detect and completely restore 1 or more complete blocks of the target file.
The more parity data is saved, the larger number of blocks that can be restored with it. But it doesn't matter which blocks get corrupted, or how they are corrupted. For example, if a given parity data supports restoring 10 blocks, any 10 blocks of the source file can become corrupted, in any ways (even completely missing), as long as the corruption is contained in those blocks - the complete original file will be possible to restore.
AFAIK, some other codes are possibly better suited for other tasks, such as stream redundancy measures in radio networks, or other applications where the data corruption has certain known and restricted parameters. But for general purpose file protection, bitrot, bad blocks, bad drives, scratches on an optical disc, download errors etc - Reed-Solomon is very well suited for the task.
That said, having recently looked at some 20-year-old discs, professionally pressed discs held up well, and bargain basement CD-Rs had a lot of physical failures, even when stored indoors in a dark case. So par2 is good, but make sure you're getting high-quality media, too. Maybe M-Discs?
For long term storage, I would recommend ZIP files + PAR2 over an EXFAT partition
https://jrs-s.net/2016/05/02/zfs-copies-equals-n/
"zfs: copies=n is not a substitute for device redundancy!"
The point of copies is belt + suspenders: you both have your typical raid protection as well as two individual copies of the file so if one copy becomes corrupted the second can be referenced.
Additionally it's useful in the case you have a single disk (like a laptop). That way if you end up with a corrupted block you can still recover your file.
I was pretty annoyed by the whole thing.
Also interesting to find out AWS uses their customers' resources for temporary storage in the transit process, instead of elastically using process-only-bound ephemeral resources outside customer space in a more cloud-native fashion. Temporary consumption in the customers' resource space in a solution pattern gives me nightmare scenarios of stray code that scribbles the temporary objects into customer-owned data, or accidentally dropped into the wrong location and read by customer processes. Would be curious to hear the trade-offs involved in that decision, they could not have made it lightly, I always try to choose fail-safe design modes and at that level of solutioning I'm sure their teams are way smarter than I am so I'd love to learn from this use case.
What are people using it for these days?