Zfec – Efficient, portable erasure coding tool
github.com
github.com
I believe it's compatible with the Java implementation used by Backblaze and others.
I wrote an MIT-licensed ReedSolomon module for Node [1], with native C++ multi-core throughput. The internal benchmark (node benchmark.js) provides latency and throughput measurements for data=10, parity=4 for a variety of shard sizes, and the number of data/parity shards can be adjusted for comparison.
It's not yet using optimized assembler or SIMD like Klaus Post's implementation, but it is compatible and also based on Backblaze's implementation, as well as including a pure Javascript implementation for browser use.
http://nisl.wayne.edu/Papers/Tech/code-pf-fast09.pdf
Luigi Rizzo is the original author of the fec C library.
I co-maintain zfec (and the Debian package of it) along with Zooko. A new release of zfec is planned to be out soonish.
edit: removing the <> delimiters around the URL. Sorry!
In my library, the SSSE3 code also has AVX and AVX2 equivalents, providing higher throughput. There are also implementations with specific optimized instructions for ARM NEON, ARM64, PPC64/AltiVec and Power8. All of this in C (with intrinsics), not ASM, so making use of your C compiler's instruction scheduling, loop unrolling,... See https://github.com/NicolasT/reedsolomon/blob/master/cbits/re...
BSD licensed
Pluggable erasure coding backends, and supports Intel's ISA-L
There is also a Python library available https://bitbucket.org/kmgreen2/pyeclib
For backups, what you want to generate additional files, par2 style, which can be combined with the regular archive files (eg like in duplicity, GPG-encrypted tarballs) to repair any bitrot or lost fractions of those archives. But with zfec, as I understood it, it only works if you turn them entirely into erasure-coded shares and assume you'll lose entire shares and wind up using m-of-n shares in recovery/restoration (as opposed to PAR2 which will handle losing individual archive files but also arbitrary corruption in any of those archive files, as long as the total does not exceed the redundancy %). So you would have to do something like create a single giant archive file, and if you wanted the equivalent of 50% redundancy, you'd use zfec to turn it into '50-of-100' shares and backup the 100 shares (or do 500-of-1000 etc).
Then any shares which aren't bit-identical to get tossed out at restoration and you hope you have enough shares left to reconstruct the original giant archive file. But you might not - your shares might be mostly good but missing a few chunks, which is something a PAR2 setup could cope with (if the chunks are in the archive files, they get repaired normally, and if they're in the PAR2 redundancy, they don't matter). This is fine in the Taho-LAFS or datacenter setup where you assume your storage nodes will store perfectly or fail entirely (since given any live errors, you can just scrub the entire machine and rebuild an additional copy from the erasure-coded shares). Not so much in a backup setting.
Ultimately, while zfec is way faster than par2, I found this setup squirrely enough that I wasn't convinced to switch.
Imagine a high-level, straightforward imperative language X whose only purpose is to specify libraries and data structures. This X has strong typing, memory safety and maybe also memory alignment directives but abstracts away as much as possible from any actual environment/platform/CPU layout. It has a precise semantics and even some ways of specifying pragmatics (e.g. big O runtime behavior of functions). In X you write "pure" algorithms that do not access (m)any OS-dependent features. Libraries in this language are then transpiled to various target languages and platforms.
To use all libraries that have ever been written for X, someone only needs to write a transpiler-backend for the target language/platform once.
LLVM intermediate language is similar but too low-level. Think about a high-level counterpart.
This approach is used to some degree in producing, for example, optimized BLAS implementations and other similar routines:
http://view.eecs.berkeley.edu/wiki/Autotuners
Of course, you do then need to juggle multiple languages. Different languages can be better at expressing different things, though (e.g., "what the function does" versus "how the function does it"), so it's not inherently crazy.
You could develop one universal way to create and deploy functions to different languages, but now you've introduced one more language to get people to use.
And in practical terms, migrating this sort of thing from one language to another when necessary is generally a good enough solution.
According to this comment: https://news.ycombinator.com/item?id=2684488
Qualcomm now owns the patents.
From wikipedia: https://en.wikipedia.org/wiki/Software_patents_under_the_Eur...
> Under the EPC, and in particular its Article 52,[1] "programs for computers" are not regarded as inventions for the purpose of granting European patents,[2] but this exclusion from patentability only applies to the extent to which a European patent application or European patent relates to a computer program as such.[3] As a result of this partial exclusion, and despite the fact that the EPO subjects patent applications in this field to a much stricter scrutiny[4] when compared to their American counterpart, that does not mean that all inventions including some software are de jure not patentable.
Going back to the original question, I certainly wouldn't want to start an open source project knowing that Qualcomm might sue me for patent infringement. Defending their intellectual property is part of their core business strategy.
http://www.europeansoftwarepatents.com/howtos/software-paten...
https://en.wikipedia.org/wiki/Software_patents_under_the_Eur...
The standard: https://tools.ietf.org/html/rfc6330
http://www.computerweekly.com/feature/Erasure-coding-versus-...
Is there any empirical data which zfec parameters are recommended for which device? AFAIK reliability is DVD < Blueray < SSD < HD < tape < cloud.
Anyone knows more details?
Originals version of the license: http://zooko.com/tgppl.html
Seems like the goal was to keep the product under copyright for some time, giving it a opportunity to recover costs or integrated exclusively inside a product for a while, giving some advantage over competitors, but making it a public-license after the period.
Also found this explanation: https://tahoe-lafs.org/~zooko/tgppl.pdf
They were used for reliable data distribution, being able to recover from loss of packets without requesting the packets be sent again.
The only way to durably store data so that it survives a hardware failure (e.g. drive dying) is to store more than one copy. Full replicas are the simplest way to do this, but you've got a relatively high overhead (e.g. Store 1GB of data with 3x replicas, and you store 3GB of data). Erasure codes are a way to effectively store fractional replicas, so you only use 1.5x or 1.7x of the original data.
Erasure codes are great when you've got a lot of data and you need high durability but don't want to pay for the storage space required for full replicas.
Why don't we always use erasure codes for everything? EC isn't great when you've got small bits of data, and since there's a bit of math involved in reading and writing the EC data, EC has higher latency than simple replicas.
https://www.swiftstack.com/blog/2015/04/20/the-foundations-o... is a great into to how erasure codes work.
Basically, spread your data across many cloud storage providers and built a super fast soft-RAID.
There seems to be no way of using it to simply attach and have a list of your data only.
I'm favoring Camlistore over Tahoe, though Camli has no good encryption yet and no erasure encoding. The basis of Camli seems to be better.