"cksum [-a crc] is now up to 4 times faster by using a slice by 8 algorithm, and at least 8 times faster where pclmul instructions are supported."
Implementing that portably is a bit tricky, as one must consider:
- support various compilers which may not support intrinsics
- runtime checks to see if the current CPU supports the instructions
- ensure compiler options enabling the instructions are restricted to their own lib to ensure the don't leak into unprotected code.
- automake requires using a separate lib for this rather than just a separate compilation unit
BTW we also introduced avx intrinsics for `wc -l`For embedded, goodluck finding a 12-bit CRC algorithm, and even if you find it, it's for a non optimal, non Koopman CRC [1].
While rolling your own is doable, it's also a risky endeavor, as any error hurts the product like forever.
With this generator, the problem that non Koopman CRC's were, is finally solved.
Indeed. It's all about generating lookup tables from a CRC's polynomial definition.
The CRC is a polynomial / Galois field, and today's CPUs have polynomial multiply (aka: pmul on ARM), or carryless multiply (aka: PCLMULQDQ on x86). These instructions can implement the "tough" part of the CRC in just one clock tick.