RISC-V Bitmanip Extension [pdf]
raw.githubusercontent.com
raw.githubusercontent.com
But ARMv8 has the more prosaic Bitfield Operations which fit in 32b instructions and are single cycle on A75. I use these a lot.
What makes these formats hard to do fast with regular code is that they may have been designed to work through dedicated hardware; or they just really pack in the bits.
The original title on this was "RISC-V Bitmanip Extension proposal is an education" because, besides a list of proposed instructions and precise definitions, the new draft explains in some detail how they are useful for real programs, and in many cases shows how they can be implemented with minimal additional gate count.
It's an education because many of these operations will be unfamiliar to most readers, and may suggest to them new ways to solve problems. Some of the instructions are also implemented in recent x86 cores, sometimes in an AVXx extension, but Intel's assembly language mnemonic offers little hint at how powerful it is, and its reference documentation sheds little more light.
So, previous drafts looked like a huge bolus of largely doubtful instructions that looked to bloat the RISC-V spec but be little used. Now we can see that most need only an extra gate here and there on such existing subunits as barrel shifters and multipliers, yet open up whole new vistas of operations they could be used for. The more powerful ones turn a whole family of O(N) operations (N the word size or number of 1s or 0s) to O(1). In turn, they are building blocks for fundamental signal analysis and encoding algorithms that may then run up to N times faster.
The document is maintained at https://github.com/riscv/riscv-bitmanip/ .
Note that the flexibility aspect means that it's relatively easy to design a chip that uses RISC-V (of which the OP article is an extension for). Designing chips is still out of reach of most hackers, but it does offer hopes of a more accessible landscape of General-Purpose CPUs, DSPs, GPUs, MCUs, etc...
[1] https://en.wikipedia.org/wiki/David_Patterson_(computer_scie...
[2] https://californiaconsultants.org/wp-content/uploads/2018/04...
Note that I'm specifically not talking about microcode, nor emulation in software.
So I gather the answer is "no", and the instructions not directly implemented must be implemented in microcode instead of trapping them to a library.
Of course this is a lot slower than the instruction might have been, but at least the code runs and produces the right answer.
Sometimes the trap does not change the execution sequence of externally visible instructions, but instead directs the machine to execute a sequence of internal microcode, while externally it looks like the machine has stalled. Indeed, on some machines all instructions trap this way.
Modern CISC chips do a more sophisticated form of the latter, where RISCs have typically avoided it for most or all operations.
So, whether and how RISC-V implementations emulate extensions will vary. Emulation generally doesn't produce very nice results, for this kind of operation -- the purpose of the instruction is to give users access to custom machinery wired into the core, for speed, and without the machinery the programmer might have preferred to achieve the results some other way.
Good overview of the research:
https://users.ece.cmu.edu/~koopman/roses/dsn04/koopman04_crc...
Optimal CRC polynomials:
https://users.ece.cmu.edu/~koopman/crc/index.html
Note that the best CRC depends on the length of the message it's protecting.
There is no one optimal CRC choice, the best CRC will depend on the error model of your channel and the lengths of the data that you're processing.
Content-Disposition: attachment
That's the correct way to do it. What's actually happening is that the wrong content type is being returned: Content-Type: application/octet-stream
This is the lazy way to do it.How is "is an education" inappropriate?
In fairness to the doc, it's a rationale. It should educate.