Clover: 4-Bit Quantized Linear Algebra
astojanov.github.io
astojanov.github.io
Once the data exceed L3 cache,
4-bit is fastest since the dot
product is memory bound: data
movement from RAM becomes the bottleneck.
The speed-up is up to 6x over
the 32-bit version.
Large linalg operations are memory bound.I learned this the hard way hand coding an arm32 5x5 gaussian blur. When benchmarking my first version, using two passes of the separatable 5x1 and 1x5 filters, I had a facepalm moment when I realized I was really only measuring two round trips to main memory.
Operating on 16x8 pixel blocks, which fit within the register file, required most pixels to only be fetched once and almost doubled performance.
I've compared using GCC intrinsics vs what I would have wanted. It doesn't really get instruction scheduling right. Things like efficiently using pipelines by grouping, say, multiplication instructions and not stalling by ensuring the right number of cycles between them.
I have some specialised block matrix multiplication which basically never stall the processor. We're talking 10x speedups over optimized but general c++ code.
Currently mixing CMOS and DRAM prosesses into same silicon is not cost effective for mass production. I think mixing of the processes and heating issues are the main roadblocks.
When benchmarking, I usually have a larger workload in mind. I'm usually working on and deciding between two significant algorithmic changes and I'm looking at differences in 100-1000ms and I still find the standard deviations unsettling.
Addressing the author(s): I don’t know much about the topic area but you made it interesting, and it is a trove of cool tricks and techniques (sure, bit manipulation is basic but it’s fun to see how you put it all together to solve the problems). And how you analyze the performance. Look forward to coming back to this after I get more up to speed on linear algebra.
Fundamentally, the reason they're especially accurate is that most of the usual sources of error in digital–analog conversion just don't exist in a 1-bit ADC. INL? Zero. DNL? Zero. Nonmonotonicity? Please.
Clover, by contrast, is interesting to me for a different reason: a lot of ANN inference stuff seems to work fine at 8 bits of precision. It'll be interesting to see if those results can extend to 4 bits.
Now I need to study up on the area of computation that clover lives in.
I dare say that the most common AD converters are simply delta-sigma converters, especially if 1MHz conversions is fast enough for your tasks.
To get into the 100MHz or above speed (ie: Oscilloscopes, Software defined Radios, etc. etc.) you need to use other techniques.
That gets back to my original point. The spectrum analyzer went into the gigahertz range. I don't think it was delta sigma. It was some multiple sample thing that only worked in the frequency domain.
I distinctly remember that Delta-Sigma was a slow-type of ADC. And I guess I confused it with the one in the ATMega328p.
The ATMega's ADC is pretty slow, like 15ksps at full precision, so I can see why you'd think that!
They require higher clock frequencies and aren't that accurate (though, as you mentioned, some other sources of inaccuracy are simply nonexistent)