Great project! How does the performance compare with conventional CPU/GPU based inference? Those devices are usually a lot higher power (and bigger/more expensive), but obviously do not benefit from specialization.
The only time I had to reach for quantized (integer) networks to do anything at all was inferencing on FPGAs. Are you targeting dsp slices by default or implementing full ieee754 floating point by default?
Are you saying that with Tensil you can run single precision non-quantized models with up to 2x gpu perf?
I probably misunderstood your last sentence, sorry.
Genuinely curious!