So what's the efficiency of this model? Can I use it instead of pocketsphinx on a raspberry pi?
For reference, our system runs at 0.1 RTF on iPhone 10 using Accelerate framework under FP16 precision. INT8 should be better but haven't benchmarked yet!