I've ran across this before, but haven't figured out if a GPU is required to use it. That would be a deal-breaker for any consumer facing web app right?
It will work with software implementations (e.g. llvmpipe/lavapipe is used in CI). That will be slower though. However you’d be surprised at how many devices have a usable GPU. Not all have NVIDIA RTX’es of course, but a simple iGPU will do. I tested it on a fairly low specced Dell office PC with an Intel Iris iGPU and that works just fine. Also most mobile devices will have a GPU these days.
Thanks! it looks like the wonnx CLI itself falls back to tract to do inference on CPU if a GPU is not available[0], and it also sounds like setting up llvmpipe/lavapipe on WASM is much harder (if not impossible?) than just shipping tract, so the approach I'll take is probably a wonnx+tract approach.
I'm curious if there are many scenarios where a CPU fallback, especially javascript based, would have acceptable performance when a GPU is not present. Is that even really a solution?
I'm not far enough into my project to know for sure, but I think that, for example, using tract to do speech recognition can be done in sub-realtime (ie, the inference takes shorter than the audio duration) on a Macbook M1 Pro. Hopefully on less powerful devices too, but I haven't tested that yet.