It's not so much that people want to run the models in the browser. They want to be able to write and publish one desktop app that e.g. runs a LLaMA quality model and runs decently across the hundreds of different GPUs that exist on their users' machines.