OK that makes sense; so basically if I want to give this a shot then I would just need to read llm_chat.js and the TVM docs, and translate llm_chat.js to my language of choice effectively?
Then yah the llm_chat.js would be high-level logic that targets the tvm runtime, and can be implemented in any language that tvm runtime support(that includes, js, java, c++ rust etc).
Support webgpu native is an interesting direction. Feel free to open a thread in tvm discuss forum and perhaps there would be fun things to collaborate in OSS