Is this tied to a specific framework like pytorch or an inference server like vLLM?
Our inference stack is built using candle in Rust, how hard would it be to integrate?
Our inference stack is built using candle in Rust, how hard would it be to integrate?