ParentFull threadAlyx1337·Thanks! There are ways to shave off the latency: hosting locally, using quantized/smaller models, streaming data instead of doing the tasks sequentiallyView on HN