ParentFull threadairocker·yes, not sure you can do better than that. You cannot still have one instance of LLM in (GPU) memory answer two queries at one time.View on HN