ParentFull threadwskwon·Not really. vLLM optimizes the throughput of your LLM, but does not reduce the minimum required amount of resource to run your model.View on HN