ParentFull threadunderlines·vLLM is simple to set up, use docker and make sure your backend (ubuntu or WSL ubuntu or whatever) has GPU support installed.View on HN