ParentFull threadbm-rf·You could use Microsoft's DeepSpeed to run the model for inference on multiple GPUS, see https://www.deepspeed.ai/tutorials/inference-tutorial/View on HN