ParentFull threadalex_sf·For the 65B fine tune, did you add another A100 node? Or just drop batch size?Any chance you’re up to sharing the training parameters?View on HN