Fully Sharded Data Parallel: Faster AI Training with Fewer GPUsengineering.fb.com3 points·TheGuyWhoCodes··2 commentsOpen articleSaveView on HN