YaFSDP: a sharded data parallelism framework, faster for pre-training LLMs
github.com
github.com
ꙗ / Ѧ
Or do you mean you literally can't draw them?
I see you have provided it, making it more accessible for my future use, at least on the timeframe of this thread being in my recent HN activity.
In Yandex’s pre-trainings, the implementation of YaFSDP along with other memory optimization strategies resulted in a speed gain of 45%.