> Fine-tuning and inference up to 10x faster than offloading
What is "offloading" in this context?
What is "offloading" in this context?
It turns out, Petals is faster than offloading even though it communicates over the Internet (possible, with servers far away from you). That's because Petals only sends NN activations between servers (a small amount of data), while offloading copies hundreds of GB of NN weights to GPU VRAM to generate each new token.
Several recent works aim to democratize LLMs
by “offloading” model parameters to slower but
cheaper memory (RAM or SSD), then running
them on the accelerator layer by layer (Pudipeddi
et al., 2020; Ren et al., 2021). This method allows
running LLMs with a single low-end accelerator
by loading parameters from RAM justin-time for
each forward pass. Offloading can be efficient for
processing many tokens in parallel, but it has inher-
ently high latency: for example, generating one to-
ken with BLOOM-176B takes at least 5.5 seconds
for the fastest RAM offloading setup and 22 sec-
onds for the fastest SSD offloading. In addition,
many computers do not have enough RAM to of-
fload 175B parameters.