Basic math related to computation and memory usage for transformers
blog.eleuther.ai
blog.eleuther.ai
I don't look at any model the same after head to head comparisons with full precision and quantization at 4bit have run on my machine. There is little to no perceptible change with models of the same initial weight. BUT!!! I am now able to run models that required a DGX a few weeks ago on my home computer thanks to quantization. These models are better in every way from my POV. I am now more interested in what I can "do" with the models vs. just getting them to run. 30B at 4 bits is the sweet spot for my setup.
$$ \begin{align}\text{Total Memory}{\text{Training}} = \text{memory}{\text{model}}+\text{memory}{\text{optimizer}}+\text{memory}{\text{activations}}+\text{memory}_{\text{gradients}}\end{align} $$
> GPT-NeoX achieves 150 TFLOP/s/A100 with normal attention and 180 FLOP/s/A100 with Flash Attention.
This advice implies they are using flash attention.