a) train two identical models on a large dataset, one with the +1 in the denominator for the softmax steps of the attention modules, one without
b) show that they have similar performance (doubt the +1 will make performance better, but we need to show it doesn't make things worse)
c) show that there are less "blowups" in the model with +1, and therefore they are more effectively quantized.