How does x.cos().cos() work faster than doing two cos calls separately? Like the first cos call returns a tensor either way, the only difference is that it's not assigned to a variable. But how is it even possible know that difference in python?
Non-fused:
foreach i
y[i] = cos(x[i])
foreach i
z[i] = cos(y[i])
Fused, no intermediate variable: foreach i
t = cos(x[i])
z[i] = cos(t)
The temporary "t" doesn't leave the GPU. Sweeping the array twice makes you twice as dependent on memory bandwidth.[1] https://colab.research.google.com/drive/13a4Y-ko6QLMPAhBz64c...