VS
0.000095214 minutes per third-gen TPU (0.39 / 4096)
Is it only me or Third-gen seem better lol
VS
0.000095214 minutes per third-gen TPU (0.39 / 4096)
Is it only me or Third-gen seem better lol
- 466 minutes of fourth-generation TPU time.
- 1,600 minutes of third-generation TPU time.
Using this logic, fourth generation TPUs are 3.4x better. But, comparing different cluster sizes is pointless. These things don't scale linearly.
This mean each TPU contributed to (1/256) assuming it scale linearly.
Even if it doesn’t the 256 TPU ran in parallel not sequentially right ?
W / (256 * speed_v4) = 1.82
W / (4096 * speed_v3) = 0.39
speed_v4 / speed_v3 = (4096 * 0.39) / (256 * 1.82) = 3.43
Note that this assumes that training speed is perfectly linear with the # of accelerators which is not true as you get to very large #s (like 4096!). So the true number should be smaller than the 3.43x above, and the reported 2.7x makes sense.