> CPUs have gotten much faster, but moving bits around in memory has not.
Would it then not suffice to just load 32 bit floats, convert them to 64 bits on the processor, and then store them again as 32 bits? Essentially, we are 'compressing' our data in memory, whilst doing the computational steps with full precision.