Did you actually test whether this happens? If I understand correctly, you want to rely on the optimizer constant-folding multiplications with zero to zero. But strict IEEE floating point conformance notoriously does not allow collapsing multiplication with zero to zero (consider NaNs and infinities), or addition of zero to be a no-op (consider signed zeros).
Example:
That also affects the overall goal, seeing as the type-level optimizations do assume the above properties. This means that you will get different results in certain edge cases than a "naive" implementation that always stores 16 floating point values for each multivector (and that's before getting to things like non-associativity of floating point operations).