I thought float4 sacrificed a negligible cost in evaluation quality for a 8x reduction in RAM?
Id actually say its the most important metric for most open models now, since the price per performance of closed cloud models is so competitive with open cloud models, so edge inference that is competitive is a clear value add