The trade-off between memory usage and inference time uncovers a potential flaw in prioritizing resource efficiency over performance.
This would deter real-time or near real-time applications where latency is a critical factor.
Also, the confusion over the phrase "0.5-2x slower" highlights a possible lack of clarity in communication within the community, which would hinder the accurate assessment and adoption of such optimizations in practice.