And measure. Always measure. Without data it’s all posturing
The most frequent problem is selecting the fastest solution with a microbenchmark and creating a lot of unnecessary cache eviction/pollution, sometimes on the code cache, sometimes on L1/L2 or both...
Microbenchmarks are good to explore the optimization space, but the final decision should be made with the full execution context in mind.