Low level string operations happen so often under the hood in any application that 10x more code and complexity in an implementation is worth it if it increases efficiency.
When I was at Google there was a "rules of thumb" page that described how much of your time X performance was worth. I always looked at that before micro-optimizing, but also always came out ahead. I remember a coworker and I redesigning one of our subsystems late on some Friday evening; the rules of thumb said that the performance improvement we predicted, at our scale, was worth a month of SWE time. We did it in 2 hours. So we came out ahead, and our system worked better. (I joked that my colleague and I would be taking 2 extra weeks of vacation.)
TL;DR performance matters everywhere. The one user using an interactive tool on their workstation will appreciate their day not being wasted by random pauses. The person out and about on their phone will appreciate the additional battery life. And your company's bean counters will be quite happy to hear that your cloud bill or data center expense forecast for next year is down. Finally, it's fun! Truly a win/win. Make it fast!
This was a very easy change; we just made a new main.go and an RPC to send the data to aggregate. The system was designed internally to be logically isolated across that boundary, so we just stuck in the RPC and then the other side of the boundary could be another data center.
In the end, I think we saved a few terabytes of RAM-years. Not a big deal, but it was something.
You're more-or-less programming in assembly here using these intrinsics, using C for goodies like for loops, and at the moment that's about as good as you can do while scalable autovectorization is still WIP/NIH in most compilers, so it's not really surprising or noteworthy. For experienced SIMD programmers, this is the standard approach to getting data parallelism out of many "boring" algorithms.