That said, I don't actually know allocation is the problem; it's a suspect, but unlike a native script I can't just run `perf` to get performance counters and assembly listings. All I really know is that the non-allocating order is a lot faster.
[1]: If you can read assembly, this one's hilarious: https://godbolt.org/g/81qWoq