Memory alignment is important and languages let you control that and/or make them align automatically given sensible assumptions does have an edge here.
Just a random example,
I wrote a two lines Python function using Numba jit,
because I know in my mind that some Python behavior is going to make it less performant
and it is a high performance kernel so we need that fast (and also lower memory footprint.)
Compiling using Numba jit is no brainer because I just add one more line there
with like 10 seconds effort.
But the code base I am merging to has a policy against Numba
(reasonably as we are targeting HPC platform
where Numba has some performance problem related to oversubscribing
if not setting up carefully.)
So I end up rewrote that in C++ and wrap it with pybind11.
And the result is that it is faster by around 30%.
Since the algorithm is entirely trivial,
the only explanation I have is exactly memory alignment,
where I can control that in C++,
but in Numba jit there's no way
(both to guarantee the array allocated is aligned,
or tell the compiler to assume that.)
(The 30% number is also reasonable
in textbook examples.)
P.S. But it has a cost:
> Showing 14 changed files with 230 additions and 36 deletions.
> from https://github.com/hpc4cmb/toast/commit/a38d1d6dbcc97001a1ad...
where the 36 deletions are mostly just documentations.
So it's an ~200 lines effort comparing to a 4 line effort...