Last time this was up I wrote a single-threaded version in C which I'm pretty sure beats both Julia and Mojo: https://github.com/bjourne/c-examples/blob/master/programs/m...
Sometimes "showing the code" is not enough. Show me the benchmark.
The only other explanation is if you ran the non-simd Julia version under a single thread.
Without it, Julia starts single threaded, which means the code does all this work to enable multithreading and then doesn't get to benefit from it.
Would the C compiler automatically exploit vectorized instructions on the CPU, or loop/kernel fusion, etc? It’s unclear otherwise how it would be faster than Julia/Mojo code exploiting several hardware features.
IIUC, SIMD.jl only works because it only provides what is guaranteed by LLVM to work cross-platform, which is quite far from being able to use AVX2, for example.