One of the annoying things about the x86-64 ABI is that floating-point numbers use vector registers, even if they're only scalars and not vectors. The only way you can tell the difference is if the result is the p versus the s in vmulsd/vmulpd.
So the optimized assembly here isn't using any vectorization. Which isn't surprising, since as far as I could tell, the author isn't actually optimizing the code using LLVM. Most of the optimizations happen in opt, not in llc, which only does codegen optimizations. Those sorts of optimizations are largely things like stack frame optimization, instruction scheduling, or some more powerful pattern matching in instruction selection (which, even in -O0 in LLVM, does a limited amount of common subexpression elimination and the like). You might have to add restrict to the pointers (in LLVM terms, noalias on the arguments) to get vectorization to kick in, but the number of values is small enough for some sort of loop-versioning to probably kick in anyways.
The benchmarking is also pretty unfair. The C and C++ code have the values get copied in from a volatile array before progressing each loop iteration, which the Python-via-LLVM doesn't have. That doesn't sound like much, but it's 30 million extra guaranteed memory accesses over the time of the program, which will come out to a few milliseconds. (And , gee, the Python code is only faster by a few milliseconds). A better approach is to measure the time by moving the function to an external file and calling it in a tight loop, wrapped by a high-precision clock.
Also, benchmarking a 5×5? In practice, it's the memory access patterns that kill code, so you'll want to use programs large enough to actually cause any expected cache movements at the L2/L3 level. 5×5 is small enough that you could actually stick it all in vector registers if there's a bit of vectorization.