Thank you. I ran both through lldb to see why luajit was faster and was confused how luajit was faster despite doing quite a bit more work:
luajit loop:
0x12a0effe0: movl %eax, %ebp
0x12a0effe2: movl %ebp, %edi
0x12a0effe4: callq *%rbx
0x12a0effe6: movsd 0x8(%rsp), %xmm0 ; xmm0 = mem[0],zero
0x12a0effec: xorps %xmm7, %xmm7
0x12a0effef: cvtsi2sdl %eax, %xmm7
0x12a0efff3: ucomisd %xmm7, %xmm0
0x12a0efff7: ja 0x12a0effe0
c_hello loop:
0x100000ea0 <+80>: movl %ebx, %edi
0x100000ea2 <+82>: callq 0x100000ee6 ; symbol stub for: plusone
0x100000ea7 <+87>: movl %eax, %ebx
0x100000ea9 <+89>: cmpl %r14d, %ebx
0x100000eac <+92>: jl 0x100000ea0 ; <+80>
0x100000ee6 <+0>: jmpq *0x134(%rip)
I'm still a bit surprised that that extra jumpq is more expensive than all those extra instructions luajit is executing.
EDIT: For those curious, Lua handles all numbers as floating-point; there are no ints in Lua. That's what all the extra instructions are for. Interestingly it looks like LuaJIT is storing the loop variable x as int (in the eax register) and just converting it to floating-point for comparison against "count" each loop. Interesting bending of the rules there.