Because Wren's a bytecode interpreted dynamic language much of the overall time is probably spent just executing the surrounding benchmark code. I.e. in:
FFI.start()
while (x < count) {
x = FFI.plusone(x)
}
FFI.stop()
It probably spends most of the total elapsed time on: FFI.start()
while (x < count) {
x = ...
}
FFI.stop()
It would be worth running a similar benchmark but with just the FFI part removed and then subtract that from the time. (And do the same thing for the other languages too, of course.)