Assuming this works without major regressions, this is a fantastic step forward for OCaml, opening it up to use cases that require parallelism.
In terms of performance it varies. The benchmarks in that paper it ranges from a 20% slowdown to a 20% speedup - most macro benchmarks are <10%. There's also still some more performance work we can do to chip away at that further.