The concurrency story is similar to what you get in Python. On one hand, you have event loop implementations with various conveniences for scheduling callbacks (Async, LWT); on the other, you have OS-level threads, but the runtime system has a global lock which makes only one thread execute at a given time. So no parallelism without forking[1] to separate processes. Moreover, the interactions between an event loop and OS-level threads which sometimes arise are non-trivial and in my experience very hard to get right, even with the helpers in Async.
Single-threaded performance is generally good when natively compiled. In a benchmark I did a few years back, which consisted of traversing the filesystem with `opendir`, `readdir` and friends, OCaml was minimally slower than C++ and Nim (when compiled natively) and around Python and Racket when byte-compiled. It probably would be even faster if I used its imperative features. Looking at the code today I also see some unnecessary copying of lists, which could be mitigated. All in all, I think OCaml has a really good performance considering its high-level semantics. You have to be aware of relative costs of operations, but if you do, "as fast as C" is certainly attainable. OCaml code for the benchmark: https://klibert.pl/posts/walkfiles-ocaml.html and the results: https://klibert.pl/statics/images/walkfiled_perf_test1.png
[1] Actually, I just checked and it looks that `fork` is not implemented, so copy-on-write memory sharing is impossible, even on POSIX systems. Docs: http://caml.inria.fr/pub/docs/manual-ocaml/libunix.html http://caml.inria.fr/pub/docs/manual-ocaml/libref/Unix.html