I am very excited not only for the native parallelism that's coming in 5.0 but also about the effect handlers! I am sure many others are looking to start creating very interesting things with them, e.g. Rust-alike async capabilities or Erlang's preemptive green threads / actors runtime.
More info: https://discuss.ocaml.org/t/multicore-ocaml-september-2021-e...
Really looking forward to what the community will build with OCaml 5.00! IMO it will shoot the language straight into the mainstream and I can't wait. OCaml is applicable for like 90% what's out there, including as a Python replacement. And its syntax is often much terser than Rust (although the higher-level typing constructs can be confusing to read).
After that, all that's left is a tool like Elixir's mix or Rust's cargo and the language is basically not only in the 21st century but much farther than many others! Looking forward to it.
This so much. Mix is so pleasant to work with, but getting an OCaml project off the ground requires messing with dune + esy if you want consistent and isolated package management. It's a massive pain.
We only have a few places that use "naked" pointers, which are in any case deprecated, and we should be able to fix those easily enough.
Mostly the rest are fairly ordinary extensions that just call C functions and use the usual CAMLparam stuff.
We do have a few places that register global roots. And several packages that do callbacks from C back to OCaml. We also use @@noalloc a lot.
Is there anything else? Since there's so much code, what should I be grepping for to find code that might be of concern?
Edit: Some examples of simple stuff:
https://github.com/libguestfs/libguestfs-common/blob/master/... https://github.com/libguestfs/libguestfs-common/blob/master/... https://github.com/libguestfs/libguestfs-common/blob/master/...
More complex stuff:
https://gitlab.com/nbdkit/nbdkit/-/tree/master/plugins/ocaml http://oirase.annexia.org/tmp/ocaml/
Noalloc and registering roots should all be the same, as are callbacks (for sequential code, which all existing code will be)
We have a scheduled build and test of every package in opam with multicore: http://check.ocamllabs.io:8082/ to try to shake out C API incompatibilities and that's proved fruitful.
If you do find things that don't work on 5.0, please let us know (and if you can, get it in to opam so we test it automatically!).
i.e. previously some code could've in theory assumed it only ever gets called by one thread at a time due to the global lock (and that lock would get dropped in a controlled manner by the usual released/acquire). I don't necessarily mean just global state in the C stubs (I don't recall seeing many of those), but implicit global/shared state in the C functions called. Reading manpages will usually tell whether a given function is safe to be called (or when not there are sometimes _r variants that are). This is no different than when writing regular multithreaded C code, one has to be more careful about what functions they can safely call. (and in fact current OCaml code can already be multithreaded, so if the C stub was meant to work correctly with multithreaded OCaml code with C->OCaml callbacks then it should've already taken necessary precautions).
I like the approach taken wrt to backward compatibility of sequential code though: only code that wants to take advantage of multicore will need to audit its C stubs for safety, if you don't use multicore everything will work as before.
How do you recommend to find these kinds of multicore-specific bugs in existing (or newly written) OCaml-C stubs? Would tools like parafuzz help here?
[1]: https://github.com/ocaml-multicore/domainslib/blob/master/li...
There is work on-going at the moment to bridge or even unify eio (concurrency via effects) and domainslib (nested parallelism via domains and effects) but it's a few months out.
- domainslib schedules all tasks across all cores (like Go).
- eio keeps tasks on the same core (and you can use a shared job queue to distribute work between cores if you want).
Eio can certainly do async IO on multiple cores.
Moving tasks freely between cores has some major downsides - for example every time a Go program wants to access a shared value, it needs to take a mutex (and be careful to avoid deadlocks). Such races can be very hard to debug.
I suspect that the extra reliability is often worth the cost of sometimes having unbalanced workloads between cores. We're still investigating how big this effect is. When I worked at Docker, we spent a lot of time dealing with races in the Go code, even though almost nothing in Docker is CPU intensive!
For a group of tasks on a single core, you can be sure that e.g. incrementing a counter or scanning an array is an atomic operation. Only blocking operations (such as reading from a file or socket) can allow something else to run. And eio schedules tasks deterministically, so if you test with (deterministic) mocks then the trace of your tests is deterministic too. Eio's own unit-tests are mostly expect-based tests where the test just checks that the trace output matches the recorded trace, for example.
The Eio README has more information, plus a getting-started guide: https://github.com/ocaml-multicore/eio/blob/main/README.md
Now let's imagine I want to do the same program in OCaml. I think my options are:
- on current OCaml, thread-based concurrency but no parallelism
- on current OCaml, monadic concurrency (Lwt, Async) but no parallelism
- on multicore OCaml, direct/effect-based (I'm not sure what's the right word) concurrency with eio, which is deterministic. If I want parallelism here, I have to explicitely create and use a shared job queue, while the Go runtime does this implicitly. Since the standard library Queue is not thread safe, I would have to use Mutex to avoid concurrent access.
Is this correct? I've read the eio documentation but it's hard to wrap my head around all of that without examples. I've found the Domain_manager which looks like what I want. For example, I could have the main thread fill the queue and for each core available, I could launch a Domain_manager.run toto, with toto taking jobs from the queue that would be shared between all domains?
The README shows an example of a pool of workers pulling jobs from an Eio.Stream:
https://github.com/ocaml-multicore/eio#example-a-worker-pool
We're still exploring what APIs to provide for this kind of thing, and in particular how to unify domainslib and eio.
See the example near the end of https://raku-advent.blog/2021/12/01/batteries-included-gener... for what I mean.
There were three big reasons pushing me towards other languages (C++, Rust) for a certain class of throughput-focused workflows:
1. Lack of multicore execution
2. Lack of control over how many bytes a datatype is (useful for many things, like ensuring something fits in a machine word or can avoid chasing a pointer in a loop)
3. Difficulty controlling where memory is allocated/copied vs moved (though some of this is less relevant when everything uses reference semantics and mutability is tracked with ref/mutable)
I’ll be glad to see (1) fading into history! Do you have any personal tips/anecdotes for memory optimizations in practice, or any suggestions for something to read/follow/search?
2. Domainslib is the library developed alongside multicore to aid users in exploiting parallelism. It supports nested parallelism and is pretty highly optimised (https://github.com/ocaml-multicore/domainslib/pull/29 for some graphs/numbers). The domainslib repo has some good examples: https://github.com/ocaml-multicore/domainslib/tree/master/te...
3. We've not tested against other forms of parallelism. There isn't anything stopping you exploiting SIMD in addition to parallelism from domains.
4. No, we've not compared performance by OS.
5. No plans for the multicore team to look at accelerator integration at the moment.