It's used to manage code fragments that need to be accessed in signal handlers.
5,054 karma · joined July 19, 2007
It's used to manage code fragments that need to be accessed in signal handlers.
Whether it's beneficial or not is unclear though.
There's still a lot of work to be done before 5.0 is released, it's just it'll happen through the usual PR and review process rather than a separate one.
Work on this is on-going via the sandmark benchmarking suite: https://github.com/ocaml-bench/sandmark
In short the expectation should be that single-threaded code performs roughly the same (single digit percentage changes) as on the sequential runtime.
Parallel code on multicore can see close to linear speedups on 64 cores, though it depends significantly on your workload. If you're interested in parallelising existing OCaml code, I gave an example-driven OCaml workshop talk in 2020: https://www.youtube.com/watch?v=Z7YZR1q8wzI
We've written a couple of papers detailing the internals and the trade-offs involved: https://arxiv.org/abs/2004.11663 (for parallelism) and https://arxiv.org/abs/2104.00250 (for effects)
Added to that is the complexity of tracking a moving target. Multicore had to be rebased through 12 releases of OCaml, which in itself was a non-trivial amount of work.
https://github.com/ocaml-multicore/effects-examples has links to tutorials and examples for how effects can be used.
There's also some slides from KC's talk on effect handlers https://kcsrk.info/slides/handlers_edinburgh.pdf and materials from the CUFP 17 tutorial: https://github.com/ocamllabs/ocaml-effects-tutorial
https://gopiandcode.uk/logs/log-bye-bye-monads-algebraic-eff... this is also a great introduction
(for maybe the second to last time)
We have a scheduled build and test of every package in opam with multicore: http://check.ocamllabs.io:8082/ to try to shake out C API incompatibilities and that's proved fruitful.
If you do find things that don't work on 5.0, please let us know (and if you can, get it in to opam so we test it automatically!).
2. Domainslib is the library developed alongside multicore to aid users in exploiting parallelism. It supports nested parallelism and is pretty highly optimised (https://github.com/ocaml-multicore/domainslib/pull/29 for some graphs/numbers). The domainslib repo has some good examples: https://github.com/ocaml-multicore/domainslib/tree/master/te...
3. We've not tested against other forms of parallelism. There isn't anything stopping you exploiting SIMD in addition to parallelism from domains.
4. No, we've not compared performance by OS.
5. No plans for the multicore team to look at accelerator integration at the moment.
There is work on-going at the moment to bridge or even unify eio (concurrency via effects) and domainslib (nested parallelism via domains and effects) but it's a few months out.
More info: https://discuss.ocaml.org/t/multicore-ocaml-september-2021-e...
2. That is unclear at the moment. There's a lot of useful history in those commits (which link out to issues and PRs on ocaml-multicore's repo) but at the same time, it also includes a lot of experiments that were ultimately backed out.
Also you can start playing with effects today using the 4.12+domains branch on https://github.com/ocaml-multicore/ocaml-multicore
The original plan was to upstream only the multicore GC. This was sped up on the suggestion of the core developers and now 5.0 will bring parallelism and effect handlers (though without syntactic support for the latter).
https://discuss.ocaml.org/t/multicore-ocaml-september-2021-e... has a good explanation of effect handlers, syntax and what will be available in 5.0.
For what it's worth you can actually start using Multicore OCaml today, there are installation instructions on the wiki: https://github.com/ocaml-multicore/ocaml-multicore
In short there should be very little performance difference against stock.
As KC has linked in a sibling comment there's a project going to explore using effects to write high performance IO libraries and that has some early benchmarks.
The overview Anil gave at the end of last year on the OCaml Platform should give you an idea of the many other strands of work that are going on: https://ocaml.org/platform/
1) As I mentioned in https://news.ycombinator.com/item?id=27142502 there is support for parallelism and concurrency.
Giving an example of where these might be useful in a webservice.
The addition of shared-memory parallelism is beneficial where you might have a great deal of shared state that needs to be used to service requests. An in-memory cache is a good example - with a processed-based approach managing read/writes and avoiding significant overhead from marshalling the data is difficult.
Concurrency via effects at a minimum can make writing network-based services much more pleasant (and debuggable!). See the examples in https://arxiv.org/abs/2104.00250 where programs can be written in a direct-style similar to blocking IO but using effects are transformed to use asynchronous interfaces. There's work going on in the project at the moment to build fast cross-platform IO implementations that sit atop of uring/gcd/iocp.
3) This is a good question and one we're still working on. I think one of the lead developers KC has a few good ideas about instrumentation we can do to enable detecting races to global state. It's certainly going to be an issue for people porting large codebases.
4) This is the place to start: https://github.com/ocaml-multicore/multicore-opam#install-mu... .
Multicore adds parallelism via Domains (which are essentially heavyweight threads) and concurrency via Effects and fibers. There's a multicore GC that supports both of those.
We plan to upstream things in two parts. First domains-only parallelism and then effects as a follow-up. When the latter lands users will be able to define their own effects and handlers, yes.
Performance is pretty good, you can see our PLDI2021 paper for a proper performance evaluation and loads more details: https://arxiv.org/abs/2104.00250
Our paper last year covers most of why this is tricky: https://arxiv.org/abs/2004.11663
That's >2x higher throughput than an Optane at equivalent queue depth and (as far as I can see in the UK) at less than a tenth of the price: https://www.scan.co.uk/products/2tb-wd-black-sn850-m2-2280-p... vs https://www.scan.co.uk/products/15tb-intel-optane-dc-p4800x-...
And that's 7 GB/s from _one_ SSD. Aggregate memory bandwidth on something like the Zen3 is roughly 40 GB/s. These are also first generation PCIe 4, plenty more to come.
Doesn't require a huge improvement before you end up in a position where you simply don't have the memory bandwidth or cycles to deal with more than one drive.
I suspect the Optane wins most of the benchmarks because of it's outrageously good low queue depth random read performance - that's very effective for software that's not written for modern NVMe SSDs which benefit from very high queue depths. Check out the 4k random read performance from the SN850 at high queue depths:
https://images.anandtech.com/doci/16505/rr-s-sn850-1000.png
If you can keep the queues deep, it manages to beat the throughput of the Optane. You've got to design algorithms and data structures to exploit that kind of concurrency though.