I'm under the impression that the implementation is based on algebraic effects. Does that include extending the language in a way that lets users define their own effects & handlers?
Also, how's the performance? Last time I looked this stuff up (and played around with Eff), which was quite a while ago, I was told that regular usage of effects may impact performance quite noticeably.
Multicore adds parallelism via Domains (which are essentially heavyweight threads) and concurrency via Effects and fibers. There's a multicore GC that supports both of those.
We plan to upstream things in two parts. First domains-only parallelism and then effects as a follow-up. When the latter lands users will be able to define their own effects and handlers, yes.
Performance is pretty good, you can see our PLDI2021 paper for a proper performance evaluation and loads more details: https://arxiv.org/abs/2104.00250
Too many programming languages, too little free time, sadly.
As understand that currently, the situation is analogous to Python. The GIL allows for process based concurrency, with the well known disadvantage regarding memory consumption. Also, my guess is that OCaml relies library based solutions at the moment?
1) What does the introduction of MultiCore potential mean for Ocaml? Will Ocaml be a much better fit to run a webserver backend? Could you perhaps give a comparison to other programming languages?
2) How will Ocaml stand out with Multicore in the PL world, and what solutions would be it uniquely suited for?
3) What tools are going to be present to deal with bugs introduced by your new runtime e.g. race conditions?
4) If one wants to have a go at Multicore and play around with it, where/how do I start?
Many thanks!
1) As I mentioned in https://news.ycombinator.com/item?id=27142502 there is support for parallelism and concurrency.
Giving an example of where these might be useful in a webservice.
The addition of shared-memory parallelism is beneficial where you might have a great deal of shared state that needs to be used to service requests. An in-memory cache is a good example - with a processed-based approach managing read/writes and avoiding significant overhead from marshalling the data is difficult.
Concurrency via effects at a minimum can make writing network-based services much more pleasant (and debuggable!). See the examples in https://arxiv.org/abs/2104.00250 where programs can be written in a direct-style similar to blocking IO but using effects are transformed to use asynchronous interfaces. There's work going on in the project at the moment to build fast cross-platform IO implementations that sit atop of uring/gcd/iocp.
3) This is a good question and one we're still working on. I think one of the lead developers KC has a few good ideas about instrumentation we can do to enable detecting races to global state. It's certainly going to be an issue for people porting large codebases.
4) This is the place to start: https://github.com/ocaml-multicore/multicore-opam#install-mu... .
What in Ocaml makes this so much harder to implement?
[1] https://www.researchgate.net/publication/2774662_Concurrent_...
Our paper last year covers most of why this is tricky: https://arxiv.org/abs/2004.11663
We are working on developing an effect system, which will ensure effect safety i.e, the compiler ensures that all the effects performed are caught. You also get a nice inferred type that says what effects a particular function may perform; if it performs none, then it is a pure function! This implementation would still use the current fiber support in the runtime. Leo White, one of the developers of Multicore OCaml had given a talk on this new effect system a few years ago [2]. That's the best place today to learn about the new effect handlers.
The plan is to first add the fiber runtime support to OCaml without the syntax extensions for effect handlers, and then introduce syntax along with the effect system.
[1] https://arxiv.org/abs/2104.00250
[2] https://www.janestreet.com/tech-talks/effective-programming/
Didn't see any mentions of critical sections (mutexes) with C++ examples in the documentation ("Bounding Data Races in Space and Time"). I'm not sure I understand the comparisons the writers are presenting.
However, in the multi-core setting, data races pose bounds and limits on how much you can trigger those optimizations. The program doesn't generally execute sequentially in a way that can be entirely reasoned about. Instructions might be reordered for the sake of the program to run faster.
Programmers can't work with that. So one proposes a memory model. Follow these rules, and our optimizations won't alter the behavior of the program. They kind-of describes what happens "in between" the critical sections of the program, hence the lack of a mutex mention.
The paper presents a local property and then shows, formally, that this property is enough to guarantee an efficient memory model. That is, a model in which you can perform optimizations, while programmers can still reason about the programs behavior.
The crux of the paper is that the property is local. This is new, because memory models which came before it are global: to reason about correctness, you have to consider the whole program, rather than consider a small (local) subset. OCaml requires more safety than most programming languages, so this is good for the fact that you can now compose local fragments of OCaml programs, without having to worry about a global safety property.
The property is also simpler for programmers to reason about.
The way you "use" the paper is that you adapt your optimizations to follow the property, and you make sure that the virtual memory model is implemented the same way on different architectures.
Finally, the examples: they explore the idea of a local reasoning. In particular, they show why the (existing) global properties fail if you view them under the stronger requirement of local reasoning. It's the setup for the paper, since it means you can't just use the existing models. They need to be adapted if you want a more localized property.
Firstly, if you are using high-level synchronisation mechanisms such as mutexes and condition variables, or higher-level concurreny libraries such as java.util.concurrent, you shouldn't worry about the memory model. C++, Java and OCaml ensure that properly synchronised programs do not exhibit surprising behaviours. Such programs have sequentially consistent semantics i.e, the observed behaviour is one of the permitted interleavings of the threads in the program. Rust inherits C++ memory model [2], but if you are using the safe subset, then you will never have to think about it. Memory model is important only to those who write the concurrency libraries. If you are still keen, read on.
The OCaml memory model is certainly simpler than the C++ and Java memory model, but being more efficient is not one of our goals. C++ memory model permits a partially-ordered lattice of stronger memory accesses starting from access to non-atomic memory locations to sequentially consistent access with increasing cost as you move up the lattice. OCaml memory model only provides two -- atomic and non-atomic, representing approximately the top and the bottom of the lattice.
OCaml memory model is also stronger than Java in that our data races are bound not-only in space like Java (data races on certain variables don't affect behaviours on other variables) but also in time (surprising behaviours stop affecting the program after the race ends unlike Java; see example 2 from [1]). This permits modular reasoning of racy and non-racy parts of the program which is not the case with C++ and Java.
The catch is that we have to disallow load-to-store reordering to get the stronger guarantees. Relaxed memory models such as ARM and Power do in fact permit these reorderings, and we have to compile OCaml code (including sequential one) such that the load-to-store reordering is disallowed. This can be done fairly cheaply. It is free on x86 which doesn't perform load-to-store reorderings, and has a small cost (up to 3%) on ARM and Power architectures whose memory models permit load-to-store reorderings.
Small correction, atomics are part of the safe subset of Rust. At a certain point it's important to have a work-a-day knowledge of the memory model with regard to atomic numerics. Dealing with allocated types, now, that's a whole different area and is specialized knowledge.
For example if I implement an algorithm that ought to be Acquire / Release (e.g. to build my own custom mutual exclusion) but I tell Rust it's OK to have Relaxed semantics, it sounds like on this PC (an x86-64) it will work anyway, but on some other systems it won't.
And I'd have achieved this goof without writing unsafe Rust, just the same way as if I screwed up a directory traversing algorithm because I relied on semantics not present in all file systems?
[0] https://preshing.com/20120930/weak-vs-strong-memory-models/ [1] https://github.com/tokio-rs/loom
Fwiw, OCaml itself can be compiled natively on M1.
The overview Anil gave at the end of last year on the OCaml Platform should give you an idea of the many other strands of work that are going on: https://ocaml.org/platform/