Draft of OCaml Scientific Computing book
discuss.ocaml.org
discuss.ocaml.org
I don't see Rust being the best option for anything other than low level systems code.
My point was more that I think OCaml competes mostly with other functional and more generally GCed languages, not a language like Rust.
Multicore benchmarks are on github and are used to drive the changes. When used properly you can see large speedups without affecting much the speed of single core OCaml (which is quite fast).
If you want more details there are monthly updates now on the discuss.ocaml.org forum (just search for multicore). We usually post them also on HN when there is some meat.
Even though rust (or Julia if we talk about numerics) can clearly be many times faster, it is an imperative language with a functional feeling (don't take it in the wrong way, I think both are great languages for different reasons). For less performance critical applications I think OCaml is still worth, it is reasonably fast for a GCed language, has a great type system and the compiler is egregiously fast.
Multicore makes quite a difference in performances for certain numerical code, however it will be hard (if even possible at all) to reach the speed of fine-tuned Julia, C/C++, Fortran or Rust (which does not yet have a proper numerical framework though, afaik). The approach of owl has been to introduce an engine (wol-symbolic) that compiles computation graphs to ONNX format, so that they can be executed on GPU, FPGA, different engines, releasing from it the burden of supporting those various backends.
Personally I find the ease of maintenance and refactoring with reasonable speed a good compromise, I think its sweet spot is for exploratory code that does not require bare metal speed and yet should be fast enough.
Essentially you can already have multi-threaded OCaml programs as long as only one thread is using the OCaml runtime at any point in time. For numerical code where you might be spending the vast majority of your time in external libraries this ends up not being a major problem. It's not a dissimilar story for where Python is.
What Multicore OCaml adds is the ability to run multiple threads of OCaml code at the same time (we call them Domains, to avoid confusing them with existing Threads - which can coexist).
There's an entry on the Multicore wiki that gives some more depth: https://github.com/ocaml-multicore/ocaml-multicore/wiki/Conc...
In terms of the project you can also follow progress in the Multicore Monthlies: https://discuss.ocaml.org/tag/multicore-monthly as well as see the in-progress and merged multicore PRs that are hitting upstream ocaml: https://github.com/ocaml/ocaml/pulls?q=is%3Apr+label%3Amulti...
If you want to know more about how the multicore runtime works the recent ICFP2020 paper has a lot of detail: https://arxiv.org/abs/2004.11663 and KC's presentation is worth a watch: https://www.youtube.com/watch?v=ASX79I0jm6M&feature=youtu.be...
I'll just note here that this paper is hot off the presses! ICFP is taking place as we speak, this week, Monday to Friday.
The biggest problem with OCaml is the tooling (and Windows support is pretty rough). It is getting better, but it's certainly not as easy as using something like Cargo.
Probably for the exact same reason each year you think about learning OCaml you decide not to (I mean this sincerely, not trying to be snide).
I think the main reasons it doesn't see mass adoption in industry:
* There are only two major companies that do a substantial amount of OCaml that I can think of off the top of my head, Jane Street and Ahrefs. Facebook does some OCaml too but I don't think it's a core part of their stack.
* The tooling is lacking.
* People have an easier time learning Python or Java so you'll have a larger pool of candidates if you use one of those languages.
Using an ML or a Lisp for language tooling is the way to go. Nothing else in that league.
Seriously, Haskell is massively impressive, both as a research language and as an implementation. But it does not shed that certain research attitude. Every known problem seems to be boring. "Oh you want a proxying http server? No problem, this is just the inversion of the endofoo over the category of abstract Monobars!". Sometimes I get the feeling that no one focuses on shipping actual software with Haskell.
https://engineering.fb.com/security/fighting-spam-with-haske...
With just a modicum of effort you can easily disprove this feeling.
Care to elaborate on this point? Maybe it's just because I'm coming from Haskell (lol), but my experience with OCaml's tooling has been pretty darn good.
While I don't think there's a heavyweight IDE for OCaml à la IntelliJ, in Emacs I get my error messages inline, on-the-fly checking (including type inference and checking, which is huge), and pretty good completion. All of this seems to "just work".
The build system also seems to "just work" and utop is pretty nice.
It is also true that the standard library is growing faster recently, so maybe this will become less of a problem I. The future.
Have you tried using OCaml's monadic let expressions? It was introduced on verion 4.08 and they share similarity with F# computation expressions.
* https://jobjo.github.io/2019/04/24/ocaml-has-some-new-shiny-...
* https://caml.inria.fr/pub/docs/manual-ocaml/bindingops.html
* More production ready (including insanely good multithreading - probably the best of any language except perhaps Erlang)
* Better for learning - there are more advanced topics you can go into with Haskell, especially around the type system. There is nothing I can think of that you can learn in ocaml that you can’t learn in Haskell
* More fun - if you want to, there are lots of entertaining/aesthetically pleasing things you can do with writing concise/elegant programs or proving properties using the type system
Also, a lot of the putative advantages of OCaml over Haskell (e.g. “Monads seem complicated”) disappear when you use OCaml in practice.
This isn’t to shit on OCaml - it’s better than 95% of languages out there. But if you’re starting from scratch with no prior investment, I would not prefer it in any scenario I can think of.
I would use Haskell, but I don't believe in tracking effects with a type system and also lazy by default causes a lot of problems. Of course I guess I could use unsafePerformIO everywhere, but that seems wrong.
I've said this before and I'll say it again, modular implicits would be a game changer for OCaml.
That would be the second time we disagree about that on HN but I still fail to see how modular implicits are supposed to be game changing for OCaml.
The situation was different before 2011. Now, first-class modules and local open have made modules really easy to use. Modular implicits would mostly be sugar most of the time.
All of that could be done with OCaml: typeclasses, parameterized behavior, configurable functionality, testing.
All of that stuff can be done explicitly right now, but I really do think there's something powerful about changing an include/import and having all your code change behavior.
For example, you could have two logging modules: one that logs to stdout the other that logs json to elasticsearch. Simply change an import, and the logging behavior of your entire app changes.
You gotta admit there's something just really awesome about that.
But you can already do that with parametrized modules. Just change the module you pass as a parameter and the behaviour cjanges. It is not more cumbersome than changing an include and it is much more clear (dare I say explicit) in its intent.
Maybe people should just use OCaml's object system. It solves a lot of these problems and is very powerful.
Not sure how it came to be that everyone totally ignores it.
Local open with let+ and let* make writing monadic and applicative code very easy when needed. Anything more abstract than that is over complex abstraction.
OCaml is not Haskell. Trying to write Haskell code in OCaml is always going to result in something highly unidiomatic.
> Maybe people should just use OCaml's object system. It solves a lot of these problems and is very powerful.
I have written a fair bit of OCaml and never actually encountered these problems. In practice, module prefixing and local open are just fine. You rarely use more than one module per function anyway.
I agree that the object system is nice but structural typing is rarely what you want and it leads to convoluted error messages.
I don't understand comments like this. This is obviously not something that Haskellers actually do, so the simplest conclusion (or, it seemed simplest to me starting Haskell) is that thinking this is necessary is simply indicative of not understanding how to program functionally/monadically. Indeed, once you figure out Haskell idioms this isn't ever an issue.
You might also say that OCaml suffers from using unsafePerformIO everywhere, all the time.
> modular implicits would be a game changer for OCaml.
They would be very nice!
In my experience, there's little value in the monadization of all effects and also a non-trivial cost. I much prefer the traditional FP style of avoiding mutation and side-effects but not lifting them into the type system when they are necessary.
This does mean you can screw up and accidentally mutate things, but I've personally never had the problem.
Having worked with both, I really like using OCaml but avoid Haskell like the plague. Being lazy by default makes it very difficult to reason about Haskell performance really fast.
Admittedly another benefit of using OCaml is that you don't have to deal with the Haskell community and its obsession with looking smart. From experience, OCaml users tend to be a lot more oriented towards practical matters like, you know, building softwares rather than lenses library.
> Being lazy by default makes it very difficult to reason about Haskell performance really fast.
I think this is a ~complete non-issue mostly brought up by people who haven't actually used Haskell much at all.
> practical matters like, you know, building softwares rather than lenses library.
This is the classic completely idiotic "haskell isn't practical" argument. Let me tell you what's not practical: not having multithreading. Not having a well-functioning asynchronous programming system (Async sucks).
Also, OCaml does have a clone of Lens, but it was only written in the last year. Another example of how OCaml is very far behind Haskell in terms of development/production readiness.
It's not persecution. It's just that recently you can't have a discussion about OCaml without having an Haskell devotee dropping by and implying you should use Haskell instead generally for somewhat spurious reasons (same with Rust but the languages are so different, it's easier to ignore). It gets a bit tiring.
> think this is a ~complete non-issue mostly brought up by people who haven't actually used Haskell much at all.
We might have to agree to disagree on this one.
> Not having a well-functioning asynchronous programming system (Async sucks).
It's a chance everyone uses Lwt then. Pretty much no one uses Async apart from Jane Street.
> Also, OCaml does have a clone of Lens, but it was only written in the last year. Another example of how OCaml is very far behind Haskell in terms of development/production readiness.
OCaml has multiple lenses libraries including one written by Jane Street which are not used a lot because mutations are fine. This is not being far behind. It is being practical.
Ask yourself if you want to learn OCaml for practical reasons, or because you want to satisfy an internal itch. Both are perfectly good reasons to learn a language.
When I went to evaluate OCaml, I realized I wouldn't get much practical use out of it despite my curiosity, and focused on learning skills that I would get practical use out of. I'm happy with my choice, and plan to revisit the language to satisfy my curiosity at some other point.
Pretty interesting. Reading it, seems closer to a tutorial in using Owl, an OCaml-written package for technical computing (what e.g. Matlab is; though architecture differs according to post).
Leaving ocaml aside, the connection between scientific computing and hardware is the one thing I miss the most in "scientific computing" books and courses, because it sooner or later limits the science that any researcher doing scientific computing can do.
To give an example, earlier this week, one of our scientists was waiting 10 minutes between each interactive iteration of their data-set, so I was called to help, and the only feedback they gave was that "its slow", to which I replied "slow with respect to what? how fast are you expecting this to be and _why_?".
The answer to these questions is the difference between "maybe they just need a faster computer", "maybe they need a different algorithm", or even "maybe this problem cannot be solved today because computers this fast do not exist".
From their facial expression, it looked to me that they actually had never thought about any of this, probably because whatever they did before was always fast enough, but now this issue was limiting their science and they were lacking the bare minimum set of tools to even get proper help.
If you are doing scientific computing, chances are that the problems you are going to be dealing with are going to be getting bigger and harder as you advance in your career. For many scientists, the first problems will actually be big enough for the hardware to matter.
I wish scientific computing courses and books will at least provide the most basic tools to these scientist for them to at least be able to get meaningful help. Having someone on call for when this matters is quite expensive.
That's like saying "most programmers would never read a book about science, or finance...", or whatever field they are writing software for.
Observation in research support, I'd guess. It typically no longer seems to be the case that you do whatever you need to for your data.
So I'd expect that every year, there will at least be a class of 30 scientists taking this course.
It gives scientists a lot of information about how to perform low level optimization on code, e.g., if your code is "slow", use SIMD, OpenMP, BLAS, or do this or that trick.
But it does not provide the scientist with even the most basic tools to answer the question: "Is my code fast or slow?" (i.e. should I optimize it at all?), much less "_Why_ is it slow, and what's the best way to address that?" (e.g. if it is slow because its using 100% of the peak FLOPs of the CPU, but your hardware has a GPU, so you end up with 1% total FLOP utilization, then none of the "tricks" there will help).
It also completely avoids the issue that, in practice, a O(N) algorithm beats a OpenMP+SIMD-optimized O(N^2/p) algorithm pretty much all the time.
The chapter kind of assumes that scientists OCaml code will be slow, and gives them a "bag of tricks" that they can try to make it faster.
So we end up with the irony of a book on scientific computing that completely ignores the scientific method.
What would really be a bigger deal is some limited dependent typing to avoid errors from mismatched array sizes. Until then though, Julia is a bit more practical choice for me.
Shape in Computing [https://dl.acm.org/doi/10.1145/234528.234749 ]
A Semantics for Shape [https://www.sciencedirect.com/science/article/pii/0167642395... ]
https://www.semanticscholar.org/paper/The-FISh-language-defi...
https://link.springer.com/article/10.1007/s100090050037
The page for FiSH used to be online. I cant find it now.
This paper was my introduction to dependent typing, so if you have a little Haskell background, you should be able to grok its gist too.
...
> Indexing and slicing is arguably the most important function in any numerical library.
These statements are undoubtedly true. The first question any practitioner familiar with other systems will ask is: what does basic arithmetic, array manipulation, and linear algebra look like?
But from what I can tell on a very quick skim, that question isn't really answered until the section starting with these sentences, on page 123. I've noticed this situation every time I look at the Owl docs webpage too, FWIW (have not looked recently though).
I understand the need to be perceived as fully-capable for modern tasks -- and that's fine for a 2-4 page set of teaser examples up front -- but I think this book would become much more approachable if the basic mechanics of doing math were presented first.
An older thread on HN on slicing https://news.ycombinator.com/item?id=20457884
As a user of both, I think they have different treadoffs. I tend to use OCaml when I am playing around with the code because I find it infinitely easier to refactor (and to figure out what I was doing if I leave the code rotten for too long)
Library ecosystem seems better on F#, but I must admit I'm somewhat wary of the behemoth that is .NET .
What else should I consider?
https://devblogs.microsoft.com/dotnet/announcing-f-5-preview...
I had originally been learning ocaml but switched to f# because it has much better tooling and more uses (e.g. better web server support since I use .net libs)
From just preliminary research, F# seems both loved by its devs, but also found to be a bit of an unloved child on a sidetrack - that was probably another reason I hesitated in getting started. Do you think that is justified?
When I started I was pleasantly surprised how easy it was to download .net core on a mac and not do anything else to use F#.
Its less popular than C# so its there's not as much documentation from 3rd party sources, youtube videos, etc, but you can use any C# modules in F# so you'll get used to reading C# docs.
I would highly recommend this overview of the .Net ecosystem: https://www.youtube.com/watch?v=bEfBfBQq7EE
I really like how it showed a preview of "modern C#" at 12:20 which is being influenced by F#.
After that, this talk opened my eyes to the power of f# philosophy and type system: https://www.youtube.com/watch?v=2JB1_e5wZmU
I am curious as to whether the F# ecosystem is similarly fragmented by Windows-only libraries.
Almost anything with a CPU has some OEM shipping JVM implementations for them, while Microsoft focus only on the major desktop and mobile OSes (one of them with its own coffee brand).
https://devblogs.microsoft.com/dotnet/introducing-net-multi-...
Plus it doesn't help that some teams are adopting React Native, even alongside .NET.
Too much chaos going on sorting out WinRT, while most of us aren't willing to say how high when Microsoft says jump.
I have spent a couple of years doing .NET for life sciences, as many labs are mostly focused on Windows due to their laboratory robots and data readers.
So many researchers end up using a mix of Excel, VBA and MFC (old tech) and Forms/WPF (new tech) based tooling.
If that is your data science domain, I would definitely advise F#.
Would you then dis-recommend F#, or merely not recommend it for Windows compatibility?
Also you can use .NET Core.
Note that VS for Mac is no longer Xamarin Studio rebranded, nowadays it shares several common code with VS, which is one of the reasons why the plugins have moved away from COM to being .NET based.
If you get a feel for F# I'm sure it will be an easy leap to OCaml if you decide to do it later.
Incidentally I've just been running through the Advent of Code in F# over the last few days - I do think the toolset for F# is very kind for a new learner.
I went from Ocaml to F# because I wanted better tooling and a much larger pool of available library (any .net library can be used from F#), I did not look back.
Both languages are close cousin but F#'s syntax feels a bit more streamlined at the cost of less powerful type inference (the type inference in Ocaml is a thing of beauty that I have not found elsewhere).
There are also some academic projects with industrial uses. Directly to mind come Coq, Frama-C, Mirage and the Zélus compiler.
EDIT: added Inria and Frama-C
Both INRIA and the CEA uses OCaml heavily (Coq, CompCert, Frama-C). Cambridge uses it for MirageOS, Facebook to write software analysers and now web applications (the web version of Facebook Messenger), Citrix in XenServer. Bloomberg developed a compiler from OCaml to Javascript.
Languages like R, Python, Matlab, and Mathematica have a LOT of built-in capability in this area to do symbolic and numerical methods and data analysis kind of stuff (large sparse matrices... etc). You can do a ton in the high level language without ever dropping down to C or Fortran. So a scientist can just do their job without having to worry about how the plumbing works as others (open source community for Python and R) or vendors (Matlab and Mathematica) take care of that, which is a huge win.
Back when I thought about doing scientific work in Ada, Rust, OCaml, and Haskell, I was dissapointed to see that either there were many libraries missing that I would need, or the existing ones were incomplete. The solution is usually just to FFI to Python or C, but at that point, I think you lose a lot of the magic of sticking to 1 high level language. Just my 2 cents. If you're comfortable in doing all the scientific bits in C and just calling from OCaml, then you should be fine. On the data analysis side though, R or Python + Pandas just have sooo much inertia.
I'll add that this book's table of contents seems to have some really neat topics, but a lot are just a single page long and I still think a lot of areas (Ex: linear algebra) are going to be a lot less complete than what Python has to offer. With that being said, the scientific scene I'm seeing in this book is much farther along than I would've guessed, so that is a good sign.
You are right that the ecosystem is small, and will likely be always behind Python, R and Julia. But considering how small the community around numerical methods in OCaml is, I think the work that went into owl and this book is incredible and may help raising awareness and growing the userbase.
I think one of our current limitations is the lack of a high quality, full-featured plotting library. There are some which are used for publication level plots, but they are nowhere near the experience of plotly, matplotlib, ggplot2 or julia's plots.
In any case, the community seems to be slowly growing already, so there is some hope. I think the best would be if we can start integrating it with SciML and similar projects.
I think y'all are definitely raising some awareness and I agree the lack of a plotting library is a major hurdle. You can always call out to GNUPlot from the terminal, but most folks probably don't want to do that.
There is actually a GNUPlot backend for Owl. Works very well, but is a bit lowlevel: https://github.com/hennequin-lab/gp
Similarly, I use https://github.com/mseri/ocaml-gr but it is a safe lowlevel binding. I hope one day to have the time to wrap it into an interface similar to Julia's one.
I don’t think there is work going on on that side of owl at the moment if this is what you are asking.
But you should be able to call those languages from OCaml. For instance, RInside
https://cran.r-project.org/web/packages/RInside/index.html
lets you call R from any language with a C FFI. A very simple solution (even if not quite as convenient as a full OCaml solution) that lets someone write OCaml if that's their preferred language. I helped add that functionality to RInside and I've been doing that for years so I can use D for my research.
Of course there are reasons why you might integrate the two. Perhaps you have an R expert doing the statistics work and the back end developer gluing everything together. That's fine, but at some point the system becomes pretty confusing and is a giant leaky abstraction.
It's harder to do generic polymorphism in OCaml since the language has nothing like C++'s parametric templates or Haskell's type classes. It does have a very nice module system that has the required flexibility, but compared to these languages, I think it is clunky for the kind of genericity seen in scientific computing.
I think Julia has shown us that strong typing is not really needed for scientific computing.
Are parametrized modules really clunkier than parametric templates ?
I think you meant static typing, in opposition to dynamic. Julia is strongly typed like OCaml.
A far cry from OCaml's great type inference and compile time static checking.
>You can specify types, but they don't really do anything except help with performance.
For Julia it's the exact opposite though, types don't help with performance since the compiler infers them anyway regardless of declaring them or not (you only really should to specify in very particular cases where inference is not possible to avoid the compiler being too conservative, and if you overspecify types you might even end up with a slower program). You declare them for their main purpose of controlling dispatch.
Julia's type system is a core aspect of it's paradigm (multiple dispatch [1]), which is the key element in both it's performance and polymorphism. It's different from a gradual typed language in which the language can exist without the types at compilation time, but those can be optionally added to enhance safety or performance, Julia cannot be compiled without knowing the types, which will happen regardless of declaration.
[1] https://www.youtube.com/watch?v=kc9HwsxE1OY
That said, types in Julia are not used for compile time static checking, which is a compromise that for many tasks it's not worth it. Ocaml ML and Julia ML can coexist well exactly because they have vastly different compromises, with Julia focused on interactivity/fast prototyping and Ocaml in safety (like Lisp x Haskell, but with somewhat more pragmatic languages).
The company I work for uses Haskell for a network packet parsing engine, web api servers, CLI executables, build infrastructure (alongside Nix and NixOS), an interpreter for a custom programming language, a gateway/proxy service for AWS services, etc.
The Julia code is going to be shorter, faster, and more elegant. The libraries will be sooo much better. The static typing of OCaml doesn't really help in this area and sometimes actually hurts (statically typed DataFrames don't work so well).
r = [0.0001; -0.00002...]
I might be interested to find the cumulative returns: cumprod(1 .+ r) .- 1
Or just apply a function f to it: f.(r)
In Ocaml I can't vectorize any notation: List.map ~f:(fun x -> x - 1)) @@ List.cumprod(List.map r ~f:((+) 1))
List.map r ~f:f
Julia allows you to write code in a very vectorized, array language style. OCaml does not. This is big big issue, imo. Also with multi-dimensional arrays and slice notation, Julia is just very convenient for working with higher dimensional data. OCaml, to put it mildly, is not very good at this.Scientific computing and array languages go hand in hand. Also the lack of polymorphic functions is a big problem in OCaml. For example it would be impossible to define an addition function in OCaml that transparently worked with arrays:
1 .+ [1; 2; 3] == [2; 3; 4]
1+1 == 2
[1; 2; 3] .+ 1 == [2; 3; 4]
[1; 2; 3] .+ [1; 2; 3] == [2; 4; 6]
Julia makes this easy. A real array language like kdb+/q or J is even better.This is great for linear algebra (matrices and tensors). You can write very math-like equations using very high level functions. With OCaml you will always be burdened with the nitty gritty of mapping and folding over the lists/arrays.
Owl is fairly new, and it has hopefully been able to learn from other ecosystems like Python.
Specifically. OfS is an introductory OCaml book (first half) which has example usage of interest to scientists (second half). OSC though it has some introductory text per section is mostly concerned in showing Owl, which is what you'll end up using.