OCaml 4.03: Everything else
blogs.janestreet.com
blogs.janestreet.com
http://roscidus.com/blog/blog/2014/02/13/ocaml-what-you-gain...
I'm putting my notes on github, in case anyone wants a head start.
https://github.com/melling/ComputerLanguages/blob/master/oca...
Since I'm coming from Haskell I'm used to a vast Applicative and Monad vocabulary and making do with just `bind` and `return` is rather painful. So having a generic library of Monad combinators that one could use with any Monad would be great. Also, being able to just write `show x` is so nice!
Edit: while I often see "modular implicits are being worked on" it is not very clear whether there is a concrete plan to add them to the language. Is there any place in the official OCaml repository / issue tracking system / wiki etc where one could check the status?
You can see some things on:
https://github.com/ocamllabs/ocaml-modular-implicits
but you shouldn't take a lack of activity on there as a sign nothing is happening. For instance, Frederic is actively hacking on the prototype at the moment but hasn't pushed anything to that repo."I'm expecting there will be more news about those in the next six months".
I'd be very surprised not to see modular implicits in 4.04 now that the foundation has been put in place with Flambda[1]:
"Even if you're perfectly happy with OCaml's performance as is, Flambda is still an exciting change. That's because various upcoming language improvements like modular implicits (a feature that brings some of the same benefits as Haskell's typeclasses) will only really perform acceptably well with a good inliner in place".
All you need for a monad is `bind` and `return`. Is the problem that there are not polymorphic monad combinators like `sequence`, `traverse`, etc.?
* Since I am only dabbling in OCaml and only use `lwt` and `async` occasionally I'm not entirely sure. But judging by the usual structure of Haskell libraries, it should help
https://github.com/Chattered/ocaml-monad
and use it regularly when writing Ocaml. I'd like to get back to it and try to implement some MTL style interfaces (MonadState etc...). Edward Kmett says this is not likely to work, but that sounded like a wager to me.
The packages that I needed and that didn't compile in the standard repository were working thanks to patches applied in his repository.
Also, the maintainer was very prompt and helpful when dealing with my pull request.
There are other forks of that repository, but when googling for "opam windows" they do not appear on the first page, which is unfortunate since they seem to me the best currently available option.
Ideally, there'd be an official Windows port and installer that doesn't depend on Cygwin. Certainly that takes time and effort to develop and maintain, so it's nice to see projects like this in the interim.
So aside a short gig to port old Caml Light code that I had lying around into OCaml, I actually spend my ML like coding in F#.
Sure I was disappointed, but sponsoring the people working on OCaml is more important.
The writers of the book have bills to pay.
On the Haskell side of things--it's biggest seller (functional purity) is also the biggest downside. To give you an idea, papers have been written about the best way to implement common data structures purely functionally.
From what I've seen, OCaml has less "abstraction overhead" because it can dip into imperative code. While Haskell is really fun, it definitely has moments where you want to beat your head against your desk.
(0) Haskell starts with a bad default - laziness.
(1) Haskell uses special annotations to introduce strictness. The presence of these annotations isn't tracked by the type system, which reduces the usefulness of the type system as a tool for understanding your program.
(2) More sophisticated evaluation strategies (e.g., memoizing functions, which can be seen as a generalization of laziness) are difficult to achieve in Haskell, even though it's completely straightforward in ML.
+ OCaml is eagerly evaluated while Haskell is lazily evaluated. This makes it easier to reason about things like memory use in OCaml.
+ IO is reflected in the type signature of Haskell functions. You may find this annoying because it stinks to have to change a lot of type signatures just because, e.g., you want one of your utility functions to make a log entry when called.
On the other hand, IO is often a huge deal either semantically or from a performance perspective. For big projects having IO reflected in the type system can be a huge help.
+ The ecosystems are different. Haskell's is bigger, though I've heard the quality of OCaml libraries tends to be very high.
What you mean is, "effects are reflected". Haskell has a monadic effect system (IO is very rough-grained part of it), which can be combined via monadic transformers.
Situations where you don't want to use OCaml:
- you need good parallelism and using multiple processes are not enough (concurrency is fine, though)
- you want a large pool of developers (also applies to Haskell, but less so)
- you like monadic effect systems
- you need a number of libraries which are not present in the OCaml ecosystem
- you want a build system that doesn't suck
- you want a good standard library
- you want typeclasses
Why you'd want to use OCaml over Haskell:
- fast compilation (especially if you compile to bytecode)
- no monadic effect system
- no awful, ridiculous record field name collisions (OCaml lets you have two distinct types with an "id" field, imagine that)
- best-in-class package manager
- faster than Haskell (I think?)
- eagerly evaluated (Haskell's informal motto is "if it compiles it works", complemented by "until you get a memory leak") and therefore makes it much easier to reason about performance
- high-quality ecosystem (though not always very well documented)
- labeled arguments (many Haskell libraries have these functions with lots of arguments which are quite confusing in the absence of labeled arguments)
- great support for Vim/Emacs (IMHO, the Haskell equivalent to Merlin/ocp-indent are not nearly as good, or at least were not as a few years ago)
- great REPL via utop (just don't use the standard one, it's terrible)
- an object system if you really need one
- functors
- pleasant "printf debugging" option, with a possibility of using a real debugger if you need
This is a very surprising claim. Do you have any examples how Ocaml specifically induces systems built on it to suck?
It already starts with the compiler, calling them manually is quite a pain, especially if you want include any kind of library, which is why ocamlfind is such a huge win.
In fact with its current compiler (AFAIK) it can only be run as a web server, which means it's not even possible to write a program which just writes Hello World to stdout and exits. This was frustrating enough to me that I soon gave up on the language. :(
Lack of global type inference is far from the only problem with Scala.
A better module system for Haskell is in the makings (Backpack by EZ Yang), but may or may-not be what you are looking for. To my understanding the module system of Haskell is not very limiting at all.
In a strict language, a lazy thunk can be represented as a single mutable cell. In ML, thanks to type abstraction, the mutation is confined to the module that implements laziness, and, in the rest of the program, there's no way to tell that anything impure is happening.
> To my understanding the module system of Haskell is not very limiting at all.
Even Haskell's own designers admit otherwise. Haskell actually originated as an attempt to standardize a lazy, purely functional language for the research community. Modules and records weren't very high in the priority list, and it really shows. But sometimes behind a black cloud there's a silver lining, and the limitations of records inspired Haskell programmers to invent lenses, which are very, very, very awesome.
But, to this day, Haskell still doesn't have a decent alternative to ML-style modules and functors. As far as I can tell, in vanilla Haskell (no extensions), you can only encode functors on modules with a single type component. With either type families or multiparameter type classes, you can encode modules with more than one type component, but the result is very awkward.
Unfortunately, the general story isn't quite so nice, due to thread-safety.
Immutable data structures are great for thread-safety, as there's no problems with data races or generally invalidating things, because there's no writes... but laziness introduces secret writes, among other problems. This means a type providing laziness either can't be used across multiple threads simultaneously, or is forced to have extra overhead (e.g. atomic operations, blocking & registering for being woken up when trying to force a thunk that another thread is already forcing), and users of a library definitely need to know if a module is using mutation in the former way, and may want to know about the latter.
Of course, if shared-memory concurrency/parallelism is forgone, this isn't a problem, but that omission has its own downsides. And... I would suspect the overhead of just implementing it in a thread-safe way can be considered negligible in many cases, especially if data is guaranteed to be pointer-sized (or less).
(Concurrency/parallelism are sometimes described as "abstraction breaking" for this sort of reason.)
* Idris : modules?, strict by default, effect tracking by default, higher-kinded types, parallelism I think (at least on some platforms)
* Ceylon: modules, strictness, experimental higher-kinded types (possibly only on JS), parallelism (possibly only on Java)
* F#: modules, strict by default, parallelism
* Scala: some modularity, strict by default, higher-kinded types, parallelism
(0) “In principle”, you could encode modules in Idris using dependent sum types, but... Good luck with that! It isn't going to be terribly usable. In general, Haskell and related languages don't consider it worth the effort to give modules proper types. I guess we can be thankful Agda's modules don't suck as badly as Haskell's - at least they can be nested.
(1) Ceylon and Scala's type systems are sophisticated enough to encode some of the use cases for modules with objects, but invariably the result is awkward. And other use cases are just impossible. For example, how should I encode an ML functor that takes as arguments two modules with shared type members? If it's possible at all, I don't even want to imagine how horrifying it will be to manually turn all the sharing by fibration into sharing by parameterization.
(2) F# doesn't allow any encoding of ML-style modules. Nothing. Nichts. Nada. It's the only so-called “ML dialect” where the most important feature from ML is completely missing.
(0) The presence of a value component.
(1) The presence of a type component.
(2) The concrete definition of a type component, while still remembering its presence.
I am not sure you can do this in Idris, other than manually shuffling data between multiple dependent record types. Which is rather inconvenient: If you have a module with 20 components, the last thing you want to do is manually shift 15 of them to another module.
Also, unlike vanilla Haskell (no extensions) type classes, which can only have a single type parameter, ML modules can have more than one type component.
As someone else mentioned, Backpack is trying to "fix" the module system in Haskell[2], which should be interesting. That said, I don't have much of a problem with it, but then again I haven't seen the light of the OCaml way yet :)
[0] https://ghc.haskell.org/trac/ghc/wiki/StrictPragma
[1] http://blog.johantibell.com/2015/11/the-design-of-strict-has...
data Front a = Nil | Cons a (Stream a)
type Stream a = Lazy (Front a)
Haskell doesn't do this, not even with the Strict or StrictData pragmas.Technically I'm on solid ground, IO _is_ reflected in the type system and not all effects are (memory use being the big omission, which we both mention). Saying "effects are reflected" is probably more helpful though.
> monadic transformers
There are ways combine capabilities without using monad transformers. For instance with typeclasses (if any non-haskellers are reading this here's an example: https://github.com/commercialhaskell/stack/blob/8b010060b0d7...).
> IMHO, the Haskell equivalent to Merlin/ocp-indent are not nearly as good, or at least were not as a few years ago
Editor tooling is still a big weakness of haskell:( A lot has been done on it though and progress is starting to pick up.
I agree 100% with your other points.
My general experience with OCaml hasn't been great. To be blunt, it seems sort of like a less-well-thought-out version of Haskell.
The biggest downsides that I remember were:
* IO sequencing via (;) : () -> () -> (). This is sort of a relic from before we had monads, and is relatively pretty clumsy.
* No type classes. "Generic" operators like (<) work by binary comparison (!!!). This is terrible and leads to nonsensical ordering semantics for non-trivial data. There's also no generic (+), (), etc. There are separate functions for ints and floats. Again, clumsy.
Questionable unboxing techniques. On a 64 bit processor, ints are 63 bits because ocaml has to use a bit to tag them as non-pointers. On the other hand, Haskell differentiates boxed/unboxed values statically at the kind level, so there are no weird pointer tagging tricks or related overhead.
I have limited experience with OCaml, so take this with a grain of salt. I also think OCaml has valid use cases. In Jane street's case, I can see OCaml being a better choice than Haskell for the reasons I mentioned. The people there definitely know the advantages and disadvantages of both. One of the hardest interview questions of my life was a very interesting Haskell question from Jane Street!
Java programmers would be confused the most, because they would have to unlearn the most. In Java, a class is a type on its own right. An object of class (and, hence, type) `Foo` has a very specific data structure, even if this data structure is unknown to the user. On the other hand, in OCaml, a class is just a constructor for objects of a particular structural type. Two objects with the same structural type can have completely different underlying data structures. Making things even more confusing, inheritance doesn't entail subtyping - a class `foo` can inherit from a class `bar`, without `foo`'s type being a subtype of `bar`'s type. This happens if `bar`'s type contains negative occurences of its self-type. As a result, few intuitions about Java's class system carry over to OCaml.
Python programmers would be somewhat less surprised, but if their idea of a typed object-oriented language is “something that looks like Java”, they could encounter all the difficulties mentioned in the preceding paragraph. On the other hand, if they have no experience at all with typed languages, OCaml's class system can still be a source of pain, in that inferred object types can be utterly incomprehensible for someone not familiar with how the type checker works. By contrast, the inferred type signatures for the non-object-oriented subset of OCaml are usually very tame.
---
Sorry, e_d_g_a_r, I can't reply to you directly, because the website says I'm “submitting too fast”, but here is an example: http://pastebin.com/yUmNnFD5 . By the way, this isn't a bug in OCaml's type checker - it's the way it's supposed to work.
lolcathost% ocaml
OCaml version 4.02.3
# class foo =
object (_ : 'a)
method test (_ : 'a) = ()
end;;
class foo : object ('a) method test : 'a -> unit end
# class bar =
object
inherit foo
method other = ()
end;;
class bar : object ('a) method other : unit method test : 'a -> unit end
# let f = new foo;;
val f : foo = <obj>
# let b = new bar;;
val b : bar = <obj>
# f#test f;;
- : unit = ()
# b#test b;;
- : unit = ()
# f#test b;;
Characters 7-8:
f#test b;;
^
Error: This expression has type bar but an expression was expected of type
foo
The second object type has no method other
# b#test f;;
Characters 7-8:
b#test f;;
^
Error: This expression has type foo but an expression was expected of type
bar
The first object type has no method other
#OCamal also suffers from portability issues but haven't tried that hard. I love the idea of OCamal but the practical/pragmatic side of me gets rubbed the wrong way.
Personally Racket has a ton to offer and has been growing its user base a lot recently. I love this language and you can amke it whatever you want to do. Racket is a progamming language for programming languages. Clojure gives you a more pragmatic functional language and running on the jvm with a very large community.
There are a number of much anticipated features that haven't made it into this release. In particular, the multicore GC, which at one point had been expected to land in 4.03, has been pushed back, likely to 4.04.
~~ Well, I know, it's hard. ..
Jane Street certainly contributes a huge amount to the ecosystem, both in terms of code and other support, but they're not the only people pouring effort into it.
For a programmer, they allow you to be more precise about the point in time where an object should be regarded as a weak object. That is, an object which can be collected when under GC pressure. A weak references edges itself toward this goal, but often require some manual intervention and/or knowledge of the GC world to manage by the programmer. An ephemeron removes this additional knowledge from the programmers mind. It allows one programmer to make an interface which is truly not leaking in GC abstraction. So other programmers don't have to know about it at all.