Standard ML in 2020
notes.eatonphil.com
notes.eatonphil.com
The syntax is super-easy to learn (The BNF for the whole fits in a mere 2 pages[0]), but contains a lot of features in that small package. Rather than tacking on functional features (eg, Java with lambdas), these features have been carefully considered and streamlined and include bits like proper tail calls and currying.
You get nice bits like actually sound types (hindley-milner types as Milner was also one of the SML spec authors), generics (way better than typical interfaces), type inference that actually works, and modules (super-powerful encapsulation). Pattern matching in all it's awesomeness is also on display.
SML has an amazing concurrency story (CML is rather like golang channels, but better with better typing and a bit more flexibility) and compilers like PolyML or Mlton are very fast (once again, around the same as golang).
Despite this, the language doesn't have the academic flaws of its descendants like Haskell.
SML isn't a lazy language, so reasoning about performance is much easier than some other functional languages. SML doesn't pretend the world is a pure function. You are free to make functions that have side effects. While most primitives and data structures are immutable by default, they either have mutable variants (eg, vector and array) or can be used as if mutable with refs (something like typesafe pointers without all the reference/dereference bits).
I used it, just for fun, in a static site generator. That was fun and worked great.
Can you elaborate on what those flaws are?
Personally, I wonder if uniqueness types in languages like the sadly forgotten Clean (a close relative of Haskell) would have been better than the whole Monad/Arrow thing. All of that abstract theoretical stuff is certainly fascinating but it's never seemed worth all the bother to me.
On the other hand, in SML you get all the benefits of strong HM type inference/checking, immutability by default, etc, etc.. while also still being able to just `print` something or modify arrays in place. SML's type system isn't higher-order like Haskell's so its solution to the same problem Haskell solves with type classes isn't quite as elegant but otherwise SML is the C to Haskell's C++.
Ignoring maths that underlie computation when writing software is like ignoring math that underlies mechanics when building houses: for some time you can get by, but bigger houses will tend to constantly fall or stand a bit askew, and the first woodpecker to fly by would ruin the civilization, just as the saying goes.
Uniqueness types were a nice idea indeed. I suspect linear types (or their derivative in Rust) do a very similar thing: data are shared XOR mutable.
Analogy to monads is this: yes, there are mathematical formalisms that can describe complex systems. Everything can be described by math! That’s literally its one job. But will the formalisms actually help you when the job you’re doing has insufficient achievable precision (eg finding market fit for a website)?
Writing code is a much more precise activity that cutting wood. Using "cheats" can get you somewhere, but due to the precise nature of the machine logic, the contraption is guaranteed to bend and break where it does not precisely fit, and shaving off the offending 1/8" is usually not easy, even if possible.
In my experience, this is exactly the problem Haskell's core abstractions—monads among them—help with!
What do you need when you're trying to find product-market fit? Fast iteration. What do Haskell's type system and effect tracking (ie monads) help with? Making changes to your code quickly and with confidence.
Haskell provides simple mathematically inspired tools that:
1. Make clean, well-factored code the path of least resistance.
2. Provide high-level, reusable libraries that keep your code expressive and easy to read at a glance despite the static types and effect tracking.
3. Give you guardrails to change your code in systematic ways that you know will not break the logic in a wide range of ways.
If I had a problem where I'd need to go through dozens of iterations before finding a good solution—and if libraries were not a consideration—I would choose Haskell over, say, Python any day. And I say this as someone who writes a lot of Python professionally these days. (Curse library availability!)
Honestly, Haskell has a reputation of providing really complicated tools that make your code really safe—and I think that's totally backwards. Haskell isn't safe in the sense that you would want for, say, aeronautical applications; Haskell programs can go wrong in a lot of ways that you can't prevent without extensive testing or formal verification and, in practice, Haskell isn't substantially easier to formally verify than other languages. And, on the flipside, Haskell's abstractions really aren't complicated, they're just different (and abstract). Monads aren't interesting because they do a lot; they're interesting because they only do a pretty small amount in a way that captures repeating patterns across a really wide range of types that we work with all the time (from promises to lists to nullable values).
Instead, what Haskell does is provide simple tools that do a "pretty good" job of catching common errors, the kind of bugs that waste a lot of testing and debugging time in run-of-the-mill Python codebases. This makes iteration faster rather than slower because you get feedback on whole classes of bugs immediately from your code, without needing to catch the bug in your tests or spot the bug in production.
Knowing a little bit of maths, but not much category theory, the wiki entry for Monads(Category Theory) is pretty clear. First sentence: A Monad is an endofunctor (a functor mapping a category to itself), together with two natural transformations required to fulfill certain coherence conditions. Easy.
Knowing a bit of programming, but not much Haskell, reading the entry for Monad (functional programming) or any blog post titled "A Monad is like a ...", it almost seems as if the author is more confused about what a Monad is than me. The first sentences of the Wikipedia article for example are a word-salad. With dressing.
From an outsider's perspective, it's almost as if a monad in functional programming is not a 1:1 translation of the straightforward definition of category theory, leading to an overall sense of confusion.
The formal definitions are straightforward enough, but the definitions alone don't really motivate themselves. A lot of mathematical maturity is about recognizing that a good definition gives a lot more than is immediately apparent. Someone without that experience will want to fully understand the definition, and fairly so -- but a plain reading defies that understanding. That is objectively frustrating.
> As a bit of an aside, it's interesting that the CS (Haskell, mainly) descriptions of a Monad are much more complicated than the math.
I actually do agree with this, though. I feel like monads are much simpler when presented via "join" (aka "flatten") rather than "bind"; and likewise with applicative functors via monoidal product (which I call "par") rather than "ap". "bind" and "ap" are really compounds of "join" and "par" with the underlying functorial "map". That makes them syntactically convenient, but pedagogically they're a bit of a nightmare. It's a lot easier to think about a structural change than applying some arbitrary computation.
Let's assume the reader knows abut "map". Examples abound; it's really not hard to find a huge number of functors in the wild, even in imperative programs. In short, "map" lets us take one value to another, within some context.
Applicative functors let us take two values, `f a` and `f b`, and produce a single `f (a, b)`. In other words, if we have two values in separate contexts (of the same kind), we can merge them together if the context is applicative.
Monads let us take a value `f (f a)` and produce an `f a`. In other words, if we have a value in a context in a context, we can merge the two contexts together.
Applicative "ap", `f (a -> b) -> f a -> f b`, is "par" followed by "map". We merge `(f (a -> b), f a)` to get `f (a -> b, a)`, then map over the pair and apply the function to its argument.
Monadic "bind", `(a -> f b) -> f a -> f b`, is "map" followed by "flatten". We map the given function over `f a` to get an `f (f b)`, then flatten to get our final `f b`.
It's a lot easier to think about these things when you don't have a higher-order function argument being thrown around.
[0] https://wiki.haskell.org/wikiupload/d/df/Typeclassopedia-dia...
- Monad is just an interface (with nice do-notation, to be sure), but it's elevated to an almost mythic status. This (a) causes Monad to be used in places which would be better without (e.g. we could be more generic, like using Applicative; or more concrete by sticking to IO or Maybe, etc.) and (b) puts off new comers to the language, thinking they need to learn category theory or whatever. For this reason, I try to avoid phrases like "the IO monad" or "the Maybe monad" unless I'm specifically talking about their monadic join operation (just like I wouldn't talk about "the List monad" when discussing, say, string splitting)
- They don't compose. Monad on its own is a nice little abstraction, but it forces a tradeoff between narrowing down the scope of effects (e.g. with specific types like 'Stdin a', 'Stdout a', 'InEnv a', 'GenRec a', 'WithClock a', 'Random a', etc.) and avoiding the complexity of plugging all of those together. The listed advantages of SML are essentially one end of this spectrum: avoiding the complexity by ignoring the scope of effects; similar to sticking with the 'IO a' type in Haskell (although even there, it's nice that Haskell lets us distinguish between pure functions and effectful actions).
Haskell has some standard solutions to composing narrowly-scoped effects, like mtl, but I find them to be a complicated workaround to a self-imposed problem, rather than anything elegant. I still hold out some hope that algebraic effect systems can avoid this tradeoff, but retro-fitting them into a language can bring back the complexity they're supposed to avoid (e.g. I've really enjoyed using Haskell's polysemy library, but it requires a bunch of boilerplate, restrictions on variable names and TemplateHaskell shenanigans to work nicely).
The ironic thing is that these algabraic effect systems have a feel/control flow pattern that's quite similar to exception handling from the OO languages or interrupt handlers in low-level code.
It's much easier to just say "this piece of code is impure, it may do X, Y, and Z and if so then ..." than to try and shove everything ad-hoc into the abstract math tree. But then you lose the purity of your language and it's really awkward in a language whose primary concern is purity. That may be a reason why algebraic effects seem a bit more natural in OCaml.
It's used pervasively in low-level code because the idealism of monads just doesn't cut it. In my opinion, this proves that the pragmatism of side effects is a necessary evil for actually getting things done in a performant way.
Lazy evaluation has the same kind of issues. Most humans don't think that way, so performance suffers. This may be a universal problem as in my experience, Haskell doesn't have good tooling to help with this issue.
Always immutable is another large issue. Yes, there's a ton of safety in immutability and it's the right tool for MOST code. Compilers aren't prefect and there seem to be an endless stream of situations where the compiler can't figure out if it is safe to mutate "immutable" data to gain performance. For the foreseeable future, the ability to mutate can have huge performance dividends.
Finally, Haskell and it's libraries are prone to rather academic programming styles and techniques. These are amazing and beautiful. They also can be hard to grep even if you know the math. When you consider that the overwhelming majority of programmers don't know the math, it seems plain that these constructs are a detriment to pragmatic, non-academic usage.
> It's used pervasively in low-level code because the idealism of monads just doesn't cut it. In my opinion, this proves that the pragmatism of side effects is a necessary evil for actually getting things done in a performant way.
It seems like you aren't very familiar with the ways Haskell programmers deal with side effects and mutation. UnsafePerformIo is sometimes needed, and there are a few common idiomatic ways to use it, but there are usually better options. What at first looks like an impenetrable wall between the IO monad and pure code is actually surprisingly permeable, just not (usually) in ways that violate expected semantics. For instance, you can read a file into a string in the IO monad and pass the string into pure code to process it. The string is lazy, so the pure code actually triggers the file to be read as it's consumed.
Another escape hatch is the ST Monad. It allows you to run imperative computations (i.e. use mutable variables and arrays) within pure code. Since the ST computations are deterministic, they don't violate any important guarantees about pure code. You can't read and write files from the ST Monad, but if you want to do array-based quicksort or something similar it's available.
Some aspects of Haskell are hard to work with. You're right that laziness introduced some performance issues that can be tedious to fix, and some of the libraries are hard to use. However, it's a perfectly reasonable tool for many general-purpose programming applications. It's not a replacement for C, but you could say the same thing about C# or Java or any other language with a garbage collector.
(edit: fixed grammar typo)
Anyways, the problem is that you might open a file, read it into a string, close the file, and then pass the string into some pure function. The problem is that the file was closed before it was lazily read, and so the string contents don't get populated correctly; you get the empty string instead, or whatever was lazily read before the file was closed.
That was a design oversight from the early days. If you know about it, it's pretty easy to work around it in simple applications. There are some more modern libraries for doing file IO in a safer way, but I haven't used them and don't really know what the details are.
On the other hand, in a program that handles a small number of files and doesn't live long after file access anyway (say a small script to do grab a couple things, crunch a few numbers, and throw the result at pandoc), there's nothing wrong with lazy IO and it can be quite convenient.
{-# LANGUAGE OverloadedStrings #-}
import qualified Data.ByteString.Lazy.Char8 as BS
main = BS.interact go
go :: BS.ByteString -> BS.ByteString
go input = -- Generate output here
If you don't mind Haskell's default String implementation, then it's just: main = interact go
go :: String -> String
go input = -- Generate output herehttps://existentialtype.wordpress.com/2011/05/01/of-course-m...
The module systems have some small differences, but are very similar (I believe OCaml's was inspired by SML's).
Standard ML is also standardized (hence the name), and there are multiple implementations, whereas OCaml is basically defined by its single implementation.
But OCaml does have a larger ecosystem, Opam, more features (first-class modules, polymorphic variants, etc.).
At a first approximation, Haskell pretends the world is a piece of state, not a pure function.
Mercury literally pretends the world is a piece of state. You explicitly say "Here, run this code. The world before it runs is on variable `a`, place the world after it runs on variable `b`".
Rust's performance is almost certainly much better, with a larger library ecosystem.
(Common) Lisp has macros, which allow me to implement pattern-matching; an even simpler BNF (as short as one line, depending on how you define it); dynamic typing, which makes generics unnecessary, and for all the hate that it receives is regularly deployed to production systems on large scales; tooling that is almost certainly better than SML's; and a condition system, which beats any other type of error-handling system, bar none (meaning that I can make my programs more robust than yours).
* Doesn't have macros, which in my 16 years of experience is a plus, macros pollute, should at most be an implementation detail
* Has types
* Doesn't have parantheses and the annoying prefix notation
* Your knowledge of SML can be translated easily to OCaml, Java, Scala, which are more or less part of the same family
SML over Rust:
* I don't have to reason about lifetimes, I'm not in a constrained environment
* TCO (tail call optimization)
* At its core, Rust is still an imperative language, it was influenced by ML but not that much as many believe
SML over Scala:
* It doesn't suffer of Haskellism and Extreme Bipolarity (oop/fp) and gets the job done.
> * Has types
Your words betray your ignorance. Common Lisp has a type system. Have you ever written it? Do you actually know anything about it? Moreover, macros are a plus - you clearly have not actually used Lisp macros, which are categorically different than macros in other languages.
> * Doesn't have parantheses and the annoying prefix notation
Spelling mistake. Additionally, the parenthesized syntax is a choice, and one that you get used to quickly. Subjective, and therefore not valid as one of the items you listed.
> * Your knowledge of SML can be translated easily to OCaml, Java, Scala, which are more or less part of the same family
Meanwhile, knowledge of Common Lisp transfers to almost every dynamically-typed language ever made, and a good many static ones - including Java and Scala - because Lisp influenced all of those languages. That is, Lisp knowledge transfers to far more languages than SML knowledge does.
This isn't really okay. It's fine if you disagree, but I don't get why such aggression is needed.
I don't know if Lisp, the language, influenced Java and Scala so much, it was more about its runtime, the garbage collector, etc. The only thing that I liked about Lisp is the composabiltiy, and coupled with objects from Simula created my favorite industrial languages, Java and TypeScript. I'm passionate about parsers, compilers and transformers and this is why I have a fost spot for SML, OCaml. I simply dislike Lisp, interesting language but not for me.
Can you show me a significant body of literature or a whole community that uses "types" to mean "a strong type system"? Because I've never heard that aliasing done before.
> I don't know if Lisp, the language, influenced Java and Scala so much, it was more about its runtime, the garbage collector, etc.
Lisp pioneered lambdas and first-class functions, in addition to garbage collection, all of which were adopted by Scala and Java later.
> I simply dislike Lisp, interesting language but not for me.
I understand! I have things that I like and dislike, too. However, it's extremely frustrating when you ask for an analytical comparison between several languages, and some random person lacking knowledge on at least one of the languages (I suspect you don't know Rust either) comes in and, instead of actually providing relevant information, gives fallacies and opinions.
In Common Lisp objects are values, so runtime tags are attached to data objects => which are actually values => which are stored in variables => which are not typed.
You can make it optionally typed at compile time to help the compiler but that means to annotate your code with 'type' keyword and it's verbose. The compiler does not enforce it.
Rust is very much a ml variant. It's all the things you claim and probably more. It's also at least an order of magnitude more complex. If you know rust, you'll probably find SML super easy to learn and refreshingly easy to code.
People gravitating toward SML likely want a few things: easy to learn, simple infix syntax with functional support, compile time type static type checks, multithreaded, very fast without a drastic recode.
Common lisp libraries are better, but not drastically so (well, I've never had 5 grand to throw at lisp works, so I can't comment there). CL is more complex to learn (must people can probably learn the entire SML language in the time it takes to master the loop macro). Macros are powerful, but mean that even if you can find a lisp dev, it'll take a long time for them to learn your custom variant of the language.
CL offers unofficial threading, but last I checked it was much harder than the SML solutions. CL is pretty fast out of the box, but if you ever need peak performance, the code changes a lot and can become very finicky (though CL makes seamlessly hiding those bits easier than most languages). Even at it's most optimized, I don't know that it can match milton in performance or memory usage.
Dynamic vs static typing is a battle that's all but over. When you have devs coming and going on a team, type make transitions easier. Loads of dynamic languages have started adding them and even CL sort of does.
CL has loads of interesting features from well known ones like macros/reader macros or metaobject protocol to less known ones like optional dynamic scoping.
I'm just saying why I enjoy SML (not trying to argue that it's the end all, be all of programming). There are other great languages too. Do what you love.
[1] https://www.coursera.org/learn/programming-languages
Edit: nevermind! I checked out their docs again and it seems to be decently translated. I'm still not sure where their code is actually hosted, just that you can download releases. So added it to the list of major implementations!
The tradeoff is that it can be more laborious than other functional languages. My code would be shorter if it had record updating, a single-character lambda syntax, contextual operators for real or int arithmetic, etc. Not only does it lack syntactic sugar for associative containers, it doesn't even have them in the standard library. Nor sorting, nor random numbers, nor complex numbers or matrices, nor uniform type-to-string formatting or string interpolation. Some of these holes can be filled by libraries, but then library choice becomes a problem.
I can live with an awful lot of that though, given the relative clarity of the language, the really practical module system, and the delightful feeling you get when using a language standardised over 20 years ago that, once written, your code can stay written.
[1] https://people.mpi-sws.org/~rossberg/papers/sml-defects-2013...
I'm also pretty jealous of OCaml's modular implicits, just because it saves you typing the module name before an operator (i.e. MyModule.+ vs just + implied by context).
It also kinda sucks that operator precedence is defined at the module level and cannot be exported. You have to always redefine a library's operator precedence yourself.
Have modular implicits already been implemented? I thought they were still trying to figure out the details of the resolution...
But I think Rossberg has since moved on, and his latest attempt at a "better ML design" should be 1ML[0].
1. Both Vim and Emacs are horribly annoying with their automatic indentation for SML. There is a lot of fighting against the editor in this department. In Emacs e.g. you have to delete whitespace all the time, because otherwise you'd have top-level function definitions shifted 80 characters to the right.
2. I've used Poly/ML and SML/NJ so far. Both of them are purely interactive, meaning I can't just compile a program into ELF and ship it somewhere else without the compiler. That makes it a no-go for me for real-world use.
3. The interactive modes of Poly/ML and SML/NJ don't support readline shortcuts. They are the most cumbersome REPLs I've ever used.
4. Inline type declarations (as opposed to Haskell-style type declarations on a separate line) are very noisy - they make reading the code harder. Omitting them (which is the rule in SML in practice) leads to hard-to-decipher compilation errors when you write a new piece of code and you made an error somewhere which confused the type-inference about your intentions. Suddenly forgetting about a word or a set of parentheses in one function results in errors in another perfectly-good function. It's the horror of C++ templates all over again.
Regarding your points:
1. Yeah editor support sucks. I normally edit in text mode with my own minimal keyword highlighting or ocaml-mode.
2. Poly/ML can definitely generate binaries! Most distros ship with `polyc` that will build the binary for you. But this is just a shell script around opening the REPL and calling some dump image function (like how you build a binary on SBCL). MLton and Poly/ML definitely allow you to build binaries. I don't know about SML/NJ.
3. For sure a pain. I use rlwrap [0] to work around this, which is ultimately simple enough!
4. Interesting! I personally find Haskell-style decorations so much more a pain since they're not inline. SML is very much like other major languages in the way it does inline types (TypeScript, Go, C#, etc.).
3. Thanks for the tip! That will save me a world of pain.
To these points, I'd say
1. I find Emacs sml-mode is just about up to it. I do have complaints - I wish it would pull back the indentation more in lines that continue an expression, it doesn't handle multiple "where" clauses elegantly, it misaligns anonymous function alternatives - but every time I consider trying to fix them I decide they don't quite upset me enough. It's a slight pity though because I strongly believe a language should be auto-indentable (life's too short to indent code yourself), and SML is, just not quite with the existing mode.
2. I like to use Poly/ML for automatic builds during development and MLton for "production" builds - both producing executables. There are still problems on Windows, which doesn't have a properly native MLton port - the existing one uses MinGW which is ok-ish but not what I would prefer. (MLton has a code generator that produces C, so the limitation is that the runtime hasn't been ported rather than with the compiler itself.)
3. Agree, "rlwrap poly"
4. I like inline type decorations, but I also like to omit them most of the time. I think you do get some feel for when it's a good idea to add them, to clarify things for the call site or check your own intuition about the deduced types. Module boundaries (with signatures) also form a natural firebreak for out-of-control type errors.
Do you have/use any dependencies? Like a http/2 (or tls capable) web server, a database client or a gui library? If you do, how does that work with two implementations - if not... What kind of programs do you write/problems are you solving?
It doesn't seem to send the right message these days if the idea is that working programmers are building lambda calculus interpreters or theorem provers? Even interpreting a lisp would seem more practical.
I think Practical Common Lisp was more in the right track for content but today I'd focus more explicitly on language design or backend system design.
Andrew Appel's book does already fill the language design slot though nicely.
I've used sml-mode in Emacs for years, and I've never had the problem you're describing. On the contrary, I find sml-mode to be quite adequate for my needs. Can you elaborate a bit on what's causing you problems?
Edit: Ah, I see what you're saying. After finishing a function definition and starting a new, the cursor is wildly indented, that's true. But if you just type `fun` and press tab, sml-mode automatically indents the definition correctly.
> 2. I've used Poly/ML and SML/NJ so far. Both of them are purely interactive, meaning I can't just compile a program into ELF and ship it somewhere else without the compiler. That makes it a no-go for me for real-world use.
I can recommend MoSML for interactive development. It can also produce compiled binaries. When I want to produce efficient compiled code, I usually use MLton.
> 3. The interactive modes of Poly/ML and SML/NJ don't support readline shortcuts. They are the most cumbersome REPLs I've ever used.
Use the REPL in emacs, it's excellent.
> 4. Inline type declarations (as opposed to Haskell-style type declarations on a separate line) are very noisy - they make reading the code harder. Omitting them (which is the rule in SML in practice) leads to hard-to-decipher compilation errors when you write a new piece of code and you made an error somewhere which confused the type-inference about your intentions. Suddenly forgetting about a word or a set of parentheses in one function results in errors in another perfectly-good function. It's the horror of C++ templates all over again.
I agree, this is a pain point. I usually leave type declarations in comments before the definitions, but that of course has obvious drawbacks.
Have you tried rlwrap(1)? I use it with the OCaml REPL and it works well enough.
For fun and educational purposes, I began working on a Standard ML compiler [1] and VSCode extension in Rust as well - with the end goal being a psuedo-clone of MLton. Currently taking a short break from it to work on some other stuff, but I'm mostly done with monomorphization
should I consider taking a look at some description of standard ml instead of rust? It seems like rust has pattern matching and immutability which matches to F# pretty good?
My use cases are mostly desktop class machines, with optional extensions to single board computers and maybe, maybe, mobile devices of some kind. Utility programs, soap and rest web services, guis for the rest apis etc.
I have projects that generate native binaries on Linux. Here's an example Makefile [0].
> should I consider taking a look at some description of standard ml instead of rust? It seems like rust has pattern matching and immutability which matches to F# pretty good?
Library and tooling support in F# and Rust is significantly more developed than in Standard ML.
To me the best use case for Standard ML today is in language development.
If its ecosystem were more mature it would be a great choice for application development. But today that's just not the case unless you use a version like Morel that is backed by the JVM (and its ecosystem).
[0] https://github.com/eatonphil/dbcore/blob/master/Makefile#L3
The article mentions that certain SML implementations have seen better support for parallel and concurrent programming than OCaml, but Merlin and ocamlformat for OCaml are great as far IDE tooling go. Which SML implementation has the best tooling for things like go to definition, reveal type at cursor, and format file?
One of the biggest missing things for SML tooling is parsers for SML written in SML. They tend to be written in the implementation language and not exposed as a library. So you don't see SML formatters or documentation generators so much.
There have been a few attempts at this but they haven't seriously caught on.
https://github.com/ProjectSavanna/autoformat/blob/master/run...
Keep your standard library standard.
Module Typeclasses
One, immutable string type. Make it UTF-32 by default (you can always store in a smaller format behind the scenes and extend). Make it multiline. Allow variable interpolation (typeclasses should help).
SML structural typing and anonymous records are better than nominal typing.
Skip on syntax bloat like named parameters.
Overload your math operators instead of having special float operators like Ocaml (once again, typeclasses make this much easier).
Rethink imports. F# requires you to add tons of files in the correct order to an XML. Ocaml imports all the things from the folder and blows up if you use the same module name twice. SML needs to standardize `use`. I'd recommend something like JS static import statements.
If you build good documentation, good compiler errors and good tooling, more people coming to the project and help make a better compiler (a reversal of this is the big issue with SML right now in my opinion).
Have a standard iterator type. Syntax wise, something closer to rust would help newcomers (see ReasonML maybe).
- you want `deriving`
- you might want to look at staging (MetaOCaml)
- you might want to look at resource management (uniqueness typing fe)
- you might want to look at paralellism and maybe you can even release it before Ocaml/multicore arrives ;)It was a really neat idea, and it felt like ML got out of the way and just let me program without having painful syntax or semantics. That said, my biggest ML programs would be five functions spanning ~30 lines loaded into an smlnj interpreter. I never got to deal with functors and modules and who knows what else, and am sorry that I don't know how to "program in the large" in ML. I'm now a professional Clojure dev and would love to see how ML does things.
I own a copy of ML for the Working Programmer (which I see @eatonphil is panning below) and the ML compiler book by Appel, neither of which I've read. Maybe I'll make the time one of these months…it'd be nice to have a Motivating Project though.
Judging from the activity on Stack Exchange, it seems to be pretty dead. Maybe a few universities around the world still teach it? It can't be more than I handful I guess. https://stackoverflow.com/tags/sml/topusers
Also, Standard ML is, well, standardized. The standard isn’t perfect but it at least allows for a strong ecosystem of independent implementations.
MLton is a whole program optimizer (rather than function at a time like Ocaml) and wrings out a ton of performance though compile times are quite a bit longer.
As of 3 years ago it wasn't, but maybe that has changed.
Whole program compilation isn't without its downsides though. It seems pretty common to develop in SML/nj because of fast compiling then doing final performance profiling and deploying with mlton (having a standard helps). Function at a time is also more amenable to caching and reusing part of the previous compile.
I find the `ref` syntax of SML much nicer to use than the mutable record syntax in Ocaml.
SML doesn't have all the dotted math operators for floats.
SML strings are immutable which lends itself to a lot of potential optimizations (and there's always byte arrays if you actually need to mutate).
Ocaml let vs let expressions are annoying to me.
SML has a standard. "The implementation is the spec" in my observation always leads to problems.
My wishlist for SML:
FIX USE COMPATIBILITY. It's intentionally not specified by the standard. This makes portable code very hard. I'd love to see an approach more like JS, but without dynamic imports (useless) and with support for something like GO_PATH plus maybe the ability to recognize and import from URLs (especially .git and .sml).
JS style template strings with interpolation (the type system is at least smart enough to convert primitives to strings) and multi-line capabilities. This would also be an easy backward-compatible way to add support for unicode. Guarantee the ability to implement them as UTF-32 with the knowledge that an advanced compiler could choose to internally represent them as UTF-16 or even Latin1 to save space (JS implementations have optimized to convert their ropes from UCS-2 to latin1 when possible with huge memory savings as even Chinese sites are 90+% ASCII).
Module Typeclasses should keep typeclasses from happening everywhere (looking at you Haskell) while still allowing them to be used for more than equality.
Unofficial first-class functor support needs to be made official.
Pick one of the 4-5 slightly different concurrency models and standardize it.
Things I want that are in SuccessorML:
Guards, "OR" shorthand, optional leading pipe in matches
Line comments
Record punning, extending, and "updating"
do declaration shorthand
Maybe I'm missing something in SML, but OCaml has `ref` as well and it works just like SML's?
let x = ref 0
x := 12
print_int !x
> SML strings are immutable which lends itself to a lot of potential optimizations (and there's always byte arrays if you actually need to mutate).OCaml's strings are also immutable. They have been immutable by default since 2017 and prior to that one could opt into this behavior via a compiler flag.
2017 was just 3 years ago. Loads of Ocaml stuff rely on versions much, much older than that. In any case, the idea of outright changing a formerly mutable structure to an immutable one would be unthinkable in most languages due to all the breakages it is likely to cause.
In OCaml ref is defined as
type nonrec 'a ref = 'a ref = { mutable contents : 'a }
so it is an immutable record which contains one item which is mutable.> In any case, the idea of outright changing a formerly mutable structure to an immutable one would be unthinkable in most languages due to all the breakages it is likely to cause.
I agree with you. I started using OCaml in 2018 but I believe many linux distributions stayed on OCaml 4.05 (the release before safe-string was default) for a while because of breakages.
I was under impression that Haskell's typeclasses plus a couple of most common extensions minus the namespace pollution are pretty much equivalent to ML modules+functors, so could you elaborate how exactly that should work?
Here's some info if you're interested.
http://smlnj-gforge.cs.uchicago.edu/scm/viewvc.php/?root=sml...
A bit unfortunate as recent releases have fixed one of the main for smlnj/reasons to use other compilers in that it now supports x86-64.
Reading the comments here I have zero clue what you're talking about.
Is this what my less technically inclined colleagues feel like when people talk about e.g. machine learning and Python?
Haven't seen any heavy Haskell-type monadology talk in this comment section for example...
At worst, you might not know the name of specific SML compilers/implementations mentioned.
But I don't see anything here that a Python guy with a general grasp of PL wouldn't know...
Still, my answer would be OCaml without questions. It has better tooling, better performance, a larger community, a larger ecosystem, more features and more work being done on it with basically no downsides.