A better build system for OCaml
blog.janestreet.com
blog.janestreet.com
> I must not segfault. Uncertainty is the mind-killer. Exceptions are the little-death that brings total obliteration. I will fully express my cases. Execution will pass over me and through me. And when it has gone past, I will unwind the stack along its path. Where the cases are handled there will be nothing. Only I will remain.
¥https://signalsandthreads.com/build-systems/
(Hard to present glyphs that would actually be used in print - I've chosen what I think is close - but they are not that close because wrong font)
I still occasionally hear things about how the more academic-styled functional languages can't work in production, but Ocaml shows that it absolutely can work, even with high performance requirements.
Mercury also uses Haskell for their backend https://mercury.com/
I was fortunate enough to work at Jet.com before Walmart completely destroyed it, which was an F# shop. I really liked it, and I never really felt "limited" by it.
The stuff I was working on didn't have nearly the same requirements as high-frequency trading like Jane Street does though. I never did any super low-latency stuff with F#, so it's tough for me to say how well it fair with that kind of environment.
Would honestly be a lot more interesting than Haskell.
It feels like the trend right now is to bolt on one or two "functional libraries" into your "normal" language and pretend that that's the same as writing Haskell or Ocaml. People have actually expressed such sentiments to me because Java has the optional type and a "map" function for the Streams API. When I suggest writing something in a functional language, the response is always "it's too hard" or "we won't be able to hire for that", as if engineers are somehow unable to learn new things.
But I think what OP meant was more about the "functional programming" side of things than the "HM-typed" side of things. Naively, anyway, you might think that "the FP-style" of avoiding mutation and preferring recursion would require lots of garbage and high-latency garbage collection, copying, function call overhead...of course, that's not the whole story, but having Jane Street to point to as a crushing counter-example is nice.
I think it's purity that actually is a big difference, but OCaml isn't pure.
It complicates things that are simple to express in an imperative language
Many jobs in finance are updating 20 year old Java code, or figuring out new ways to load data in and out of Excel files for custom reporting.
I wonder if they disable all the fancy exploit mitigation protection in linux kernel just for a tiny performance hit
go fish
Definitely doesn't, you can just slap unsafe and manipulate raw pointers if that's what you want
Yes, so it means GP is right! Modern capitalism in tech is about rewarding the two aforementioned tasks.
Because endless growth is the only reason these once fun spaces have been hyper focused to be as addictive and stressful as possible to the "whales" of scrolling. That's why, when their own internal reports say "people spend unhealthy time on our platform and it's making them unhappy," it gets passed up the chain of command and whittled down by internal incentives until it dies as an issue.
Individuals hold some blame, but to put most of it on them is to ignore what growth demands. You're supposed to doomscroll and engage and worry. That's the business model. Facebook is in the same business as Cigarettes and Casinos. When I see someone on an air tank playing slots, literally crying when they spend their last dollar, I will not waste my breath blaming them. Just like I won't blame the doomscroller, anxious that they need to stay "informed," who hasn't met the basic needs in their own life.
> The minds that want to scroll are the same as the ones that make money on the scrolling
No? Where are you getting this? I don't think the people guiding these companies want to spend 6 hours scrolling TikTok. This is not the way most people live their lives.
I meant they are all humans. We're all sort of the same. If you disagree just think of your opinion of any other species, or about a group of people a thousand years ago, and you see what I meant.
No, sponsor contracts, advertisements.
So the story here is that over the last twenty years they stole the lunch from the traditional market makers like eg banks.
Of course, they got rich in the process. But they started from relatively modest means, compared to the companies they took on.
Michael Lewis's 'Flash Boys' is an hilarious account of this process. Well, it's involuntarily hilarious, because to tell his story, Lewis needs to cast Goldman Sachs (!) and other big banks as the victim. See the rebuttal 'Flash Boys: Not so fast' by Peter Kovac for more insight.
Obviously, the people who own the automation will want a cut of the rewards, like any other business.
Obviously NASDAQ and electronic trading systems are a good innovation. But firms basically doing arbitrage or exploiting uneven network latency are not that economically productive.
> Inefficient market spreads and network latency is not worth remediating.
Well lowering market spreads is all about increasing the returns for capital, and incenctivising overfinancialisation. It's hardly curing cancer is it?
At worst it's actively harmful if you believe that the current state of turbo-financialised capitalism has its drawbacks.
> Network latency
Not really sure what you're talking about but surely spending billions of dollars to bring rtt latencies to 50 micros or whatever is not really a great use of money and top engineering talent. Again, it's playing an arbitrage game but not really delivering any value.
I want liquidity, low spreads, price discovery. You seem to forget that “not delivering any value” is just like y’know according to you…
EDIT: The funny part is even the exchanges and hft firms agree with me see PLP/speed bumps on exchanges like Eurex lol
I honestly can't tell what your concrete points are. I come from the position that economies are naturally occurring phenomena which cannot be centrally planned or controlled. If people can find ways to profit off market inefficiencies, they should! The HFT/Quant firms make their arbitrage money (value for them) and all market participants in return see: (non-exhaustive list)
1. Better price discovery 2. Tighter spreads 3. Higher liquidity
Which is value for everyone else.
If your bar is that "all smart people should be working on curing cancer or andrepd-approved endevours" then almost nobody in the economy is providing value. Is my lowly SecEng job at $MEGACORP good enough? What about my buddy who writes firmware for toothbruhes? Are professional starcraft players wasting their talents?
> EDIT: The funny part is even the exchanges and hft firms agree with me see PLP/speed bumps on exchanges like Eurex lol
This debate has been going on for ages, and it's silly to pretend that it's been settled and everyone agrees with you.
This is a challenge to untangle. It sounds like you're saying that there is no point trying to regulate, legislate or control what happens in the economy at all. But that sounds bonkers to me.
For starters, there are (and should definitely remain) absolute limits to business activities. We've moved on from Victorian-era child and slave labour for good reasons, even though such a situation was "naturally occurring" at the time. Moreover economic activity is dictated by cultural mores - if your service is morally reprehensible in some way then you won't get much business whatever your matgins are. Economies are inherently subject to the laws and customs of the agents.
Secondly, some regulation is pretty clearly beneficial. For example, there's a recurrent tendency for market power to concentrate in modern economies; we need robust anti-trust regulation to prevent consumers from getting ripped off and to prevent fragile supply chains. A well-conisdered balance of public and private provision supports the least well-off in society while allowing room for the fruits of individual flourishing.
Thirdly, we must consider what makes one economic system better than others. One way to measure this is to look at how efficiently it converts resources to social utility. I'm far from convinced that it's efficient to employ our brightest minds to build trading models with brief lifespans so that investors who are already well-off become slightly more so. It's worth investigating what regulations and incentives could put those minds towards things of greater value - solving climate change, cancer, sending humans into space etc... .
I really do not appreciate this mischaracterization of my position. Focus on my actual words. I don't care about 'winning' this online argument. I take effort to engage because I am disturbed by the number of intelligent people who believe if only _THEY_ were in charge (or at least the right person), we would be able to fix all of society's problems.
> For starters, there are (and should definitely remain) absolute limits to business activities...
I agree with everything that follows. Government needs to be around to keep the peace. I want to be explicit: When I say "centrally planned/controlled economies" I am NOT talking about the general concept of regulation. If you are debating in good faith, this should be obvious. Look at all the history of failed states who tried to implement top-down control of their economies.
Also, YSK that not all regulators are government entities.
> Thirdly, we must consider what makes one economic system better than others. One way to measure this is to look at how efficiently it converts resources to social utility.
Never before in history has mankind been so prosperous. What system would you like to emulate? The US capitalist system is not perfect (and never will be)...but it blows all of its peers out of the water in terms of economic prosperity. Here's a couple data points: (Please read the technical definitions if you are truly interested in this subject)
- https://en.wikipedia.org/wiki/Disposable_household_and_per_c...
- https://www.numbeo.com/property-investment/rankings_by_count...
> I'm far from convinced that it's efficient to employ our brightest minds to build trading models...
This is where my "closeted dictator" quip comes from. Nobody is "allocating" these minds...they are acting on their own free will. Why should you or anyone else be the arbiter? What if individuals disagree with your beliefs? Space exploration is a great example of a debatable "worthy endeavor"
A lot of jobs are extremely mundane though, compliance, regulations, legacy code bases, etc.
Yes, all mitigations get disabled
What's the threat model where your HFT application's running hostile / untrusted code?
I'd like this stuff enabled on my desktop because I'm not sure what hideous javascript is being dumped into my browser by some advertising network.
But my trading platform? If the badguys are able to execute these attacks there, it's because they've got full access already.
Processes reading and writing directly to FPGA/NIC ring buffers.
Shunning TCP in favour of UDP based protocols that are easy to optimize for your particular usecas in userspace.
Removing cores from the Linux scheduler entirely and pinning processes to those cores.
This stuff isn't even novel, it's been standard practice for a couple of decades.
"How do you live with yourself making high 6 fig / 7 figs a yr?"
Quite easily in fact.
Work on real problems. Try to make real people's lives better and happier. There are real problems in finance but my feeling was it's all very simple and solved decades ago, now it's just pointless complexity that isn't solving anyone's problems. I recommend John Kay's Other People's Money for a primer on what finance is actually good for and where it's gone wrong.
The real big problem in finance IMO is digital cash. Bitcoin started out trying to solve that problem, and there are still some people in the community interested in it, but it's mostly of interest to the finance guys now. Just another "instrument" in their "portfolios".
If the last few years have taught me anything it's that a large % of the population will actively aim to make their own lives worse long term because they are told lies. What benefit is there really in trying to undo their own self-inflicted damage.
Let's see if they last at least a couple of centuries.
Please tell me a better system that would cure cancer. I don’t think it’s Leninism or its derivatives. You need a macro system that is rich enough to allow for significant investment in medical research.
He also claims they're full of elitists from top universities and are not receptive to ideas outside that bubble.
I still rank it above making people click on ads though.
Also, it's apparently significantly harder to land a position at JS than at a Google/Meta.
Actually, I applied there a while ago, the interviewer was actually pretty unpleasant, which hasn't happened to me at big tech. Didn't leave a really good impression.
It was probably the first and only time I would have rather have been ghosted lol
There are a ton more people working in tech in finance that don’t quite have it as fun (or lucrative) as Jane Street let alone your average tech company.
"go learn this awful new language"
I don’t mean that as an insult to Lua’s creators. They seem like really smart fellows. It’s just that the language is (to my eyes, with my background) viciously ugly. And 1-based arrays, of course, are evil.
It has some neat ideas, though, and it is supposed to be very easy to integrate into a project. But man, that syntax …
>> With nowhere to go, it has to roam everywhere in your system, both in your code and in people's heads. And as people shift around and leave, our understanding of it erodes.
>> Complexity has to live somewhere. If you embrace it, give it the place it deserves, design your system and organisation knowing it exists, and focus on adapting, it might just become a strength.
- Fred Hebert, https://ferd.ca/complexity-has-to-live-somewhere.html
And even when the complexity is essential, IMO it's better off not in the build system. I'll gladly accept more complex code for the sake of a simpler build (even though that theoretically means worse performance). Worst case if I need to do something complex at build time I'd rather model that as "the build system invokes a program that does something complex" than try to express the complex thing in some Turing Tarpit "configuration" language.
The trouble is when people say "complex" you don't really know what they mean, though. They often just mean "difficult". Every programmer who wants to use that word needs to watch this: https://www.youtube.com/watch?v=SxdOUGdseq4
I strongly disagree. Most software is insufficiently complex to adequately represent reality.
That may be so; what I'm claiming is that most of the complexity in software as it currently exists is accidental.
Very respectfully, I think you may be missing the author's point. When you fail to make a home for necessary complexity, it rears its head as unintended complexity in unexpected parts of the system. The source of 'accidental' complexity is unaccounted for complexity.
If that's what they're claiming then I completely disagree. No, that's not the reason, that's got nothing to do with it. If that were true we would expect e.g. projects with more complicated builds to have simpler code, and IME that's not true.
Instead, each domain space has some degree of inherent complexity, which varies from problem to problem. Failing to account for this inherent domain complexity appropriately will cause it to bubble through at unexpected points throughout the system.
Build systems inherently have a very complex job. A good build system grapples with this complexity and tries to harness it; a bad one pretends it isn't there, and becomes a tangled mess once the (inevitably complex) demands made of it exceed its limited assumptions.
I don't think this is true. I think that when looked at in the right way the job of a build system (when used appropriately) is actually fairly simple, and most build system complexity is either accidental complexity (either just straight-up bad design, or misguidedly overengineered flexibility in directions that don't matter) or comes from trying to accommodate things that the build system shouldn't have been doing in the first place. When I've seen overcomplicated builds they've never been because the build system made assumptions that were too limiting.
I am painfully aware of bitbake. I’ve probably written 3-400 recipes.
Most of them are about 20-30 lines long, because I refused to hide the compilation mess inside a recipe. I fixed the problem _before_ getting to the bitbake part. Most of my recipes at this point need only a repo name, the recipes are identical after that.
I'll eat at least a bit of shit if it means I can get more than one platform's-worth of build process out of a single set of human-editable configuration files.
The world does not work like this though. CMake is weird but once you've learned some non-intuitive stuff, it works very well and there's a reason why pretty much everyone is using it.
(And two minor annoyances that people determined to hate CMake will never shut up about. But then, if you're determined to hate something, there are worse things than CMake to do it to! So I can't be too critical.)
Sounds like the perfect match for C++.
Ocaml never clicked for me, I have a rare form of semicolon allergy and Haskell just looked a lot nicer to me.
But then I recently tried Reason and enjoyed it A LOT, so everything Ocaml is suddenly interesting.
https://reasonml.github.io/ looks cool, OCaml with javascript.
The Javascript-oriented part of ReasonML got forked to be its own language: Rescript. https://rescript-lang.org/
I'm sure once you get the zen of it, it's fine. Like Lisp, I guess, you learn to think in its structure. But looking at a screenful of Haskell to me is intimidating.
Of the bunch I found SML/NJ to be the most readable.
But in any case, you can use curly braces and semicolons in Haskell just fine. You can also write your Haskell like Lisp, and add lots of parens everywhere, and use all operators in prefix-form.
https://learn.microsoft.com/en-us/dotnet/fsharp/language-ref...
I took a brief look at these things, and my impression is that their stuff isn't "ready" for anyone outside Jane Street, even though they put a lot of effort in building the ecosystem and open source their code.
I've used some of their other libraries too, their logging and unit test ppx are common maybe even de facto standards as much as the ocaml world has such a thing. I've also used, off the top of my head, their code formatter, one of their test frameworks, their implementations of some advanced data structures.
Sometimes you do run into one like the other commenter said, where that shit just does not work. It depends on an undocumented something they shipped separately, or needs a secret bit of config or whatever. These aren't malicious, I open a ticket and come back in a year or two often they'll be working.
It's not zero frustration but I appreciate their approach of just throwing everything over rather than spending more resources testing and polishing fewer releases. Their code quality is generally very high and even if I can't get something working directly, it provides a rigorous & vetted example implementation.
I love Lean 4, but good luck getting help with it from AI. Today's project-in-progress is digesting their reference manual to fit well within a 200K context window. We'll see if that helps.
How is the design of the APIs? How stable are they?
Does Jane Street respond to bug reports/pull requests (if any) quickly?
I worked on Bloomberg DLIB which is basically an implementation of https://www.cs.tufts.edu/~nr/cs257/archive/simon-peyton-jone...
JS does have a bit of a NIH culture, but I'm not sure if that was really at play here. There just...weren't very many good build tools available at the time, particularly for a company using an unorthodox tech stack.
But Dune started (according to this blog post) in 2016 and JS started seriously improving and adopting it last year. So to me Jenga sounds like a reasonable step in 2012, but pouring significant effort into migrating from Jenga to Dune (and improving Dune) in 2024 sounds more weird
Dune is a rename of Jbuilder (2016). Jbuilder uses Jenga (2012) configuration files.
> By 2016 we had had enough of this, and decided to make a simple cross-platform tool, called Jbuilder, that would allow external users to build our code without having to adopt Jenga in full, and would release us from the obligation of rewriting our builds in OCamlbuild [...] Jbuilder understood the jbuild files that Jenga used for build configuration.
So in 2012 it made sense for them to build Jenga, because there weren't any good alternatives - Bzl etc. didn't exist, so they couldn't have solved their problems.
And in 2016 they had open-source code they wanted others to be able build; those people didn't want to use Jenga, and JS didn't want to rewrite their builds so that they could use something else. Thus, Jbuilder was a shim so that JS could still use their Jenga builds and others could build JS' code without using Jenga. Bzl etc., even though they existed, wouldn't have solved these problems either.
In particular I was under the impression one needed to be able to run ocamldep before hand (or compile twice) - buck2 can do this, bazel needs hacks iirc.
The major pain point is LSP integration, which has to be closely tied to the build system, since it's only by building that the LSP server can know a file's dependencies. Everything is all neatly available with dune. We've cobbled something together a bit with buck2 but it's not as nice.
"He who controls the [build system] controls the universe."
The lsp requires you to run "dune build" first, bad already.
If you add a new file, the lsp wont pick it up until you dune build it again.
The compiler errors arent there too.
But i loved writing OCaml, its just thats a bit more painful to learn than due to the tooling, since i didn't use many functional langs before.
Cmake mostly
(This rant more or less equally applies to other language-specific build systems.)
But I'm a new OCaml user, and actively starter using ocamlbuild because dune added layers of indirection really tripped me up at first.
What's a 'familiar Linux build system'? make?
Ironically many modern C/C++ projects use Cmake to generate Makefiles. If anything the inverse of your observation is mine.
Because if they're still using the language-specific build tools and dependency management systems, then I think you would find that the Fedora maintainer higher in this thread would not be any happier that there is a sugar coating of Make. That's not what they're asking for, based on other rants I've seen from Linux distro maintainers.
The more complex ones at $JOB actually do some caching, dependency management, code generation, and compilation.
Rust https://github.com/rust-lang/rust/blob/master/tests/run-make...
Go https://github.com/golang/go/blob/master/src/runtime/Makefil...
NodeJS https://github.com/nodejs/node/blob/main/Makefile
OCaml https://github.com/ocaml/ocaml/blob/trunk/Makefile
Python https://github.com/python/cpython/blob/main/Makefile.pre.in
Haskell https://github.com/ghc/ghc/blob/master/ghc/Makefile
R https://github.com/wch/r-source/blob/trunk/Makefile.in
CLR https://github.com/mono/mono/blob/main/Makefile.am
Ruby https://github.com/ruby/ruby/blob/master/enc/Makefile.in
Downstream projects in these languages do not typically use Make.
More to the point, I clicked on the Go one, and it's just including this tiny "Make.dist" file that does nothing except invoke "go tool": https://github.com/golang/go/blob/master/src/Make.dist
Wow. So useful.
I clicked on the Rust one, and not only did it seem to be specific to some old testing infrastructure, but I found this note:
> There are two kinds of run-make tests:
> The new rmake.rs version: this allows run-make tests to be written in Rust (with rmake.rs as the main test file).
> The legacy Makefile version: this is what run-make tests were written with before support for rmake.rs was introduced.
So, it's an obsolete system that is being migrated away from.
But, again, the main point is that what the toolchain does with its free time has little to do with how end user applications are developed, and the complaints in this thread were strictly about building applications in distros, not about building toolchains.
If an application in one of these languages uses make, it is typically just a little syntax sugar around the toolchain commands, which does absolutely nothing to absolve the project of the complaints Linux distro maintainers have about how dependencies are managed.
What people are claiming is that make is used as a build system for projects whose source code is written in C or C++.
"Familiar" is not a property of any system. It's a relation between a system and a user.
Some Linux build systems maybe be more familiar to some users, but will be less familiar to others. When picking a build system, you can't just look at the system itself and declare it familiar or not.
It's not even enough to look at the total number of users familiar with some thing. Hindi is one of the most familiar languages in the world, but you're probably gonna have a bad time if you use it for the menu in a cafe in rural Texas.
You have to look at your actual cohort of users (and potential future users) and see what's familiar to them. This is one of the key reasons why usability is actually a deeply hard problem. So much of usability hinges on familiarity, but familiarity is a human-specific highly variable property.
(this rant more or less equally applies to all os-or-distro-specific build systems)
---
This rant is only semi-serious. I do see some value in the Linux distribution style packaging. In particular, I do sympathise with the need to do cross-language builds. But goodness are they a pain to work with, and probably the biggest barrier to me shipping software on Linux.
My hope is that eventually an evolution on build systems like bazel/buck2 will lead to a truly universal build system that is both cross-platform and cross-language. But unfortunately it doesn't look like it's coming soon.
Don’t[1]. Ship source tarballs (or VCS tags).
I’ll grant that most distros’ build systems are antiquated and, in places, silly. (That includes Nixpkgs, first released 2006.) We could really use some fresh ideas there.
But they’re also not for you (or me) in your (or my) capacity as a software author. They’re there for a person who works on packaging software for, most of the time, a single distro, and their balance of complexity and flexibility is calibrated accordingly. One of the functions of that person is also to keep you honest and represent the interests of users before you, because they have more expertise than the users but not as much of an attachment to your software as you. The ecosystem is less healthy when the author tries to fill in for the packager.
“But then the users will come to my bugtracker to complain about bugs in patched versions!” Pre-Google, we used to have a solution for that: a configure option to set the bug reporting email, present in all GNU software. Nowadays it’s not clear what a good solution could be, but it does seem like, unfortunately, the author will have to maintain a table of packager contact information for the end users.
[1] https://drewdevault.com/2021/09/27/Let-distros-do-their-job....
Less popular apps or closed source apps don't work. Things with old dependencies won't work. Other language with package manager which depend on dependencies and dependency versions that may not be in distro packages may have trouble.
For the core the distro model might work. For the rest maybe something like flatpack from the devs might scale
And yet those language-specific build systems are overwhelmingly winning, in pretty much every language.
> As Fedora packager for OCaml packages,...
I honestly think traditional Linux packaging is in the wrong here and the problems are essentially self-inflicted (not in the sense that individual maintainers are doing something wrong, but in the sense that the policy that traditional Linux distributions are following is inherently unsustainable. It's designed for a pre-CPAN world)
> As an upstream OCaml developer, the whole thing falls down the minute you need to integrate other programming languages into your build (or OCaml code into a code base written in another language).
True up to a point, but frankly the worst case is falling back to a terrible C-style build, and "always do terrible C-style builds in case you need to integrate with C code" is not a proposition that has much appeal.
Much as I wish the whole world would standardise on Maven or Cargo, I can't see a realistic path to there without first eliminating C, because the C people are never going to agree to follow a standard for package repositories.
The issue is there. Not you or your fellow packagers of course, but in the idea that every linux distribution needs it's own packaging system, and each version of a distribution it's own packages.