This.
A lot of people seem to be willfully dismissive of this reality when they are fanboying for their favorite languages.
Ask for memory in C++, you'll take a HUGE hit getting it. Ask for memory in Java, and you'll get it lightning fast, but you'll take a huge hit getting rid of it. Either way, you should carefully plan memory usage up front.
Basically, one road leads to death and despair, the other leads to disease and destruction...
pray you choose wisely.
While it is hard to beat existing general purpose allocators in all the scenarios, it is easy for specific use cases.
Firstly, these trading firms probably have tight control over their server architectures. Secondly, once you start using low-level Java idioms in order to optimize your code, you've started approaching the point where differences in performance between JVM implementations across platforms become noticeable.
JIT compilers optimise based on the actual runtime characteristics of the running code, not the best guesses that a static compiler has to make. The effect can be significant. It's no surprise that the Java leaders in Techempower benchmarks match or surpass the C++ ones.
In this case though you would know exactly the target hardware.
The Techempower benchmarks mostly show how good the entrant's web-serving stack is. Java has a couple of decades of multiple skilled, well-established groups competing to build the fastest web stack. C++ has Boost.Beast, which, although its developers are smart, wise, widely respected, sexy, and moderators of a Slack i frequent, can't compete in terms of resources.
There are a few AOT-compiling JVMs out there, and it's not the case that they reliably outperform the JIT-based JVMs like HotSpot.
The significant differences are in Java mandating array bounds checks, runtime checking for null, no undefined behaviour, etc.
Whereas 90% of the time might be spent in 10% of the code image, no 10% chunk of the image can ever be found that contributes to 90% of its bloat.
1% of the image contributes 1% to the bloat, 10% to 10% and so on.
:)
Rust's future is far from certain. I would feel irresponsible using it on a project that wasn't self contained.
What?
Now I'd be happy to see high perf low latency talks about rust, go, ocaml, haskell, sbcl, whatever. I believe limit-cases like these are always very very instructive.
In my gut I would not choose Haskell for my first choice for a low latency language. I would definitely have looked at Rust before Haskell.
I always love optim talks whatever the language (granted its not a clusterfk interpreter).
But on practice, I don't think you are missing anything.
You can pick languages other than C++ or Java, but for a specialized and bespoke hardware/software combination like HFT, you'll be at a disadvantage vs other teams.
In my experience, or at least something I've seen at 3 different firms now, there's also an issue around the kind of people that pick exotic languages in terms of pragmatism. Often times the Scala/Haskell/whatever else side of the codebase is beautiful and slow moving while the hacky C++ side actually works and is agile.
Compared to the typical functional language applications, Rust is different in that it's mainly being battle tested in the kind of Firefox performance issues that are very interesting to an HFT programmer and Rust easily interfaces with C.
I'm not sure if it's still an issue since I don't use Rust daily; but one recommendation I'd make concerning array bounds checking is to do something like electric-fence and allow the compiler to be told to allocate arrays to end at the page boundary such that the hardware interrupts when you go off the edge.
If the compiler restricts such an array to only being indexed with an unsigned integer, then it can be safe but forego bounds checking altogether since the lower bound is always 0 and the upper bound check is hardware accelerated.
Yeah, a lot of those would be faster now, because SIMD was stabilized. That's the source of the discrepancy on n-body, for example. I don't think anyone has bothered to really submit new implementations.
Most bounds checks can be hoisted out of the compiler can’t figure it out on its own via asserts. In general it’s not a huge problem in real-world code.
You made that claim before, and I showed you the Rust n-body program "vector instructions using SIMD" contributed 5 months ago.
https://benchmarksgame-team.pages.debian.net/benchmarksgame/...
Edit: the correct URL is https://benchmarksgame-team.pages.debian.net/benchmarksgame/....
It shows 0.48 seconds, and 2 seconds, but https://benchmarksgame-team.pages.debian.net/benchmarksgame/... shows 13.27. Why isn't that one on the main page?
HN have now corrected the URL, as I asked.
> It shows 0.48 seconds, and 2 seconds
It shows 3 different workloads:
secs N
0.48 500,000
2.01 5,000,000
20.09 50,000,000
> … but … shows 13.27…Which we can see is the time for `n-body #2` at the largest workload:
secs N
0.54 500,000
1.33 5,000,000
13.27 50,000,000
> Why isn't that one on the main page?I don't know which you consider to be the "main page".
We can see that measurements for `n-body #2` are shown on both:
faster/rust.html
and measurements/rust.htmlRewriting some hot path parts of the stack in Verilog is a usual thing, and rewriting some other parts in Coq is the future.
https://www.janestreet.com/tech-talks/ocaml-all-the-way-down...
Java is just someone's C++ program.
> "Pointer arithmetic, placement new, and other low level features make it hard to determine pointers from non-pointers and to move things in memory."
Real applications in a garbage collected language are mixtures of the managed code, plus unmanaged components (like foreign libraries).
Garbage collected C++ can work in much the same way: there is a framework for managed code that uses GC, and then there are unsafe parts that are analogous to foreign code. (Except, the managed code can much more easily and efficiently interact with this code than the typical FFI.)
There is more than one way to integrate garbage collection into C++, in any case.
Not all garbage-collection schemes move things in memory; only copying and/or compacting garbage collectors do.