Java is better than C++ for high speed trading systems
news.efinancialcareers.com
news.efinancialcareers.com
As it happens, I have developed one algorithmic, low latency trading system in Common Lisp / ANSI C, and then was asked to rewrite it in Java which I did.
It actually traded on Warsaw Stock Exchange and was certified by WSE and was connected directly to it (no intervening software).
Yes, it is possible to do really low latency in Java. My experience is my optimized Java code is about 20-30% slower than what very optimized C code which is absolutely awesome. We are talking real research project funded by one of larger brokerage houses and the application was expected to respond to market signals within single digit microseconds every single time.
The issue with Java isn't that you can't write low latency code. The issue is that you are left almost completely without tools as you can't use almost anything from Java standard library and left scratching your head how to solve even simplest problem like managing memory without something coming and suddenly creating a lot of latency where you don't want it to happen.
Can't use interfaces, can't use collections, can't use exceptions, etc.
You can forget about freely allocating objects, except for very short lived objects you must relegate yourself to managing preallocated pools of objects.
I don't know what is the current state of garbage collection but in our case garbage collection had to be turned off and nothing outside of Eden could be collected during program execution (Java 8).
Writing that kind of optimized code is all about control and Java is all about abstracting it so you don't have to worry. As you see, both goals are at odds.
So I still prefer working low latency in C as it is more natural, native solution for managing problem where you absolutely need to control memory layouts, preallocated object pools, NUMA, compilation of decision trees to machine code, prefetcher, etc.
I have about 10 years combined experience working with C and 15 years working with Java.
For example, I had piece of Common Lisp code optimize and compile decision trees to machine code. These were used, for example, to check whether the trade is allowed to go to market. This had to be very optimized as it was running in parallel to constructing the message and had to deliver verdict before the message was actually created to have a chance at actually stopping anything from reaching exchange. This was my solution to have order verification in essentially zero time (the cost was single branch instruction that would only be mispredicted if the verdict was to stop the order).
Another piece of Common Lisp generated machine code to parse XDP messages coming from exchange. The spec for these was in XML, Common Lisp macros created extremely optimized machine code that used absolute minimum instructions to parse out the information we were interested in without having to write all that code manually. I actually used concept from Practical Common Lisp book for this (http://www.gigamonkeys.com/book/practical-parsing-binary-fil...). Thanks, Peter Seibel.
The entire application was essentially a Common Lisp application that used FFI to call a library with optimized ANSI C code. Some Common Lisp code was running in parallel to constantly update things like decision trees that required regular recompiling based on market situation but most was really just to setup everything and give control to the C part.
The application did not talk to operating system (Linux) after it started (no syscalls at all) and used full kernel bypass to talk to NIC.
That's really interesting... How does it work? How can a user space program talk to the NIC directly without using system calls? Shared memory?
Realtime means processing of at least some of the events and tasks in your application must be completed before a set deadline. For example, coin acceptor that detects what kind of coin you dropped in has x milliseconds to figure out if the coin is to be accepted or not. If it can't do it within the time you get unhappy customer.
Realtime systems don't have to be very performant, but they don't have something like 99th percentile latency shooting up. All events must be processed within deadline.
Low latency systems are soft realtime PLUS performance.
In a low latency trading system you have to respond to market within a deadline to have a chance to do something useful but if you don't normally it doesn't mean lives are lost, metal is bent, etc.
Nowadays I rather use Java and .NET, which have been picking up on the features that should have been there at v1.0, like the ones you mentioned, but hey better later than never. :)
And on the few cases where it isn't enough, I can code a tiny library in C or C++, no need to throw everything away.
The only situation I can think of where Java would be right fit is where you have an application with a lot of logic and just a small part that needs to be highly optimized.
https://www.ptc.com/en/blogs/plm/ptc-perc-virtual-machine-te...
Usually playing with numbers in these kind of deployments isn't part of the package.
Real time doesn't (necessarily) mean fast, it means the latency is bounded.
https://en.wikipedia.org/wiki/Real-time_computing
Real-time has nothing to do with performance and everything to do with predictability and deadlines.
People are put off when they find out that real time versions of Java or Linux are actually slower than the regular ones. Until you think about it a little bit and if they were faster there would not be multiple versions.
Real-time versions of products are created because sometimes performance is achieved by means of stochastic processes or amortizing costs and you need to remove these optimizations to be able to support deadlines.
Real-time programming emphasizes ability to calculate how long an operation will take as more valuable than raw performance.
I think best hypothetical development environment would allow you to specify the tradeoff and choose whether you want strong typing, fully managed memory or ownership semantics.
Out of curiosity, what was the rationale for writing the system in Java? Were the tradeoffs not obvious until too late, was the choice of language dictated by higher-ups, etc.?
If you choose Java for your project it usually means you don't care a lot about latency and performance (you care some, but not all that much compared to, for example, linux kernel development). You care more about managing larger codebase, having easy access to whole lot of frameworks and libraries, good development tools and easy development / deployment process.
Developers who have experience with Java are inevitably people who have mostly had experience with those kinds of projects of which I would say most are corporate backend systems or Android applications. It really is difficult to find any other projects.
On the other C is inflicted pain of having much harder development environment for the size of the project. You only use low level language like C when you have a specific need. Maybe you need to develop Set Top Box software, or video codec, or some device controller, or high performance quant library, that kind of stuff.
So it isn't really fault of Java developers to not know the stuff. It is natural result of Java being suited to a different type of projects than C is.
We also had custom Intel CPUs that are auctioned individually. And cutthrough switches that stream the IP packet immediately when they receive the IP header, before they receive the payload. I have calculated that if the message from exchange was a beam of light it would be delayed by less than a meter.
That's just latency to your NIC, right? 1 meter ~ 3 ns ~ 9 CPU cycles at 3 GHz, less than the integer pipeline depth last I checked.
Decade ago.
May as well just use C and never free your memory. Lots of memory management issues go away if you never free memory.
I have some sympathy for the idea that the JVM is better since it means you won’t spend all your time chasing crash reports. The thing is, like another comment hinted at, is that this issue is generally more reflective of the environment you build in than the technology choice.
Here’s a good talk on the reasons for and limits of using C++ for low latency systems: https://m.youtube.com/watch?v=NH1Tta7purM
If I had to rank in terms of reliability of trading infrastructure I’ve used/worked on,
* Low Latency C / C++ infra at fully automated trading firm. Very much run by the programmers and quants and also fairly small. By far the fastest (near limits of what you could do) and also most reliable.
* Low Latency Java execution infrastructure. Pretty reliable, not that fast, had some issues with GC battling and manual memory management to avoid GC, etc. There was a pretty clear latency floor (still quite low) even when “doing everything right” that serious native infrastructure beat.
* C++ market making infra at a firm run by manual traders. It was by far the slowest and least reliable. Echos the experience of “spent hours debugging weird crashes”. The culture was very “A trader asked for this and needs it done yesterday. Also this refactoring business doesn’t sound like adding new features, drop it”.
What I saw is that if you hire people who have a good idea of what they’re doing and keep a culture of technical excellence, C++ is definitely a better choice if you care about latency. You really have to maintain a culture of high quality code, testing, and in general caring about the technology.
This is only really possible when all of the stakeholders are involved in the technology, or at least understand the benefits. Once you start down the path of “well this feature could be done a day quicker if...” this goes down the drain, and a few years later you find yourself getting run over on latency AND with an impossible to use trading system. It’s really the worst of both worlds.
I'm still wondering if the businesses would work better if the focus on technical excellency could be instilled and especially which parts of the company would have to suffer for it and what the ultimate outcome would be.
I'm convinced that at least for some companies, it simply can't be the winning strategy.
Trading systems trigger almost all their orders based on detecting market changes, which you cannot predict. If a symbol hits a price that triggers one of your trades you want that trade in the market ASAP every millisecond matters. Randomly add 50ms on to that and your trading system is out of business. The customers will go elsewhere.
Typical pauses may be far less than a millisecond in these GCs because they do almost everything in parallel.
At the super fast speeds you start running into things like:
* why won’t the devirtualizer trigger?
* these object headers are sure wasting a lot of cache
* there’s a lot of forced pointer indirection
And you just end up spending vast amounts of time and effort trying to shave off those few microseconds you’re wasting in the JVM. Once you add in all the effort trying to get around the GC, it’s bleak picture for competing with quality C++ without spending far more effort.
The answer to both questions is that many people prefer writing in higher-level languages with more safety guarantees.
Yes, of course $HIGH_LEVEL_LANG is preferable for many many use cases. In this context, we're discussing "high speed trading systems", for which native implementations are going to be the favorite.
...but the point of the article is they aren't always the favourite.
So they use regular Java for the build system, deployment, testing, logging, loading configuration, debugging, etc etc. There's a small core written in this strange way... but everything else is easier. And things like your profiler and debugger still work on the core as well.
Depending on the goal, one has to be very careful when choosing what one uses. The good side is, C++ kept its C features: if I'm deciding how I'll do something, I don't have to follow the rules of the "language lawyers." I can do my work producing what is measurably efficient. And compiler can still help me avoiding some types of errors -- others can anyway be discovered only with testing (and additional tools). At the end, knowing good what one wants is the most important aspect of the whole endeavor.
http://www.jot.fm/issues/issue_2003_01/column1/
https://wiki.c2.com/?UmlCaseVultures
One size doesn't fit all. Some solutions to some problems could be and are provably better than those typically promoted or "generally known" at some point of time.
If that all still doesn't mean anything to you, please read carefully and very, very slowly "The Summer of 1960", seen on HN some 9 years ago:
https://news.ycombinator.com/item?id=2856567
Edit: Answering the parallel post writing "When you proclaim to ignore language lawyers, it sounds like you are knowingly breaking the rules of the C++ standard."
No. The language lawyers, in my perception, religiously follow everything that enters the standard and proclaim that all that has to be used, because it's standardized. Including the standard libraries and some specific stuff there that isn't optimal for the problem I'm trying to solve. And especially that whatever is newer and more recently entered the standard is automatically better. It's understandable that they support their own existence by doing all that -- it's about becoming "more important" just by following/promoting some book or some obligatory rituals (it's an easy and time proven strategy through the centuries, and that's why I call it "religiously", of course -- and I am also not surprised that somebody who identifies themselves with being one of the "lawyers" wouldn't like this perspective -- you are free to suggest a better name). But it should also be also very obvious that it's not what's necessarily optimal for me to follow, as soon as I can decide what I'm doing. And yes, it's different in the environment where the "company policy" is sacred. There one has the company "policy lawyers", and typically every attempt of change can die if one isn't one of them.
"Orthodox C++"
E.g. at the whole bottom of the page in some comment is a link to:
"Why should I have written ZeroMQ in C, not C++ (part II)"
where the author writes "Let's compare how a C++ programmer would implement a list of objects..." and then "The real reason why any C++ programmer won't design the list in the C way is that the design breaks the encapsulation principle" etc.
I have indeed more than once used intrusive data structures in non-trivial C++ code, and the result was easy to read, and very efficient. There I really didn't care about some "thou shalt not" "breaking the encapsulation principle" because whoever thinks at that level of "verbot" 100% of times is just wrong.
The "encapsulation principle" is an OK principle for some levels of abstraction, but nobody says that one has to hold to it religiously (exactly my point before). I would of course always make an API which would hide what's behind it. But where I implement something ("the guts" of something), I of course have my freedom to use intrusive elements, if that solves the problem better. I have even created some "extremely intrusive" stuff (with a variable number of intrusive links in the structures). It worked perfectly. Insisting on doing anything as "ritual" all the time is just so, so wrong.
I'm not exactly sure what specifically you are trying to imply here. When you proclaim to ignore language lawyers, it sounds like you are knowingly breaking the rules of the C++ standard. That takes a lot of faith in compilers doing what you meant to do, despite writing code that is incompatible with the standard those compilers implement...
In the JVM I think this can only be done speculatively (you have to double check the type), but it still matters.
In the JVM devirtualising means making a virtual object a full object again, so the opposite of what you want to be happening.
I don't think the JVM really names the optimisation you're talking about, but it does do it, through either a global assumption based on the class hierarchy, or an inline cache with a local guard.
When we say devirtualize we mean that we know animal is always a cat so we don't have to look up which speak to use.
And that makes literally no sense in-context, you're mapping those words directly to the concept of virtual methods but that's the opposite of the way chrisseaton uses them (and they're clearly familiar with the concept of virtual calls and it's not what they're using "virtual" for), hence asking them what they mean specifically by this terminology.
Yes... but let's not say 'broken up on the stack' - the object's fields become dataflow edges. The object doesn't exist reifed on the stack - fields may exist on the stack, or in registers, or not at all.
Again, thanks for the response, this was insightful and I'd love to learn more.
Also, an aside - do you have any thoughts on the use of Rust in these same systems? It's a little bit more bleeding edge, but I'm curious to hear an expert's thoughts!
Sometimes an arena may instead refer to a region of memory for allocations that are all the same size/type. Because everything is the same size, the data structures for tracking allocations may be simpler.
In a language like Java, you basically preallocate an array of objects of the same type with a bunch of fields set to zero, and then have some data structure to give you a fast interface a bit like malloc. Because you don’t do any extra allocation, the GC can be safely turned off.
People who don't like this style (or the code) leave.
The direction for quality has to come from the top and be supported by systematic mechanisms. In fact that would be a question to ask in interviews.
"No-one ever explicitly specified that the bridge shouldn't fall down and kill everyone who was on it at the time", said no engineer, ever.
A particular one that caught my eye: Cimarron River Rail Crossing Dover, Oklahoma Territory United States 18 September 1906 Wooden railroad trestle Washed out under pressure from debris during high water 4-100+ killed Entire span lost; rebuilt Bridge was to be temporary, but replacement was delayed for financial reasons.[8][9][10] Number of deaths is uncertain; estimates range from 4 to over 100.[11]
Emphasis mine. How is that different from software engineering?
As the risk of nitpicking, we already can produce reliable code, but as piaste's comment states, it's rare to make the necessary investment.
Avionics software and medical systems software is held to a very high standard, at great expense, but the average website or desktop application is not.
> Webpages will be 100GB of JavaScript though
I agree that the problem of bloat is likely to continue to worsen in the web, but that's just the web. Other sectors, like high-performance game engines, will continue to compete on efficiency.
Rather, because software engineering hasn't killed enough people to advance; or when it has, it wasn't obviously the culprit.
In which fields of software engineering do we find actual solid procedures and standards? Avionics. Medical hardware. Fields where the link between bug and death is as short as possible, and so the stakeholders have demanded solid engineering.
What are the consequences of GP's acquaintance writing poor code? Some trading firm becomes slightly less efficient at trading. Impossible to evaluate the net human loss from it (if it's a loss at all).
The civil engineering equivalent would be something like façade design or layout. If your building is ugly or confusing, it will annoy or waste the time of the people who live and work in it, but as long as it doesn't fall on their heads nobody is going to withdraw your certification.
I think the undertone is usually that the software developer at fault had no idea what they were doing, and had no place working on critical systems, rather than that they had ill intent.
and where it had, you get pretty serious control and certification process put in place to avoid it (Boeing notwithstanding)
The MCAS software was modified to read from both angle of attack sensors and to be less aggressive in pushing the nose of the plane down. The software that controlled the indicator light that illuminated when the two angle of attack sensors disagreed was also fixed.
While reviewing the software systems, a number of other software issues were found.
The wiring bundle issue was also found during these reviews.
https://www.barrons.com/articles/these-6-issues-are-preventi...
> The MCAS software was modified to read from both angle of attack sensors and to be less aggressive in pushing the nose of the plane down.
That strikes me as a design issue in the domain of aeronautical engineering, rather than in software engineering. Software engineers aren't the ones with the domain expertise to determine the right aggression parameters.
> The software that controlled the indicator light that illuminated when the two angle of attack sensors disagreed was also fixed.
I thought the issue was that the sensor-disagree warning light was not included as standard, it was sold as an optional extra. [0]
> While reviewing the software systems, a number of other software issues were found.
Interesting, I wasn't aware of that. If I understand correctly these issues aren't thought to have a direct bearing on the crashes.
[0] https://www.nytimes.com/2019/05/05/business/boeing-737-max-w...
But the biggest defect was the decision that the pilots didn't have to be told about the risk (to avoid retraining and certification costs). If they had been told the pilots/airlines may well have protested about the lack of redundancy. All those higher ups in Boeing and the FAA should be in prison.
That said, all bridges have design specifications and limits. If someone drives a convoy of trucks carrying gold bullion over a bridge and exceeds its design max weight by a factor of 10, and it doesn't hold up, that's not the engineer's fault. If a bridge is designed for a geologically active area and is designed to withstand a quake 10x more powerful than the strongest that's ever been recorded in the area, or is predicted by seismologists, and the area gets hit by a 100x quake, that's not the engineer's fault.
And if a bridge designed to last say, 5 years, is neither re-certified for use, or just closed, when its intended lifespan is up then that's not the engineer's fault either.
But the engineer is responsible for finding out what the likely max load of a bridge will be, or how powerful the strongest quake will be, or how long its intended lifespan is - and for adding in safety margins on top of that. They can't just assume that any new bridge will only be needed for 5 years, just because that no-one told them that it will be needed for longer, and then claim "well, it was only temporary" afterwards.
The law takes care of that though. The law and inspection. If it was up to the business, any engineers building bridges who would refuse to project it faster because of that would be fired.
Compare bridge building to developing software for pace maker.
Now compare number of failures in both cases.
You get what you pay for, if you pay for developer you'll get a developer.
That's way too oversimplified. You can be as good and as thorough as you want, but if the root of the problem is that software is seen as a cost not as an integral part of the solution, you get bad results.
This often starts with unclear/vague requirements that change every other day as new information and understanding is gained.
But it doesn't stop there - unrealistic deadlines, lack of defined processes and quality control as well as disregard (and refusal to budget and schedule) for background tasks (documentation, refactoring, ...) are also contributing factors that can turn even the best and most diligent software developer into a messy code cowboy.
If gaming the system is rewarded more or actually doing a good job is even penalised, why do a good job?
If it's your job to do what is requested from you, then you do it. It's not like having brittle unmaintainable code is morally wrong. It's not engineers responsibility to judge use case of the customer paying for the bridge.
Actually yes, yes it is! I am a civil engineer and there's a standard of ethics and personal responsibility among engineers that is very, very high. When we graduate, most engineers participate in a ring ceremony. They get a funny, angled ring on their right pinky. It was originally made from the metal from a bridge that fell down and killed people. It was meant to be a daily reminder, worn on the drafting hand, that every single drawing you signed carried the weight of public trust of their lives. Even pedestrian bridges get built with vehicle load standards, because you can't just assume that it will only be used by pedestrians.
Civil engineering is a licensure, and no matter what liability structure you use, when you stamp a drawing, you carry PERSONAL liability for anything that might happen if that design fails. Even small violations of ethics or operating outside your scope of knowledge are dealt with very aggressively by the licensing board.
In school, if we misplaced a decimal or got a calculation wrong or failed to adhere to a very specific standard, it was automatic failure because if you do that in real life, people die.
And that is because no one can just go to a 6-month bootcamp and call themselves a Civil engineer. And even if they did, no one would hire them or the employer would go to jail.
If we want Software Engineering to be as rigorous as other Engineering disciplines, we need to erect similar barriers to become a Software Engineer.
It's strange too, you hear in other threads about how high the demand is for software engineers, how high their salaries are, how much negotiating power they have with their companies, but when it comes to the actual product content, suddenly they have no power at all and it's just Yes boss, whatever you say, boss. How is this true?
Anyway for most systems deployed the customer usually wants to maintain a business using it. If the system constantly falls over, cant be readily changed etc etc that is going to cost the customers business compared to its competitors.
Add that a lot of developers work for the same company as the customer and that is just as much the developers responsibility.
99.999% of bridges built would be expected to hold up 10 years later and require minimal maintenance. I'd wager all bridges even temporary army ones would be expected to not unforeseeably fail. See the Morandi Bridge tragedy.
I would expect that there are plenty of engineers that look at faults in the systems they're working on and say, "management doesn't care, not my problem".
It’s not that this type of engineering is “not engineering”; it’s that it’s engineering where the engineer themselves (and their ability to actively re-stabilize the system) is considered a load-bearing element in the design, such that the system will very likely fall apart once that element is removed.
Combat engineering is still engineering, in the sense that there are still tolerances to be achieved, and it’s still a disaster if the bridge falls over while it’s in active use for the mission it was built for. It’s just not considered a problem if it falls over later, once that mission is accomplished.
Do we put the cockpit in front of the boiler or behind?
Do we make a walking steam engine? How about using gears instead of wheels?
Which gauge to pick for the rails?
We'll get there in software engineering at well. It will be very reliable and equally boring.
1. It can be rebuilt in a matter of hours by a single person.
2. The customers that ordered the bridge are also the people on the bridge and are totally fine with the bridge crashing on them from time to time.
I remember the exact turning point. We had a super buggy Windows application that had tons of crashes in it. Instead of root causing each crash and fixing them, I was asked to simply write a launcher app that sat in the background, waited for the application to crash, then re-launch it. That was the great solution. And it was totally acceptable to the customer. Arghhhhh! I remember thinking: I didn't spend four years in university to shit this kind of finished product out.
TL;DR: I think the extent to which we are engineers is a choice we individually make.
On the other hand, if someone takes software reliability seriously, either they are going to burn themselves out to meet the crazy deadlines or are going to "not deliver on time", in which case they will be replaced by someone who can "get shit done" cheaper and faster.
If you want to design the software that designs the bridge or the car...just need a 8 week bootcamp.
Here is a 1992 truth bomb rewind about it:
https://www.developerdotstar.com/printable/mag/articles/reev...
Basically engineers care about the quality of the code because they know that is the design. As soon as caring about codebase quality leaves the building, so will the quality of the product.
Software engineers often hide behind a contract saying they aren't responsible for anything.
Software engineering is real and does require you to have a real understanding of risk (including through formal methods, model based design, systematic testing etc) in regard to the consequence. If you work in a more regulated industry (e.g. railways, aviation, automotive) these things are taken more seriously. This is why things like 'partial autonomy' driving needs more regulation -- no one is requiring Tesla to have a Chief Engineer sign off on anything meeting any standard.
The whole of the trading strategy (and capital traded, return profiles, etc) between the first two was fairly different, so it’s hard to compare.
I would say though that the first firm was significantly outperforming its peers in a way the second firm did not, although both were very successful.
The technology wasn’t the whole story in either case, but the first firm had significantly more and better opportunities as a result of just having a better stack all around
There's at least a couple of MMs struggling these days.
If you'd really really care about latency and have the resources, wouldn't it be more effective to just go bare metal? Bare metal as in: no OS (talk to the hardware registers directly from code), possibly a custom compiler, use all the relevant instructions that your CPU offers you. Or at the very least mess with the compiler to optimize it more to the specific make/model of the CPU/hardware that it will run on?
What are the considerations in those companies for or against this?
It might surprise people that relatively few hfts do latency arb (and the ones that do often have some other structural advantage), and at that the ones are more and more mixing in serious short-term alpha. In that context speed is simply becoming table stakes for playing the game, and not the one true race to win it all.
Compilers already offer flags to optimize for specific processor generations (and use the newly available instructions); you aren't going to be able to do better than that with a custom compiler.
I'm not sure whether it's worth it to go through this trouble (writing for bare metal vs just using an off-the-shelve OS) in a real life trading-app scenario. Hence my wonder :)
VM - the key is to get everything mapped in at the start and avoid page faults subsequently. madvise can help with this and obviously you need to avoid memory leaks that could result in sbrk - but in any case allocating at all in the trading thread is generally unnecessary (after startup) and frowned on. This does mean that the (lock-free) queue back to the shared core needs to recycle memory and both sides need to service it regularly or you need a strategy to deal with exhaustion. Scheduling isn't an issue - you have a single thread per core and (with isolcpus) the kernel scheduler knows to avoid placing other threads on it.
There are apparently ways to do this with an OS running in parallel as well though when you have multiple cores available apparently (see another comment to this thread).
- raw, unshared access to devices.
- no more interrupts.
You can get really close to this via virtualization though.
Run a guest machine without operating system on the hypervisor. It doesn't even have to be low level code to get most of the benefits (for example MirageOS)Get to OS to set up DMA for you.
> no more interrupts
If the OS has no need to interrupt you, it will never interrupt you. The OS doesn't context switch from a usefully running application on its own core unless you ask it to.
Don't need virtualisation.
"If you have an unlimited amount of time and resources, the best solution for speed will be coded in FPGA," says Lawrey, referring to the language that directly codes field-programmable gate arrays. "If you still have a lot of time and resources and you want to create multiple exchange adaptors, you'll choose C++. - But if you want to engage with 20+ exchanges, to go to market quickly, and implement continuous performance tuning, you'll choose Java."
I remember just refactoring all of foreach-loops/iterators into for-loops and got insane speedups.
Functions with inefficient iterations would get optimized if called directly but deep in the call stack and nothing happens.
These kinds of things are impossible with C++.
It's possible to write efficient Java but you have to not use some of the language features to do so.
Wow! Can you please provide a simple example when/where it happens?
> It's possible to write efficient Java but you have to not use some of the language features to do so.
Please tell us more.
An even worse case that made massive speedups is just inlining everything. I believe modern IDEs can automatically inline all invocations of a function. You'd be surprised how much performance you can get with that.
As for writing efficient Java, it comes down to using primitive and value types.
Although, my experience is mostly with Java 8. I have no idea how streams or other new syntax works. After optimizing a really massive Java code base I've never used it since.
You just have to try it and see. I remember it was quite easy in IntelliJ, I believe there's an option to inline all invocations of a function. I just went crazy with that option but did it systematically to really find where the bottleneck is and make a minimal change.
What? Anything you can do in C, you can do in C++ at the same performance level (pay only for what you use); just be aware of what you are doing and using.
(well strictly speaking, additionally you could also have hand-written assembly infra).
The main benefits of C per se are ABI stability and availability of compilers.
Rust is worse than C++ in that it struggles to express some idiomatic high-performance code architectures, such as using DMA for I/O, because they violate axioms of Rust's memory safety model, or in other cases because Rust currently lacks the language features. To make it work in Rust requires writing more code and disabling safety features.
Rust fixes many language design and performance issues with Java, while offering some similar types of safety, but (like Java) it is missing some elements of expressiveness of C++ that are important for high-performance architectures.
I've done my fair share of C++, including low latency stuff, and in the grand scheme of things I'd say "expressiveness of C++" is completely overshadowed by its complexity and occasional ambiguity, lack of proper type system and proper generics, lack of proper module system, lack of dependency tracking, lack of a unified build systems, etc. I'm not exactly sure what you mean by expressiveness though.
Instead of using an ownership-based memory safety model, which is clearly broken for systems that use DMA, they can use schedule-based memory safety models, which are formally verifiable. These don't rely on immutable references for safety, and also eliminate most need for locking -- many of the concepts originate in research on automated lock-graph conflict resolution in databases. The evolution away from ownership-based to schedule-based models is evident in database kernels over the decades. It was originally motivated by improved concurrency; suitability for later DMA-based designs was a happy accident and C++ allows it to be expressed idiomatically.
As for expressiveness, beyond the ability construct alternative safety models, modern C++ metaprogramming facilities enable not only compile-time optimization but also verification of code correctness that would be impractical in languages like Rust.
Isn’t everything you wrote from the wrong decade?
Pure latency arb isn’t the only hft trade there is anyways.
We used Chronicle as a high performance messaging system. The secret sauce is just memory mapped files, single threaded processes, and using Java objects as flyweights that just write to the end of the file, or are read from the end of the file. No GC. When taking market data off the wire, or sending messages around a system, that is competitive with C++ and certainly has a fast time to market, and lower defect rate.
The other thing people are conflating is high frequency part: get the market data and... do a complex calculation as fast as possible, and place or pull an order. Placing or pulling orders can be further optimized too.
The fast complex calculation part is where Java isn't as competitive and you run into the GC unless you use object pooling, warming up the JVM and so on. We had a 50-50 system of Java for the bulk of the system, and C++ for the complex algorithms where we mark hot paths and check the assembly on godbolt (thank you Matt Godbolt).
Picking two significant languages as a start up is a bold choice.
Measurement and QA is everything in terms of performance.
This was over 10 years ago, so things may have improved, but it would certainly give me pause about ever implementing anything time-critical in Java.
Aws is slow. Not talking about the network.
This is using the GC that ships with JDK 11. Future improvements, such as Shenandoah and ZGC, may be able to further drop the interruptions to 5ms on average.
The radio repeater is designed for real-time, safety-critical operations dealing with voice communications. Java is up for the job; like others have mentioned, it's more about the technical team and whether management understands what it takes to write and maintain an excellent code base than it is about Java versus C++.
Tick to trade latency for a modern software HFT stack needs to be under 10 _microseconds_, which is over three orders of magnitude more stringent.
They already just said they need 10 microseconds. 10 milliseconds is far far too slow.
Another relevant aspect is what happens when one hits GC pauses again. Will there be any tweaks left? Will they conflict with the previous ones? Using such high-level levers to fix performance problems in a specific component is great when it works, but when it doesn't you're out of options.
The biggest issues are that an engineer with the knowhow is expensive, a team of them is prohibitively expensive, a lot of companies outsource this work to different countries and make them sign onerous and dubiously-enforceable contracts, pay little, demand too much, and have incredibly unrealistic deadlines for the type of work that needs to happen.
You know how it is, if it never fails, you're throwing money out. That's the mentality of a lot of management types I've dealt with.
That's certainly one way to look at over five years of engineering effort, extensive systems testing, countless hours of automated regression test development, independent product quality verification, rigorous code reviews, risk assessment processes, mandatory four 9s call quality requirements, integration analysis of JVMs with deterministic GC, and so on.
Judge not a galaxy from a single star.
Doesn't Java itself say it shouldn't be used for anything like this?
(MISRA-C and other odd dialects excepted)
Consider that large chunks of Google, Twitter, Amazon, Alibaba etc run on Java.
The other major complication was that this was doing a lot of complicated protocol conversion, so it wasn't enough for the core message processing & routing to be solid and leak-free, all the input parsers and output converters had to be too.
Source: I worked on several extremely high performance Java systems at Sun and also worked optimizing various Java servers at subsequent jobs based on my Sun experience.
At one company (after Sun) I took on a backend service which was doing GC pauses every 3-5 minutes, the code was a dumpster fire mess. After cleaning things up we didn't hit a single full GC in two months of production use (they redeployed every two months so that's as long as the process lasted).
Productivity is the point. Even while paying attention to GC, it's much faster to write good code in Java than to worry about every malloc() in C. I love C, but if I need to churn out high performance code quickly, Java is the choice.
IMO (possibly biased by how much I hated C++ in the 90s) it is easier to maintain clean code discipline in a large team with Java than it is with C++.
Definitely depends on your team of course.
A few other reasons:
The rest of the platform for review and analysis doesn’t have such stringent performance requirements and it would be nice to reuse code.
There also may be some good libraries that are jvm only that your program relies on.
Packing up a fat jar can be a lot easier, less complex and more reliable than building a binary.
Your dev team has a lot Java expertise.
Getting Java to not allocate is not as hard as it seems if you set out to do it from the start. And even if not it’s not impossible because the tooling is so good. You just have to know the right tricks and it can be easier to incrementally learn those tricks for a dev team than learn C++ and it’s ecosystem.
Allocate either locally on the stack or allocate for the life of the process so it never gets GC'd.
Of course, spend time profiling to make sure you got it. Figure out your redeployment cadence and that sets the ceiling for how much data can be an exception to these rules.
(e.g. if you redeploy nightly then it's fine to grow the memory consumption as long as you don't hit full GC in less than 24hrs.)
In a nutshell that's it.
Programmers used to say that the GC and its fuc*#=+ STW was a pure mess, but under the hood, they didn't realize (accept ?) in fact the code WAS the mess.
Don't (always) blame the tool !
10 years has seen improvements to GC, but not eliminated the problem.
Cass 4.0 finally enables ZGC, so we'll see how that goes.
If you have stateless, you could do two JVMs and monitor the GC state, and reroute requests to a low-heap JVM and let the other clean up.
Garbage collection is hard.
I'm surprised that there aren't expert-level concepts like weak references and ways to sub-partition the heap, with say a new operator that can target a partition, so major GC doesn't have to do the whole heap. Although that starts getting into rust concepts and analysis.
Of course Python is slow, but it will call into efficient C routines for a lot of things that aren't just available out of the box in C++. So what does the average programmer do? They implement their own, inefficient version of this in C++.
Now you've saved the overhead of copying your data from the Python to the C-world and back, but the main part of your computations is just not on the level of numpy.
So naturally C++ can be much faster than Python, because it doesn't have this overhead, but you have to do it right, which may be more effort.
Going from a statically typed language to using python to do number crunching genuinely makes me want to vomit. I don't understand how people convince themselves that it's productive to write so many tests and (say) check types at runtime etc.
I actually love Python's syntax but the language design seems like it was thrown together over a weekend.
Totally agree, although I think strong typing is also important, not just static typing. C/C++ has a weak type system because of the automatic conversions it allows.
The problem with this discussion however is that the people that push for dynamic languages usually do it because it's "so fast" to do something in Python, and they won't test their software anyway.
It has a Pythonesque syntax, but strong static typing with extensive user control of semantics you seem to care about; e.g. floats and ints do not convert automatically unless you explicitly "import lenientops"[0] ; You can define 'operational transform' optimizations (such as: c <- c+a\b converts to multiply_accumulate(c,a,b) - which is a big performance difference for e.g matrices) that will be applied by the compiler so that your code shows what you mean (c=c+ab), and yet the compiler gets to compile the efficient version (using the relevant BLAS routine)
It's young, but has the best FFI for C,C++,Objective-C or JS you'll find anywhere on one hand, and already a good deal of native implementations, including e.g. a pure Nim BLAS that is comparable to within 5-10% with the best out there (including those with carefully hand optimized ASM kernels).
It has the spirit of Python 2 but written by a compiler geek (in the good sense). If I were to write an HFT it’d be tempting to use. The new default GC is reference counting but without using atomic ref counts.
P.S. thanks for the lenientops tip
Other than syntax obviously.
FFI: Can nim use a C++ class and vtables? D does it all the time, nearly every language proclaims to have the best C++ interop but only D seems to be actually able to do it. Templates, classes, structs, and more all work.
We also have Mir, which I haven't benchmarked for a while but was faster than OpenBLAS and Eugene a few years ago and was recently shown to be faster than numpy (the c bits) on the forum.
Well, it depends on how you define "best". Beyond an ABI, FFI is obviously a function of the implementation rather than the language per-se.
The main Nim implementation can use C++ as a backend language, and when it does, it can use any C++ construct very easily by way of the .emit and .importcpp directives. For sure, classes, structs, exceptions all work, and IIRC templates do too (although you might need to instantiate a header yourself for each concrete type or something .... haven't done that myself). This implementation also means that it can use any C++17 or C++20 construct, including lambdas and friends. Does D's C++ interop support C++17? C++20? Can you guarantee it will support C++27? Nim's implementation already does, on every single platform you'll be able to use C++27 on (as long as C++27 can compile modern C++ code; there had been backward incompatible changes along the C++ history).
You can't just #include a C or C++ header and call it a day; You need to have a Nim compatible definition for any symbol (variable, macro, function, class, ...). There are tools that help you and make it almost as easy as #include, such as nimterop[0] and and nimline[1], and "c2nim" which is included with the Nim compiler is enough to generate the Nim definitions from the .h definitions (though it can't do crazy metaprogramming; if D can do that, then D essentially includes a C++ compiler. Which is a fine way to get perfect C++ compatibility - Nim does that)
But Nim can also do the same for JS when treating JS as a backend.
And it can basically do the same for Python, with nimpy[2] and nimporter, generating a single executable that works with your installed Python DLL (2.7, 3.5, 3.6, 3.7) - which is something not even Python itself can do. There was a similar Lua bridge, but I think that one is no longer maintained.
> We also have Mir, which I haven't benchmarked for a while but was faster than OpenBLAS and Eugene
There's quite a bit of scientific stack built natively with Nim. It is far from self-sufficient, but the ease with which you can use a C library makes up for it. I haven't used it, but Laser[3] is on par with or exceeds OpenBLAS speedwise, and generalizes to e.g. int32 and int64 matrix multiplication; Arraymancer[4] does not heve all of numpy's functionality but does have quite a few nice bits from scikit-learn, supports CUDA and OpenCL, and you can use numpy through nimpy if all else fails. Also notable is NimTorch[5]. laser and arraymancer are mostly developed by mratsim, who occasionally hangs out here on HN.
D is a fine language, I used it a little in the D1 days, and it was indeed a "better C++" but did not deliver enough value to be worth it for me, so I stopped. I know D2 is much better, but I've already found my better C++ (and better Python at the same time!) in Nim, so I haven't looked at it seriously.
[0] https://github.com/nimterop/nimterop
[1] https://github.com/sinkingsugar/nimline
[2] https://github.com/yglukhov/nimpy
[3] https://github.com/numforge/laser
Now I can't stand untyped python. Have to actually look at documentation to see what methods a class has, who has time for that?
Writing golang with tabnine is otherworldly. It's so regular, not just the syntax, but the variable naming conventions, and course structure even, that it feels like loosely nudging the autocomplete in the direction of the goal.
>pointer indirection
30+ years ago, a student at University, i loved how pointers allowed to code algorithms efficiently and expressively compare to the non-pointer languages of the time. In the industry today the situation is completely opposite as correctly noted by the parent.
Compare this to many functional languages (or lisps) where the most convenient data structure to hand is the singly linked list. It was already bad for performance when those languages were invented but in modern times they are relatively much worse than before.
Sometimes I think that the performance of C programs come more from the fact that it such a massive pain to do anything in C that you can usually only do the simplest thing and this tends to be fast and reliable. The problem is that if you can’t do a simple thing it will either take a lot of refactoring or you’ll do a complicated thing and your program will randomly segfault.
This is also my theory as to why languages like C have better libraries than lisp: it is such a monumental pain to do anything in C that people go to the small extra effort to package up their achievement into a library either for others or in case they need to solve the trivial problem again themselves. Improvements can them come later as needed but get shared. Compare this to, for example, lisp where the attitude is usually that libraries aren’t flexible enough and generally not that hard to implement yourself (so long as you aren’t so worried about data structures), and I think this is the reason there don’t tend to be so many libraries, especially for trivial things.
I guess very limited number of constructors, and most of all the methods are static.
And they still struggle to find programmers that can code that way.
Assuming those libraries do not do allocations of their own negating all the efforts.
Of course at this point the usual answer is just use Rust, but is there a language that meets in the middle? Sometimes I just want the GC to do the work and I'm okay with that.
edit, to give you an example. Let's say you have double precision X/Y coordinates. Put that into a class with two double arrays for X and Y. In order to e.g. run an Euclidean distance computation against it, provide a (double, double) -> double interface against it. Then you have either an Euclidean distance class providing that function or just an inline lambda that you can plug into that interface, giving you a distance array. Let's now say you just want to have a list of 10 closes points - instead of sorting you're fastest to brute force it. Only after you reduced it down to the 10 points, you put them into Point objects because there's probably a consumer expecting it that way.
Yes: Java
Performant Java looks mostly like C, just with no manual malloc() calls.
Sure you could write Java in a J2EE way but.. that not a good choice.
All modern VMs (not just limited to Java here!) apply two key optimizations. The first is escape analysis, which checks if references to objects will escape the current function boundary. If not, the objects will be stored on the stack instead of the heap. The second is generational GC, where memory allocation looks like this:
void *new_ptr = heap_mem;
heap_mem += alloc_size;
if (heap_mem >= max_size)
outlined_function_to_get_larger_blocks_of_memory();
It's actually likely to be a tighter allocator than C/C++'s malloc, since there's no mucking about with freelists.It really can't. Quite the opposite. Writing C-like Java is already doomed from the start. Java is faster than C when it comes to OOP. Unless you product is a highly complex OOP nightmare, C will always beat Java to the curb.
GC is not free. The problem with GC is that you pay asynchronously, while allocations are essentially free. But this asynchronous cost is very very hard to measure and to control. Java does not offer the mechanism C/C++ offer for native resource management. It also doesn't do complex optimizations, most notably vectorization. Number crunching performance will always suck in Java. Even if they add all the optimizations in the world, the simple fact remains: Java as a language simply does not allow you to express code in a performant way. That leave the compiler at a double disadvantage. It needs to essentially "convert Java to C++ and guess the most performant interpretation" (we are decades away from this), and it needs to do that within milliseconds (because its a JIT).
It just makes no sense to talk about this. Use Java for the 99% of your product that is not a hotpath, use C for the remaining 1% where you need pure performance. Simple as that.
Pre-allocate arrays, avoid collections.
Use LUTs.
Avoid string processing, operate on bits and bytes as much as possible.
Same applies to telecoms-development that has to be fast and predictable.
Funny enough, I've always thought of it more of writing Java like Erlang than C.
And C is faster still when you write it with Fortran like (column order, and "restrict" mostly) semantics.
An old mentor said to me 30 years ago that "A good programmer writes FORTRAN no matter what language they are using" (he was referring to speed of execution in the context of "good"). It seemed super funny at the time, but there's a lot of truth in that.
What I meant was: In FORTRAN, you idiomatically do Struct-Of-Array=Column-Major order by default, In C/C++ you do Array-of-Struct or Array-of-Pointer-to-Struct (both of which are Row-Major) by default. The former tends to be much more efficient than the latter in computational code.
It turns out that, often and especially in computationally intensive code, only a small number of fields of an object/record are used in every part of the computation pipelines, but almost all records are.
As a result, if you do your data in column major order, then the cache usage patterns reflect that - whereas if you are in row major order, a lot more field get loaded into cache when they are not needed (thus, much lower cache utilization and lower performance).
FORTRAN did not have structures until quite late (FORTRAN-95 has them for sure, but IIRC FORTRAN-77 and earlier didn't). Thus, the idiomatic way to do stuff is have an SoA=Coiumn-Major implementation, rather than the C/C++ AoS=Row-Major order.
Furthermore, FORTRAN didn't have any pointers. As a result, most code was written with very little pointer chasing / foreign key reference chasing -- which also contributes to efficient use of cache and memory bandwidth.
Creating liquidity has some value - though not as much as bankers tell themselves.
But HFT does not add liquidity. No one needs subsecond liquidity, but even more important, HFT traders vanish like a fart in a windstorm the moment liquidity is actually challenged.
Market makers are the ones who provide consistent liquidity, because they have to keep inventory of the stocks.
Now look at what HFT does - they flood the wires with hugely out-of-the-money bids and asks, all of which are in bad faith, because they know they won't even get hit or lifted.
So everyone else pays for this, and they get to shake absolutely everyone down for a few pennies, a trillion times a year.
---
It would be trivial to stop them without changing the markets in the least. Even just charging a dollar for every "bad faith" bid or offer made that expires without being within, say, 10% of the actual price, would put a big dent in them. Punitive taxes on trades where the trader doesn't keep the position for more than a few seconds would also be very effective.
And investment banking does?
>Even just charging a dollar for every "bad faith" bid or offer made that expires without being within, say, 10% of the actual price, would put a big dent in them.
Tell me this, what is the actual harm of these orders?
Whataboutism.
I'm tired of hearing this and other myths like "C++ is for geniuses". When are we going to acknowledge that managers only care about shipping feature fast regardless of the quality of the codebase? I work professionally with "productive languages" and I could claim the same thing about javascript and others.
On this article it is very clear that companies won't invest in proper training, quoting: "Most Java developers are web developers and are used to modelling things they can see" and "We began by hiring people from trading floors, but have more recently hired Masters and PhD students"
This is a culture problem, not a language one. If anything, C++ will always perform much better than Java. Claiming the opposite because you assume your coders don't know what they're doing, it's just ridiculous.
I mean, you can write a very fast system in Java but by the time you accomplish that goal you’ll have spent as much or more time than if you had just used C++.
But otherwise they’re probably the same.
But I get why people might feel it's worth a try, cpp probably looks more daunting than any other language to a novice. There's a lot of concepts to be mastered before you can write decent cpp, where in a lot of languages you can get started and discover things later. With cpp the abstraction gets very detailed. You might not even care what the difference between stack and heap, value or reference passing is in other languages.
Eg you can write python pretty intuitively and get a hash table as part of the language. If you're writing cpp you need to know how types work, how templates work, and how the STL containers work (allocators, traits, constness etc) just to save a simple hash table.
Working in other languages, the idea of references and values didn't kick in to me, that is until I learned basic C++. I think it's because C++ engineers (well, the ones I would watch anyway) were sticklers for performance and so they'd describe the subtle differences in implementation. Even though I don't code C++ anymore, using it for a short time really helped to develop a better understanding of how code actually works.
The trend is moving towards quicker time to market and simpler implementations and replacing cpp with java. The cpp code bases I had to deal with were old, hard to maintain and easy to break. More often than not cpp was a pretty bad lock in as well. For example a big evil bank very well known here struggled for 3 years to upgrade the compiler on one of those projects. Did not happen to this day, I'm told. Few years later they are still on RHEL6 and gcc 4.
Speed is an important factor in this business, but so is time-to-market.
Banks are not competing on latency. It's never been part of their business model.
Meanwhile if your logic isn't going into FPGA you may well just not be super competitive anyway and therefore the priority is time to market.
(Source: hft quant)
Maybe the JVM's GC is just so much faster that it's worth it?
Apparently it's production ready-ish.
bing.com/version
Nope .NET is still slower than molasses in comparison to most JVM implementations.
.net core is generally faster than java at this point.
https://devblogs.microsoft.com/premier-developer/understandi...
Joking aside, C# has only become a solid contender after the fairly recent performance improvements and better Linux support. I could be wrong but I doubt a lot of these systems are built on Windows.
.NET on Windows for HFT is a non-starter. In comparison to Linux, you have practically zero control over the OS; you need to turn off unnecessary processes, pin processes to cores, exclude processes from cores and so on.
Having said all that, I do know of at least one hedge fund in London that is mostly C#-based (though I have no idea whether their 'deep' trading code is C# though).
They are. However I'd say, based on what I saw at Pivotal, that .NET Core is sweeping all before it. It's faster, less crufty and so on. But the real win is that (1) most of it is familiar to folks with previous .NET experience, but (2) your ops teams can get out of running two very different platforms in Windows and Linux.
But still the topic ("trading system" or in our case "routing engine") is already very complicated due to the involved algorithms and C++ likely makes it a bigger challenge also regarding attracting good developers.
The memory waste in Java is the biggest challenge that you have to solve. Usually you solve this via primitive arrays i.e. do not use List<Point> but int[] x and y and before this gets complicated wrap this in an iterator or other thin Object (flyweight pattern?). And today you can use advanced GCs like ZGC or Shenandoah that work well also for >100GB of RAM and you can heavily reduce hiccups. And you need to be aware of some smaller pitfalls e.g. foreach loop or Java streaming API might be slower than the raw loop. But you should never follow rules or trust a "guru" and only use your performance tests to guide development :)
And in non-performance critical sections (70% of the code?) you just do not care and can program "idiomatic Java". Also do not underestimate the crazy work that a modern JIT can do for you.
And now after years of work we are for several aspects better or same memory-wise and speed-wise compared to other open source C++ routing engines. And the interesting thing is that IMO this is because of the tooling and habits (like rigorous testing) which reduces bugs and debugging sessions and slowly shifts your focus from language problems to the actual domain problems. But of course all very subjective.
https://www.reddit.com/r/rust/comments/bhtuah/production_dep...
Also, if memory serves, one of Lawrey's ex-colleagues has written his own version of Chronicle in D that's being used somewhere too.
First, let me preface this, the language does matter, but the larger complications revolve around the operating system. Most transactions have a hard real-time deadline of 20ms or less. JVM creates additional issues on top of what are already present from the OS. If you're on Linux, a real-time kernel is required to handle these transactions.
For Windows, there are asynchronous methods you achieve this functionality, it is possible, but very difficult to do.
After reviewing this guys LinkedIn profile, all he has ever done is Java. He is a Java advocate. I think Java is great language, but this is one place it doesn't belong.
This is "why Java is better", and the answer is "it usually is better because your programmers often suck".
I can, and have, written high performance systems in assembly (long ago). I've written many high performance systems in C. Rarely, but sometimes, still do.
But doing it in assembly takes a long time. No way to justify time to market on that one today. In C, sometimes it's worth it.
But Java productivity is so much higher and (if written properly) leaves very little performance on the table. It's the sweet spot of performance vs. productivity still today and has been for about 15 years (~Java 1.4).
If they do suck, then almost everybody in the whole industry sucks
The amount of changes you need to make due to customer preferences and regulations is continuous. If it's ok for your system to run in the 100 microsecond range then Java is a clear winner. If you are just focused on a single exchange, running in the same rack as the exchange then of course C or lower is what you will want.
The core features are stable for a long time. The days that you needed nightly for pretty much ever web framework are behind us.
On the other hand it we have some niche edge in a smaller market that's less speed competitive, Java might very well be the better choice given that we might be able to capture that edge with 8us tick to trade latency.
So the answer totally depends on what edge you're trying to capture. There are still a large number of edges where Java is good enough. GC is not really a problem, either using a commercial JVM or coding without much garbage and large memory.
You can never get around the boxing/unboxing, the larger sized objects, jvm inefficiencies, lack of direct access to assembly, etc.
But the cost of building a super low latency setup, and then doing anything useful with it is much higher than doing it with Java/c#
Does anyone have any guidelines for this type of code?
'The only problem with low latency Java is that most experience(sic) Java programmers struggle with the new paradigm.'
So...badly written C++ isn't reliable, and well written, idiomatic Java isn't performant...so the solution proposed is, rather than make sure you write your code well in C++, to instead write your code well AND non-idiomatically in Java? Seems an...interesting take.
whereas in Java you get an exception, a stack pointing neatly at where the issue was and an error that explains what's going on, and you can fix it and move on.
I'd argue that C-like Java is much easier to debug than C.
It's easier for a programmer to learn a new PL (at the start of the project) than for a 500kLOC code base on GC-enabled JVM to migrate to Rust.
It's a bit harder to write some bugs in Java than it is in C or C++. I gather it'd be almost impossible to write them in Rust.
This is where modern C++ (augmented with, e.g. the C++ Core Guidelines) and Rust are more convenient and even "better" than Java. Fancy libraries and tools, with zero additional overhead compared to the optimal low-level code.
That is a common sentiment. So much work went into this! We don't understand how we got here in the first place. It's so big! Look at all this legacy code and technical debt! There is no way we or anybody else could rewrite this in Java let alone C++ in any useful time frame! We sunk decades of work into this monstrosity!!
The text boils down to a "time to market" argument. I find that surprising because you'd think for high speed trading correctness is more important than time to market.
Each person who loses money on the market loses it to someone. They are ready to be that someone, for you.
With all that real estate on CPUs, I really would like an FPGA. If Oracle wanted to make money on Java, it could offer a commercial java that compiles bytecode to FPGA. Apple with Rosetta2 could have done an FPGA.
Emulation of lots of systems that will be abandoned by their commercial companies has always relied on the mhz and node shrink "free ride", and that's over. FPGAs could handle a lot of legitimate emulation.
Choosing Java over C++ is choosing to be that second guy.
Sure, Java coders are cheaper and easier to hire, and they produce lots more code. But if lots of code is a liability, lots of Java code is a bigger liability.
A trading house with a lot of Java code is a red warning.
[1] MIT 6.824: Distributed Systems - https://www.youtube.com/watch?v=cQP8WApzIQQ&list=PLrw6a1wE39...
When someone says "C++" instead of C, you're really looking at developer issues, not technology issues. This is a one-off article trying to make some other point which seems like management hand-wringing to make sense of project failure.
> Java developers are better than C++ developers for high speed trading systems
I would read this, especially if it was an actual study.
However, if you are a bank, which not where every developer wants to work, you have a limited budget, limited time and you have a project which isn't technically exciting (but can make a lot of money for the bank) you might find the developers you have who understand the business are more productive in Java and the resulting trading system is faster than the one they might have written in C++.
When you have some teams developing in Java and some in C++ you get to see which teams are more productive for your organisation.
Interesting to note the comment section there vs here.
Is there any open source project that one can study to understand how to code like this?
The results were basically -- with the best teams, almost identical times; the obviously slowest entries were on the C++ side. Admitting bias now, I was surprised to see Java winning by a "hair's breadth" in certain cases, with certain entries.
Later, talking to a PhD in CS who wrote their thesis on graph-analysis of code conditions for JIT optimization, I found out that Java put a lot of excellent computer science into the problem, and that works very well. As alluded to in a different comment, the response times in both camps were appproaching the physical limits of the hardware, apparently. I had to recalibrate my expectations, since I thought that C++ would certainly win without fail.
Java is pretty good, I've used it for over 20 years, but it wastes memory for basic use cases, and yet GC pause limits the utilization of huge memory cases.
Did anyone try Go? IMO Go tries to address the slowness of java startup and commands, but it punts GC/memory waste so the huge-memory systems are even worse.
no one tried go-lang, the elaborate specialty code needed for the actual work did not exist
the Java probably did use pre-allocation and flat heirarchy a lot, but I do not know the details. Some Java did win by a significant margin, in very odd corner cases, probably due to simple code-coverage miss on the C++ side
In the case of high speed trading systems, any second they are late to market, they are losing money. It might be the fastest thing ever once it's finished but while it's not running, its slower competitors have the market to themselves. Worse, any month you are late is time your competitors can use against you to learn from your mistakes and beat you at your own came. High speed trading is of course extremely competitive. So things change all the time. So, now we are talking speed of development and maintainability in addition to raw performance and throughput. The impact of most improvements you make are in any case typically temporary in the sense that they only provide you an advantage as long as your competitors don't catch up. So, getting bogged down into lengthy maintenance and bug fixing cycles while your competitors copy everything that you got right means you lose some or all of the market opportunity.
This article mentions what happens when C++ projects fail. Basically, you are late to market with something that doesn't work as advertised (slow, unstable). That's not an inherent risk of course but it's a risk none the less. Performance often comes at the price of complexity. Complexity has its own penalties as well in the form of e.g. poor maintainability, bugs, or stability. C++ is frankly notorious for this. You compensate with having skilled engineers acting in a disciplined way. And sometimes that actually works as advertised. The project mentioned in the article sounds like it had issues across the board. Probably, at least one of the root causes for that was that they were trying to take some short cuts to get it to market.
The JVM ecosystem has a lot going for when making such trade-offs. That's why it's so entrenched in the industry. And it's been around long enough that there are lots of solutions to all sorts of issues that you might need solutions for. Also, you always have the option to go native in Java; it's not an either or kind of thing.
For the same reason you see people opting for a not particularly fast language like python in a domain where you get to use fast GPUs to get things done quickly (i.e. machine learning). Python gets that job done, everything on the critical path is basically delegated to lower level native stuff running on a GPU (or custom hardware in some cases). I wouldn't be surprised to learn that some high frequency traders have lots of python code around in addition to their precious native code. It would be a pragmatic thing to do.
Understanding your domain and understanding the trade offs is what differentiates successful companies and engineers here. There are lots of valid reasons for preferring C++ for some things. Personally, I rarely encounter such requirements and I'd probably prefer Rust over C++ if I did. And since I don't, I actually have limited experience with Rust as well because other than the intellectual challenge I don't see how I'd be getting much real world advantage out of it. I'm pretty sure that holds true across large parts of our industry. It's why you see people actively not caring about performance by choosing to run on slow (virtual hardware) on slow cloud networks using slow run time environments that include slow interpreters and slow application frameworks because they can get stuff done quickly with it and can trivially scale by throwing dollars at the problem. Slow is fast enough for most.
As a game HFT is interesting but in real life its just a plain scam. Yes you may have a job and you may work on HFT but that doesn't mean you have to give your heart and support it.
Disclaimer: I worked on HFT but I still don't like it and the money HFT generate is just insane.
The smartest people I know would always try to simplify things, but many C++ developers (especially in financial industry) have this preference for complexity I can't explain. And C++ is the worst language to get clever with.
From my experience: (apologies if this looks like a flame; purely personal experience, with no offense intended).
- Python attracts all kinds of people.
- Java attracts people who prefer rigid structure, specified down to minute detail (which non Java people will often consider unneeded bureaucracy)
- C++ attracts people who like brain teaser challenges with lots of details and steps. The compiler is both their colleague and adversary in reaching the optimal speed goal (their "frenemy")
- C attracts people with more experience who just wants "the simple thing that works" without giving up the low-level control that they would have to give up using e.g. Python or Java.
- K/J/APL attract people who like brain teaser with short elegant solutions, even (and perhaps especially) if reaching and understanding that solution takes nontrivial effort -- which is basically also people who like math for math's sake.
- COBOL attracts people who like job security, good income, don't care much about job satisfaction or recent advances made in the last 30-40 years, and want everyone to get off their lawn :)
You're not wrong, the ego among C++ devs is astronomical. There's a huge amount of serious hate for python and this idea that being a C++ dev is superior.
Get this we use Nix and C++ and the nix part of the code base is just as complex as the C++ code base. The nix stuff is just infrastructure! It's insane.
If you wanted to get the same performance from C, you'd sometimes need to write your own preprocessor/code-generator, and/or abuse the built in preprocessor, so you can specaialize. To get the same performance in Python you need to go down to C. To get the same performance in Java, you need to do stuff like what's mentioned in this thread (allocate 1TB of memory and disable GC, for example).
The only place where this is considered acceptable and commonplace, is the C++ community, I guess a chicken-and-egg problem.
In most places, you only need that performance in a small part of your pipeline - which in HFT teams will sometimes be done in FPGA or ASICs or NIC firmware. For those places, C++ is also often required, and network effects (and "impedance mismatch") often mean that it will get used for a whole lot more than where it's actually essential.
C++ is not the only such community for algo trading, by the way - the K programming language enables fast yet incredibly short implementations, and thus tends to attract competitions in "how short can you get the fastest code" (aka code golfing) to manifest all over implementations. And personally, I'd much rather have a oneliner K brainteaser that does alot than a huge 100-line C++ template monster that does the same thing 10% more quickly.
What would josh.com say..?