We chose Java for our high-frequency trading application
medium.com
medium.com
I worked in the industry and it's always a little funny to see who calls themselves HFTs vs quants.
Basically, there's a bit of a spectrum of fast vs smart. In general it's hard to do incredibly smart stuff fast enough to compete in the "speed-critical" bucket of trades and vice-versa there's barely any point in being ultra-fast in the "non-speed-critical" bucket because your alphas last for minutes to hours.
Just from this read, I feel like these folks are just a hair to the right of "fast" in the [fast]---------[smart] continuum. I mostly make this appraisal based on these paragraphs:
>To gain those few crucial microseconds, most players invest in expensive hardware: pools of servers with overclocked liquid-cooled CPUs (in 2020 you can buy a server with 56 cores at 5.6 GHz and 1 TB RAM), collocation in major exchange datacentres, high-end nanosecond network switches, dedicated sub-oceanic lines (Hibernian Express is a major provider), even microwave networks. >It’s common to see highly customised Linux kernels with OS bypass so that the data “jumps” directly from the network card to the application, IPC (Interprocess communication) and even FPGAs (programmable single-purpose chips).
That's nice but that's where the cutting edge of the speed game was in 2007ish. Everything mentioned here is table stakes at this point (colocation, dedicated fiber, expensive switches, bypassing the kernel in network code, etc). The fact that "even FPGAs" is listed as "even" is the biggest thing I focus on. FPGA's and/or custom silicon is where the speed game is right now. Similarly, "even microwave networks" is also table stakes at this point (you can get on nasdaq's wireless[0] just by paying).
This is the kind of game where capex for technology is dwarfed by the margin you're slinging around every day in trading, so you see some pretty absurd hardware justified.
[0] http://n.nasdaq.com/WirelessConnectivitySuite
Edit: Also shout-out to a different comment in this thread mentioning ISLD, a story I considered telling as well: https://news.ycombinator.com/item?id=24896603
I think as with any space, if you really zoom into the details, you see a lot of diversity. You're obviously right in that people who are still hitting the CPU can't compete with people that do everything on FPGA, but it seems like there's still plenty of money to be made by people who are just a bit less fast than that.
I've heard people argue that they prefer being in that sort of space, too, as it gives them a bit more room to compete on their own alphas, and tends to be a bit less winner-take-all.
I'm not trying to gatekeep who can call themselves HFTs or not. The main thing I find funny is if you ask 10 firms if they're HFT or quant shops, it will probably not actually line up all that well with exactly how many orders they send or how speed-sensitive they are.
Depending on the overall strategy, you can be smart and fast at once. If you have occasional alpha harvesting opportunities that must be acted on quickly (very low latency but also low-ish frequency), then it is possible to spend the time in between trades modeling market context and developing optimal short-term plans for how to react in case of various market triggers.
Speed always matters, no matter where on the smartness spectrum you are, but it's relative. If your model prediction takes 5ms you're not getting much ROI out of investing $1M into shaving off 50ns in your data processing. But if your end-to-end latency without prediction is 1ms, you better invest in getting that down.
Let's say you have trading opportunities once every 100ms. They need to be acted upon within 2ms or they vanish.
You don't have the time budget to run a NNet every time the state of the market changes. You can, however, train a NNet to output a very small decision tree that can run in under 1ms. The NNet can then decide which "micro-strategy" in the form of a much faster and more reactive decision tree is more appropriate for the current market context.
DSquare is reasonably famous in London, somehow everyone knows them. They are on the smarter end of things, because the founders were plugged into the early development of the electronic FX market at the beginning of this century.
There's a sort of engineering-vs-business culture thing to this as well, but I think it blurs over time.
Why do you need 1TB of RAM in these machines? Because when you're Java based, you want to avoid stop-the-world GC pauses. These trading systems only have to be up from 9:30AM-4:30PM EST, so they simply disable GC altogether! At the end of a trading day, restart the app or reboot the system.
Of course, it's possible to write Java in such a way to minimize this type of heap-thrashing. But by that point, you're already doing the equivalent of C++'s manual memory book-keeping anyway. It's not like "just provision a ton of memory and turn off GC" is a free lunch.
@State(Scope.Thread)
@BenchmarkMode(Mode.AverageTime)
@OutputTimeUnit(TimeUnit.NANOSECONDS)
public class EscBenchmark {
Rnd rnd = new Rnd();
public static void main(String[] args) throws RunnerException {
Options opt = new OptionsBuilder()
.include(EscBenchmark.class.getSimpleName())
.warmupIterations(5)
.measurementIterations(5)
.forks(1)
.addProfiler(GCProfiler.class)
.build();
new Runner(opt).run();
}
@Benchmark
public int testEscapeAnalysis() {
int[] tuple = {0, 2}; // esc analysis? where are you?
return tuple[rnd.nextPositiveInt() % 2];
}
}
And the output of GC profiler: Benchmark Mode Cnt Score Error Units
EscBenchmark.testEscapeAnalysis avgt 5 8.234 ± 0.029 ns/op
EscBenchmark.testEscapeAnalysis:·gc.alloc.rate avgt 5 2647.216 ± 9.275 MB/sec
EscBenchmark.testEscapeAnalysis:·gc.alloc.rate.norm avgt 5 24.000 ± 0.001 B/op
EscBenchmark.testEscapeAnalysis:·gc.churn.G1_Eden_Space avgt 5 2643.140 ± 177.137 MB/sec
EscBenchmark.testEscapeAnalysis:·gc.churn.G1_Eden_Space.norm avgt 5 23.963 ± 1.613 B/op
EscBenchmark.testEscapeAnalysis:·gc.count avgt 5 157.000 counts
EscBenchmark.testEscapeAnalysis:·gc.time avgt 5 103.000 ms () -> System.out.println(1)
I lost hope in escape analysis quite frankly.Rnd is something I have written because Java's Random is slow and clunky.
https://github.com/questdb/questdb/blob/master/core/src/main...
I'd imagine other HFT companies may also pay for the Azul JVM which goes even further and comes with a kernel module that the JVM coordinates with for memory allocation. The kernel module pre-reserves x% of system memory at the time it is loaded.
Basically one main reason you need to stop the world in common GC's right now (and thus cause pauses) is because if you want to compact memory a GC thread has no idea if other threads reads/writes to an object while moving leading to worst case corrupt/lost writes.
What Azul does is to make the pages in question invalid for reading/writing so any access to them would page-fault and then proceeds to move the objects without worrying about other threads. Since the hardware provides the page-fault "for-free" this protection doesn't cost anything in terms of runtime performance in the common case (compared to the new ZGC made by Oracle that has read-write barriers that steals some mutator performance), worst case if another thread tries to read/write memory that is being moved a specialized page-fault handler detects if this was an page that was in motion and can then do a slower read-write but ONLY in the seldom cases where this occurs (compared to all the time of f.ex. ZGC)
https://devblogs.microsoft.com/oldnewthing/20180228-00/?p=98...
Even this isn't fail proof right ?
Last-drop latency improvements. For ultra-latency-sensitive applications, where developers are conscious about memory allocations and know the application memory footprint exactly, or even have (almost) completely garbage-free applications, accepting the GC cycle might be a design issue. There are also cases when restarting the JVM -- letting load balancers figure out failover -- is sometimes a better recovery strategy than accepting a GC cycle. In those applications, long GC cycle may be considered the wrong thing to do, because that prolongs the detection of the failure, and ultimately delays recovery.
So essentially you need to know the memory footprint of your applications. To err on the side of caution just get as much memory that is available in the market.
At this point like another reply mentions here aren't you just better off writing C++ code ?
I think I heard of the same pattern before 2018, so I guess there are other ways to do it. Perhaps Java has some option like memory usage value at which the GC is run.
There are a lot of tricks though to not require 1TB.
And allocation in general is a bad idea even if you don't collect because it scatters stuff all over memory and messes up cache locality. You really, really don't want to allocate in a performance sensitive jvm application if you can avoid it. It's the opposite of a lot of what I was told and taught (e.g. never do object pooling), but empirically, in my experience, allocations are the biggest slowdown. You can get an application a lot faster just by opening up the memory allocation tab in a jmc flightrecording and refactoring the biggest allocators, usually there is a lot of easy to optimize low hanging fruit that will give good performance improvements, even better than focusing on hot spots in code (in my personal experience).
By far the biggest allocator in trading is going to be marketdata and calculations on it. For reading marketdata from the exchange it's best to leave raw data in memory and access it with a ByteBuffer / sun.misc.unsafe. Under this pattern classes have 1 value, the memory address to pass into sun.misc.unsafe, then everything from there on is done with offsets onto that address. For calculations it's better to write things as static functions, or use object pooling.
In the course of optimizing a trading engine I wrote lots and lots of code to get allocations down to zero. It's definitely doable, but best done from the start, I refactored an existing trading engine to do that, it was not very fun.
Reminds me of the classic WTF "That would've been an option too" https://thedailywtf.com/articles/That-Wouldve-Been-an-Option...
The article in question was written in 2008: I was 11 years old. I guess I now fall in to the category of “heard of it by the age of 30”.
It's not about the amount of RAM, it's how fast and predictable the overall system is. Note in particular the remark about cache locality; that can be the difference between nanoseconds and microseconds.
HOWEVER If you DO _allocate_ much then 1TB seems like a great choice since GC pauses are killers w.r.t. latency in comparison to cache issues. A cache line (often 64 bytes) would probably only hold at most a couple of Java objects (minimum is probably like 16 or more bytes for each small object) so cache locality won't be improved much by a GC (Yes, the G1 GC in newer JDK's does neighbour compacting iirc so you get a little cache locality but only if the patterns were bad from the start, but avoiding GC entirely is better)
Also JNI for calling C/C++ code is fairly bad and error prone (if you want to combine code), JNA is better but i think this is one of the bigger reasons that Oracle is funding GraalVM is that it promises "seamless" interoperability and breaking up things into Java and C/C++ parts might be a good option in the future.
GraalVM is the new name of MaximeVM one of the JVM meta-circular JVMs. Other well known ones being JikesRVM and SquawkVM.
What Oracle and others in the Java community are doing is reducing the need to drop down to C with support for value types and more fine grained control over native memory.
Java 16 will have the first preview release of the new low level memory APIs.
Those unsafe bits are a tiny portion of the overall code.
Damn those smart-ass interns!
Over a decade ago, we used a third-party pre-trade risk system that was implemented in Java. Since it was a "service" (we connected to a TCP port), the underlying tricks they used to make it "fast" were transparent to us... until it was not.
They highly tuned their GC to where there was seldom any GC. One day, the third-party made a change to a supervisory service to generate more periodic monitoring emails. The file handles from this actor were apparently not GC'd and held by the process until the system ran out of file handles. That made the service stop working properly.
But, in addition to alerting, that service had a more important job: it was a post-trade risk "watchdog" to the pre-trade risk gateways, which we sharded across.
The various pre-trade risk gateways, upon not hearing from the watchdog authority then 1) began cancelling outstanding orders and 2) not allow new orders. We saw this happening haphazardly over 10 minutes and had little time and capability to recover; this also happened near the end of the trading day when some orders (MOC, LOC) are not cancellable. This was across our whole organization so it affected many strategies. So it wasn't as simple as "stop all" and although we had many checks and recovery procedures, this was a pretty special Chaos Monkey.
We ended up with a >$1B basket of random stocks that cost >$10M to liquidate over the next few days.
$10M evaporating in 10 minutes because of "GC optimizations" and poorly considered OS settings.
The #1 risk in algo/HFT trading is not financial, but operational.
[Some of the technical details may be slightly off, as it was third-party; my outlook is pieced together from post-mortems. Also, no other party (e.g. broker, SIPC, market participant, the third-party) was financially affected besides our firm.]
Also ZGC / Shenandoah can help a lot otherwise
Certain things may just be "table-stakes", but for many strategies table-stakes is the only requirement. Simply being "fast enough" may be fine if you have a smarter model, exploit niche opportunities overlooked by others, or are willing to shoulder certain risks that other HFTs are trying to offload.
To take this to the extreme, look at Renaissance's Medallion fund. It's certainly the case that much of what they're doing is "HFT-ish" in the sense that they're executing high turnover, high Sharpe strategies with short holding periods. Yet they've managed to continue to be successful from the 90s well until the age of HFT, without competing very hard on latency at all.
Markets are an ecosystem, and like biological ecosystems there isn't just one trait that predicts dominance.
you may need to rapidly adjust the expected fill price at a high frequency on a multi leg strategy being run in 100 positions, but the actual sending orders to the exchange doesn't need to be high frequency, and it is an important distinction that you wouldn't be competing with others on the frequency at this stage, it doesn't matter if this occurs in 1 millisecond of 500 milliseconds, even a couple of seconds. Very different ballgame than the wishful femtosecond game.
its probably better that people don't understand that, but it is annoying that there are this many roadblocks to a nuanced conversation
either way your servers still have to have all the authentication code, and various algorithms running to monitor the tape.
Another thing is the structure of Equity and Derivative markets vs FX. The former is fairly standardized. You'll need infrastructure at a handful of exchanges in 2 or 3 cities. Microwave networks and FPGA's are the norm.
But for FX markets, where these guys trade, the picture is much more messy. There are countless places to trade. Co-location means different things at all of them. The entire structure is much less regulated and understood. As a result, strategies and trading system architecture looks much different.
They built the Disruptor data structure around 2011 for their high performance financial exchange on the JVM: https://lmax-exchange.github.io/disruptor/files/Disruptor-1....
I used the Disruptor at a smart grid startup in 2012-2014, after LMAX open sourced it.
Martin Thompson has a lot of interesting presentations on the concept of mechanical sympathy:
- https://www.youtube.com/watch?v=929OrIvbW18 - Adventures with concurrent programming in Java: A quest for predictable latency by Martin Thompson - https://www.youtube.com/watch?v=03GsLxVdVzU - Designing for Performance by Martin Thompson
Market data is the slow monster - especially in the options market - not order entry. This is where the fpgas come in to help.
Often I have seen trading groups throw money at infrastructure when it was really just that a competitor (a virtu or citadel) has a better way of internalizing order flow or managing risk so they can pay more for the same trade.
PHP or Perl are called interpreted because the interpreter (installed on the destination machine) compiles each line of code as it goes.
Yeah, that hasn't been true of most "interpreted languages" in decades. The most common interpreted languages that I can think of where you parse as you go are shells like bash.
Languages like Perl and PHP are called interpreted because an interpreter runs the script. But the interpreter actually compiles the code first and then runs off of the compiled representation. Which is byte code - just like the JVM. If you put energy into it, you can use all the same techniques that the JVM does. As a practical example, both nodejs and pypy (a fast variant on Python) have a JIT. Julia regularly hits speeds for numerical calculations that is fully competitive with C++ and FORTRAN.
if you've got a script the likes of:
#!/bin/bash
echo 'Hi!'
sleep 20
echo 'Hello!'
... and runt it... and while the script is executing the sleep line, modify it so that instead of 'echo "Hello!"' it does something else, like 'echo "Oops!"'...
... once it resumes from the sleep, it'll execute the updated code.
Don't ask me how I know... I just wish it were more well-known.
Don't modify shell scripts while they're running!
On a common PHP setup (with OpCache today, any predecessor in the past) this happens only once on startup, thus is neglectible. (Java's VM and JIT are faster, no doubt, but where the translation happens matters only a little ... especially since Java does optimisations and JITting on the VM at run time only as well)
From your link, The Zend Engine compiles PHP source code on-the-fly into an internal format that it can execute, thus it works as an interpreter. In other works at startup itreads code, compiles to an internal form of byte code, and from then on it executes that. So it works as an interpreter (you give it code, it runs). But internally it does not go line by line as you go.
As a simple demonstration, a syntax error on line 10 of a PHP script will keep it from executing code on line 5. By contrast a syntax error on line 10 of a Bash script will not keep it from having executed line 5 first.
Of course because it’s also very dynamic, the lines are blurred, imo
When C++14 came out the project was ported to C++. It needed to go fast, but more importantly it needed to be stable and readable. An error in code could cost us, which wasn't something we could accept. Moving to modern C++ wasn't fun, but the code is clean, easy to read, works and works well.
Today, the code base has started to be ported to Rust, but unlike going from Java to C++, this move isn't as big of a deal, but the added safety is appreciated given the nature of the project.
Today it's pretty common for firms to create their own lisp like or Haskell like programming language. By having it FPP it can mirror mathematics a bit closer which reduces bugs, yet by having it in house it can be fast, otherwise these companies would use a common language. We are the only ones I know of moving towards Rust. I imagine a lot of firms who are aready C++ are going to stay with modern C++.
I mean, if you leak memory or have a use after free issue, it's not the end of the world since this isn't exposed to untrusted input.
Though, I'm sure the language features like async/await and safe concurrency are probably appreciated.
>Does rust buy much for HFT?
Going from C++ to Rust is low priority. It may not ever be finished. That low of a priority. It's more about refreshing the code base.
One thing that surprised me is Rust can be faster than C++ at certain tasks.
The big benefit is if you have a new dev coming on and they try to modify something that would introduce a bug, the compiler is all over it. Currently acceptance tests are all over it and work fine and C++ can be pretty strict so it works, but Rust is more ideal for this.
A desk trying to gain first queue position at the opening auction is probably running in FPGAs or ASICs and aiming for 100 nanoseconds. A spot market maker needs to respond fast to new level formation, and is aiming for 10 microseconds on a C++ stack. An alpha-driven liquidity taker may be using complex ML models, and is running a Scala stack at 50 microseconds. An ETF market may just need to re-mark their quotes when the underlying shifts, and is fine with 250 microseconds on OCaml.
You're taking real risk, in the sense that your buy order could get filled, your sell order doesn't, and the market moves down, so now you lose money. Latency is pretty important to market making for two reasons. One is because you want to quickly adjust your quotes before the market moves through you. If you're too slow, opportunists will pick you off in the wrong direction. Two is because exchanges mostly follow price-time priority. E.g. if me and you both enter buy orders for $0.10, the front of the "queue" to get matched, will be whoever's order arrives first at the gateway. Not only is that important for getting more opportunities, but generally orders in the front of the queue are less likely to be trade against "toxic flow", where the market moves through you.
There's definitely a baseline minimum of latency that you need to play the game at all. But there's tons of ways to compete besides just being the fastest. You can build your own signals to internally predict which way the market moves before it does. You can have superior models to tell you when the risk of "toxic flow" is lower. You can use clever tricks to get to the front of the queue when others aren't paying attention.
The industry has consolidated and any player that didn't adapt to the new, more cerebral game is dead.
Except for the one that won the speed war, and keeps investing massively to remain 10ns ahead.
The article focuses quite heavily on Zing vs Hotspot but it'd be interesting to see an analysis of a variety of the standard JVMs GC methods (namely Shenandoah).
For anyone interested in low latency Java I'd recommend watching some of Martin Thompson's talks on building LMAX and his blog Mechanical Sympathy is a great start too.
It also uses sun.misc.Unsafe to do the latency critical aspects, so yes it's Java, but most certainly not vanilla Java.
It is hard to imagine the same performance on a Pentium Pro.
What is the right processor-independent measure for latencies like this? Number of clock-ticks? That seems flawed, too.
Here is a more in-depth article on how it works, using pointer coloring.
Also, ZGC, which is considered production-ready in JDK 15.
Software security matters, in the abstract, but trading firms don't care about it nearly as much as tech companies do. The other thing is that being as fast as C++ is not compelling enough a reason to replace C++ for their use cases. If anything, Rust's dependency management story would be the thing to highlight (this is part of what's compelling about the JVM).
I think the responder's version of 'safety' really boils down to 'null pointer safety' - not 'software security'.
And even then, people keep talking about it like it's 'the thing'.
Java has null 'references' as do most languages frankly and it's just 'a thing' almost never 'the thing' to be concerned about.
'What devs want' is basically something kind of like Java, but that compiles, predictable/controllable performance and memory management i.e. no GC. That's it. It will probably end up being Rust, but that's because Rust will eventually provide all the nice, clean, modern package management, idioms, libraries etc., not specifically because of the 'safety'.
Granted I don't want to diminish that in the attempt to create 'proper memory management' you probably end up writing better software anyhow.
That being said: Jane Street is pretty avant garde in this respect. I expect there already is, or soon will be, a successful trading firm which likewise builds its tech brand on Rust and alternatives to Rust standard library primitives. But overall adoption will probably continue to be anemic.
If you're interested, check out https://www.youtube.com/watch?v=10gSoVZ5yXY for an example of the types of compile time guarantees can be had.
I am aware that dependent types exist, but quite frankly I've yet to see a popular/ergonomic language implement them. As much as we'd like absence of null pointers, error types, and immutability to be the norm it is not and it will take decades until it is.
with a modern JVM you can even disable it entirely (Epsilon collector)
I saw people comment on the "fast<->smart" continuum, and in this context, I believe they mean smart=computationally intensive. (An extreme example is a stat arb shop running its portfolio optimizer.) But there's another way to gain an edge, which is to iterate quickly on your "fast" algorithms and develop them so they take advantage of opportunities that are only partially exploited by competitors. Java seemed pretty good for that purpose for about a decade. Things like JVM warmup, GC, individual ultra-low-latency responses whenever necessary, were all dealt with after the fact.
This was a pretty reasonable choice at the time, as this was long before C++11 was around, and Java just offered a lot of advantages for developer ergonomics.
A natural consequence of this is that a number of HFT shops that were seeded by ex-ISLD devs just fell into using something like the ISLD tech stack, which at least at first involved a lot of GC-free Java.
https://blogs.oracle.com/javamagazine/the-unsafe-class-unsaf...
https://josh.com/notes/island-ecn-10th-birthday/ https://josh.com/notes/island-ecn-10th-birthday/ISLAND.PRG.T...
"Levine found a simple but hugely consequential 'trick' to speed matching. In Levine’s later paraphrase, what the 'enter2order' procedure did when the Island system received a new order was to: 'See if there was a record from a recently cancelled order that we can reuse for this new order. This is hugely important because that record will likely still be in the cache [fast internal memory] and using it will be much faster than making a new one. (Levine, n.d.)'"
http://www.sps.ed.ac.uk/__data/assets/pdf_file/0003/97500/Is...
At one firm I used to work at, the spread was several orders of magnitude. On one end, people were counting nanoseconds, and even C++ wasn't fast or predictable enough. At the other extreme, some teams didn't care about anything finer than a millisecond, and a big chunk of the stack was written in managed languages. It all counted as high frequency trading.
They used a combination of Python and highly GC-tweaked Java.
This was six years ago, but I see that similar techniques are still used to this day.
I also worked on HFT software written in Java (not the one in this article), and the mindset was completely different from other industries. F QA, F unit tests, F your F-ing AbstractProxySingletonFactory, just write the code that makes it work, see you after lunch to deploy.
Yes, in a lot of ways it is awful, but OTOH it is really great intense bootcamp to learn how to write very modular, easy to change and to read code. No time to over engineer, no time to build these fancy generic abstractions to support the one use case you have and the dozen hypothetical use cases product never asked for.
At that point the technology organisation starts taking by over, the desk developers are let go or they move into technology and start getting paid by technology instead of by the business. Things are now done in a more normal enterprisey software development way, and things are generally done more safely, with more design up front, and more documentation. Of course it also takes a lot longer to get anything done.
So traders decide to hire a bunch of developers, answerable only to them, and back to the start we go...
Specifically, the C4 looks very interesting, although I'm a bit perplexed because, if it was universally good as it's written, they would have essentially solved one of the major problems not just for Java, but for a large amount of languages.
Here is a good article on them. https://blogs.oracle.com/javamagazine/understanding-the-jdks...
In a way, low-latency GCs like C4 and ZGC have done that, but implementing them requires a large and experienced team, and not many languages even try to target high performance to begin with.
I'm not saying PHP executes just like Java, but PHP is more like Java than he realizes. PHP does not execute line by line like an interpreted language. But takes the source and compiles it down to an intermediate representation (IR), just like Java. Java calls it bytecode, PHP calls it OPCodes. Also like Java, PHP then takes that IR and executes it on a VM. Java's is stack based and PHP is register based, if I remember correctly. Many PHP accelerators / caching solutions save the IR and thus, the next time the code is needed, no compilation of source code is done.
The actual merit of the competing technology is simply not part of the evaluation. This is at least a step up from government procurement processes where the procurement officer, who has zero technical knowledge, procures based on a checklist he or she is unable to correctly assess bidder responses to, and of course price.
Also, in Java it is easy to end up with resource leaks whereas PHP always starts from a clean slate for each request. Way less headache for operations.
With a PHP application, it's easier - spin up the new Docker container, disconnect the loadbalancer, do the schema change, and connect the loadbalancer to the new Docker container.
That assumes you're running a new separate PHP processes for each request and an application built this way is unlikely to scale terribly well.
No. PHP always wipes the complete execution stack even in a long lived execution model like FPM or Apache mod_php.
When i finished my theoretical Physics PhD, HFT was THE way out of academia for making money in my field. I didn't follow this path as it felt at the time a pretty evil thing to do.
Does HFT provide any benefits for our society?
Evil is everywhere, really. Our brains are not made to really understand these things well. The easiest thing is to resign yourself to at least make you and your familys life easier. Which is acceptable enough.
i know, i live 10km from Nestlé HQ. and the banks in my country manage a significant part of the cash in this world.
> Our brains are not made to really understand these things well.
Agreed. I was recently reading several text saying that our brains are adapted to small hunter-gatherer communities of 300 ppl max.
> The easiest thing is to resign yourself to at least make you and your familys life easier.
Not now i know how the Earth System react, because a look at the latest IPCC projections doesn't make me want found a family.
> Which is acceptable enough.
Listening to some climatologists, there is a chance i see a world with less than a billion person from my own eyes. My brain cannot accept such a perspective acceptable (without actively fighting it at least).
But there can be indirect value in working in HFT, just like there is with other demanding jobs. There's interesting systems and ML research going on in the field - that's knowledge you can take with you and share with others. Some time ago, HFT was driving a lot innovation in networking, infrastructure, and systems.
Also, over time I've come to realize that "providing value" is really difficult to quantify. Does working in Academia provide value? Maybe if you get lucky and publish something that actually turns out useful. The majority of academics publish papers that are nothing more than noise - negative value. Does working at FAANG provide positive value to society? Again arguable. You may inadvertently end up promoting ads and misinformation. Not everyone does - but it's not guaranteed that you don't without realizing.
"Providing value" is subjective and most of the time it's nothing more than a story people tell themselves so they can feel good about themselves.
I don't think your comparison to academia holds up there.
As I mentioned, I'm not an economics or finance expert, so I'd be interested to hear what more experienced people would have to say.
At a short time scale, yes there tends to be a winner and a loser, but even then it's not so clear cut. Most of the trades happening at a short time scale are institutional rather than retail (i.e. people who have the risk tolerance for this type of trading & who have profited enough from it to keep doing this type of behavior). The activity of most HFT firms is some combination of market making & arbitrage, where the profit can be thought of as a fee for trading for other participants (e.g. Robinhood doesn't charge commission but sells order flow, which tends to be cheaper than commission per trade). In general, price improvement is also possible since these make markets more efficient, so as an example it would not be unusual to see firm A sell to firm B who sells to firm C, where both A & B see a profit from their sales. This is possible because prices across markets are not automatically synced, so A has access to a better price than B who has access to a better price than C. And because institutional trading dwarfs any kind of retail trading, the trades can appear high frequency from the perspective of the HFT firm while only occurring, say, once every six months from the perspective of the retail trader.
I can buy the idea of high frequnecy trading serving as market making even if it is not something I can fully appreciate. But it still seems like these high frequency traders (and others invovled in investing based on fluctations) are mainly skimming money from the market without providing value.
For example, a trade where someone needs to convert his assets into cash due to a family emergency benefits both sides. The person with the emergency takes liquidity from the market and pays a premium because the trade is time-sensitive - he needs cash the next day. Other liquidity traders may profit from such "uninformed" flow in the long term, but both parties are happy because they got what they want.
Another example is trading off risk and hedging against certain changes in the world that would affect you.
But you know, i'm more willing to write application for Earth Sciences (like im doing now) and living a misery than transferring cash & stocks with obscure algorithms..
That might just be stupid, maybe you're correct, that's just a story i tell myself.
When there isn't enough volume HFT strategies tend to not run, and so to give an example, my first time trading options was an iron condor on /GC (gold). I bought at a reasonable price, but when I went to sell, where I would have netted a cool 50%, because of low volume and I didn't know to set a limit order, I lost about 50%, so even if my trade was "successful" I lost. If there was a bit more volume HFT could step in and we could have normalized prices where limit orders are not required. On the other end, without HFT, it's easy to play that game and setup orders of opportunity where someone else might mess up and you can make a pretty penny, not that I've personally tried this.
When someone can profit from the same decision everyone else would make by doing it 0.000000012 days more quickly, IMHO he’s exploiting a design flaw in the way the market clears. That’s amplifying noise rather than signal.
i understand any civilisation needs funding to grow (do we need growth is another core question), and possibly market to do so.
yet any meaningful decision won't be made at the millisecond scale, sure you can train a model to learn short term rules and generate long term profit. but does that profit society as a whole, i doubt it.
If we’re going to paint anyone as “real” villains in the world of trading then I’d argue it’s the informed big block traders. ie hedge funds, banks trading exotics with a high barrier of entry, etc. At least MMs are largely market neutral.
The sharp v recovery in the markets was surprising, but not so much when you realize what's backing it. People gave credit to the wrong group. They said it was the fed, it wasn't, it was the HFT. All the fed did was add some zeros to accounts and relax some interest rates. HFT had to figure out where the money goes.
Without HFT, it likely could have taken a year for the market to recover. It would have taken forever to unravel all the changes (some ephemeral, some not) and figure out how to efficiently allocate all that money.
If I was the democrats, I'd find a way to connect the market to main street rather than try to create ftt. I believe the first country that does this will win the 21st century.
Imagine being able to come up with a business idea, partner with some friends, do a bit of proving out and publicizing your team and money just flows into it based purely on the merit of your real potential. No politics, no long lunches and dinners with VC, no irrational idiots. Just pure, unfettered meritocracy.
Or imagine having a long standing mainstreet business that gets hit with a one time, out of its control, disaster. No need to go the bank with hat in hand, the money just shows up in your account, with hyper competitive interest rates. The same could be said if you have an idea for expansion or efficiencies. Just send out a press release with a few specifics, and next day you are underwritten.
SPAC is sort of like that. Sadly, democrats have their own brand of xenophobia. They tax what they don't understand. It's very unfortunate that the GOP didn't vote with the democrats to impeach Trump. Historically, I think it will go down as the most colossal political mistake since Nixon.
One way to think about all this, is DeepMind / starcraft. AI could already out-micro any competitor, but now it's getting to a point where it can out-strategy as well.
Having tools like that to manage our markets is a very powerful tool and makes the US not only hyper efficient but a far more competitive than any other nation.
That said, the whipsaws of computers figuring out things too fast can be tough on people. pure, unfettered meritocracy can be hard for those without merit to offer.
Yang and his UBI isn't a bad idea.
Where it provides less value are traders who intentionally inject volatility. Some guy went to jail over that (flash boys)
Unsophisticated investors only using market orders will also pay less spreads than they used to.
Going out on a limb, I think the losers are mutual funds and similar who are less able to monetize their fundamental analysis by buying or sell big chunks of stock without moving the market.
And FPGAs are the norm now, not the exception.
Sure, Java still has its place, but implying that it is anywhere in the hot loop of a contemporary ULL system is wrong at best and disingenuous at worst.
I mean just read the abstracts from this years STAC conference. https://www.stacresearch.com/spring2020
There are simply no material external gains from HFT etc, it's completely zero-sum other than maybe a 'job creation program'.
The site ridiculously talks about 'levelling the playing field' by extending HFT opportunities to other players, which is rich, considering the playing field could be 'levelled immediately' by some really basic regulation and a few smart queuing algorithms for placing orders.
Most traders I've met have been fairly honest about contributing nothing to society, or that the world would probably be a better place is large segments of the finance industry disappeared tomorrow.
I can't say the same for the majority of developers from Facebook or Google I've met.
I run ads and campaigns for my company and they are an absolutely essential form of communication.
... as almost everyone running and actual business knows.
Yes - it can be very inefficient - but it's not inherently wasteful.
You're making the wrong comparison - the 'Digital Ad Industry' could be compared to 'Digital Banking Industry' perhaps - i.e. useful overall but deficiencies exist.
HFT is one of the ares in which there are no reasonable externalities.
Better/faster tech for CDNs etc. can be developed without the HFT industry.
Running an ad campaign, or having to 'get the word out' for something possibly meaningful can be an enlightening and transformative experience for so many people who live adjacent to, but otherwise 'totally outside' the ostensibly opaque world of marketing.
It's one of the most stupidly represented issues in pop culture, like 'reefer madness' memes form the 1960s' where 'smoking weed will make you gay!'. Well - just 'smoke weed' once and you realize it's frankly just not that big of a deal.
Ads are ancient, normal, universal. The industry has serious efficiency problems, we're starting to face privacy issues (which are overstated but nevertheless legit), and I don't like some mediums (I for one, think we should dump most billboards) - but the nature of the industry is not wrong or evil, it's actually beneifical.
'Good ads' are a win for everyone. HFT is only a win for those directly involved and it's zero sum.
Can you prove how? Thats what the previous poster was asking.
>'Good ads' are a win for everyone.
Again, how? I don't like ads. I think they're annoying. They're not good for me.
I think you're drinking the koolaid.
The ad tech in question is the issue; the data collected, the platform itself. Think Facebook - in pursuit of creating a platform and getting data that allows for very specific targeting, they have created bubbles that magnify extremist messages.
For your company, not for the World.
Business absolutely cannot exist without advertising. The need must be filled one way or another.
The same cannot be said about HFT and stock markets in general.
But advertising is not negotiable, it must exist in one way or another if the economy exists.
If you want to go down this road, argue about the "how", not the "what".
The idea of banning advertising outright is too silly for words.
The biggest manifestation of this is that FPGAs really struggle with sizing. Say there's a given market event that produces a very obvious opportunity, but it's not so obvious how big that opportunity is. Archetypical case is when the price changes, and there's a new level formed on the book.
It's almost certainly the case that adding an epsilon of liquidity at the first position of the queue is profitable. What's harder to determine is where the breakeven point is. Should you add 1 lot? 100 lots? 10,000 lots? That's a lot more complex because you generally have to be sensitive to the typical sizes in that particular market and that particular time, as well as how strong the level formation event was. E.g. was it in response to something that looks like a very big order, or just some random liquidity depletion.
What happens is that the FPGAs know they should quote an epsilon, but can't really confidently quote anything bigger than that. So if the right answer is a thousand lots, the FPGA will quickly capture first queue position with something like a hundred lots, then the software-based players with take the rest of the 900 shares for themselves.
What you see with FPGAs is a segment that makes very consistent, very good money relative to the volume that it executes. But fails to capture the sizable majority of the market share.
The majority of programming tasks don't need anything better than Java, and don't need anyone more expensive than the typical Java coder.
Low-latency trading is not one of those tasks. In low-latency trading, as in poker or any other zero-sum game, you either rip the other guy's heart out and stamp on it, or get your own heart ripped out and stamped on. Using any but the best tooling and programmers means you will be that second guy.
So, anybody serious is using C++17, coded by people who understand cache interactions, branch prediction, TLB shootdowns, and core isolation, and who avoid lockless queues because of the cache-line pingponging. If you don't know what all that means, you are not the right person to code low latency trading.
They are not the hardest core: the sharpest players farm the simplest, most time-sensitive operations to FPGAs. But FPGAs take tradeoffs that make them not right for everything, and FPGA coders are not just programmers.
This statement as a definitive "this is how it works" makes me think the person who wrote this doesn't really work in anything at the intersection of networking and HFT.
Much of the serious HFT for transoceanic stuff at low bitrates, but very important signalling messages, moved to shortwave radio with giant-ass aimed yagi-uda antennas several years ago.
You can go out to some places not very far from the CME datacenter in Illinois and find big, suspicious shortwave band things aimed at Tokyo.
I get a very real feeling that the people writing this article have no idea what they’re talking about.
Take a look at https://github.com/OpenHFT to get a taste of what Java + HFT looks like.
- Until late in it's lifecycle, .net runtime performance was worse than JVM.
- Also, the Windows TCP stack, re-architected with Vista, had severe latency issues, which probably led to wholesale abandonment of that platform for low-latency applications. Some of these problems were fixed in Windows 8, but it was probably too late by then.
(These two points are based on my own experience with Windows/CLR/JVM in the mid/late 2000s and some anecdotes I heard second-hand. So I can't really cite any public sources for them)
I'm sure the fact that you can customize Linux to the extreme and the lack of .Net on Linux in those days probably had a lot to do with the choice of Java, in addition to the above.
However, based on my reading of blogs like Mechanical Sympathy [1] a while ago (~8-10 years), there does seem to have been some demand for .Net-based solutions in the HFT space. They had code for both Java and C# in their public code.
However, we moved away from “pure C#” about 10 years ago, to use FPGAs, especially in the ultra low latency stuff.
A lot of strategies can still be profitable with C#, colo’ed servers and kernel bypass NICs.
I'm a bit surprised by all the chatter about garbage collection because it's not as much an issue as it sounds.
The thing is not to do allocations (i.e. GC) on the _critical path_.
In a trading strategy you react to a stream of market data. It appears constant to the human eye, but in fact it's a series of "blips", sporadic events.
When you get a new message you need to react as fast as possible. The fastest code is, of course, no code, so it doesn't really matter if it's not written in C++, C# or PHP. Here you're just going to do the very minimum required. You will find that generally there is no difference in the code compiled whether it' done by the C# JIT, the Java JIT or a C++ AOT compiler, because the code is simple: you read a few fields from the message and select a pre-computed outcome to decide whether or not to send a (buy|sell) order.
Now, after that event has been handled (and perhaps you sent an order), you need to recompute stuff. You generally have plenty of time (a few milliseconds) before the next event arrives, and that's where you want to do all your fancy calculations, and here a high-productivity language like Java or C# helps, just because it's quicker to iterate and deploy modified trading strategies (typically deployed daily, or at least weekly).
Many people assume that you need to do all your fancy calculations very rapidly right after a message arrives. But you don't have to because the message is not completely random. The order book is not going to change completely after one message. You can pre-compute a few likely scenarios, and when a message arrives and confirms one of those scenarios, you just mechanically go with the precomputed result.
Ideally, you would want to have the benchmark developers cooperate with the team creating the JVM you are evaluating in order to be as far as possible with the benchmark sort of like how the TPC benchmarks are done.
Both issues related to low latency metioned here (JIT and GC) can be solved with very simple and common methods with the JVM provided by Oracle.
We had similar issues in my old company, where we implemented a graph database in Java. It was very slow and somtimes freezed for alomost a minute (512 GB heap).
Here are the strategies we used for solving these issues.
- performance - we built warm up logic in our app, which called the most time sensative parts on startup. It delays the start up process by 10 sec, but after that the application is very fast.
- GC - we pre-alocated big chunks of memory outside of the heap and put there the 'heavy' objects - those which require most effort from the GC. We also optimized the code so that all other allocations are in the stack (escape analysis). This requires some discipline, but eventually we ended up in situation where there was nothing left for the GC to do.
I'm not sure if it's worth paying hunderds tousands $ in license fees for something that can be achieved with good programming practices and discipline, which most companies need anyways.
Something about warmup: this is an area where JVMs still have catching up to do versus JavaScript VMs. JVMs have suboptimal startup time in the non-cached case (I.e. not what they are doing with Zing) because it’s just not been an optimization target for JVM hackers. In the JSVM world, we were forced to optimize for this and optimize we did (otherwise page loads would suck). I suspect that Zing with the ReadyNow thing is way better than normal Java, but I don’t know if it’s better than what JSVMs do. It’s probably still way worse if you’re not cached. It might be better in the cached case but then again JSVMs also save some profiling in the cache (JSC definitely does). That suggests that JVMs can probably improve further. No reason why Java - an overall very performant language - should start slower when uncached than JS. I suspect this will get fixed since AOT is a thing now, but getting start times right is about way more than just JITing less.
It would surprise me more to find an HFT that doesn't use Java at some part of their execution lifecycle. Not everything needs to be in microseconds, nanoseconds, or picoseconds, especially if you can factor it out of the critical path and do the work/build some data structures ahead of time.
Are GPUs really any use here? Aren't they quite far from the NIC?
I mean yeah those GC pause times are extremely impressive, but you're clearly operating against a different set of performance requirements than those teams that care about several microseconds of difference and it's a bit disingenuous to compare your own tech stack to them.
That being said I do appreciate articles like this that show just how far Java performance has come over the last few decades.
They used some of the intel instructions meant for running VMs to get fast concurrent GC. It’s probably a lot easier to court customers with a software only solution than with custom hardware they can’t really repurpose...
I'm curious. How much heap memory do HFT applications typically allocate in a day?
I'd expect them to be doing most of the work, or at least most of what actually needs low latency, in locally allocated primitives which would be on the stack. If that's the case, you might be able to turn off garbage collection and just stick enough RAM in your computer that it takes a day to fill up with garbage. Once a day, reboot.
Assuming the code can be made not to allocate much, this approach could probably be made to work, [0] but I imagine it would make more sense to just configure the GC to run once a day.
Alternatively, you could have a pool of Java servers and coordinate GC with your load balancer.
That is, some system would:
(1) coordinate with the load balancer (or directory service of endpoints) to ensure that a Java node is completely out of the pool
(2) coordinate with the Java node to ensure all its in-progress requests have completed
(3) run a full GC on that node
(4) return it to the pool
(5) move on to the next node, cycling through them.
You probably wouldn't need to increase the number of Java nodes very much, either. If a GC is required every 10 minutes and it only takes 30 seconds to do this, then you only need 5% more Java nodes.
[1] I'm aware of many other benefits, but this is the one most unique to Rust and the one I care about most by far.
https://github.com/ixy-languages/ixy.swift/blob/master/perfo...
I feel like they've buried the lead here. If the above statement is really as it sounds, I find that way more impressive than the fact that they were able to make java fast.
If you avoid overloading yourself with unnecessary complexity, and you as an engineer spend time understanding the problem/business domain, there's no reason why this speed of turnaround should be unattainable.
I have frequently had the experience of sitting in a meeting with (internal) customers and/or a product manager, and shipping discussed changes during that meeting.
If you can't do this, it's because you've put barriers in your own way. Ask yourself if they're really worth it. (They might be!)
I suspect there are some significant caveats to the claim, and even considering that I still think that's a far more impressive accomplishment than just making java fast enough.
In a sane market, there would be no advantage to making an investment and then selling it a few milliseconds later. It introduces unnecessary volatility and incentivizes insider trading and other illegal activity[1]. It defeats the original purpose of the market: to make long term investments. It also leads to some unbelievable waste:
Traders’ need for speed has grown so voracious that two companies are currently building underwater cables (price tag: around $300 million each) across the Atlantic, in an attempt to join Wall Street and the London Stock Exchange by the shortest, fastest route possible. When completed in 2014, one of the cables is expected to shave five to six milliseconds off trans-Atlantic trades.[2]
The problem is that our 20% of our economy is "financial services". It's doubled in the last hundred years or so[3]. Technology isn't shrinking this sector because the people who own it also own lobbyists who own Congress. They write the rules so they can play stupid financial games, and then force the government to bail them out when the latest ponzi scheme hits maturity.
[1] https://www.zerohedge.com/article/first-hft-casualty-finra-f...
[2] https://www.motherjones.com/politics/2013/02/high-frequency-...
[3] https://www.washingtonpost.com/news/monkey-cage/wp/2016/03/2...
>In a sane market, there would be no advantage to making an investment and then selling it a few milliseconds later.
And yet exchanges (NASDAQ, NYSE, CBOE, LSE, CME, Euronext, etc.) actually _PAY_ HFT firms to do this exact behavior since it provides actual, quantifiable benefit to the exchange and its customers.
>It introduces unnecessary volatility and incentivizes insider trading and other illegal activity[1].
The firm mentioned (Trillium) is not a player in HFT, they are a day trading proprietary trading firm where the firm basically gives individual traders (borrowed) money to use however they want. In this case, some of the traders decided to break the law because there is no real oversight at this kind of firm. When looking at reputable proprietary trading firms (i.e. the kind that don't just hire random people and give them money to play with) this type of behavior simply does not happen. I agree that these types of places are extremely scummy (and are probably responsible for a ton of insider trading crap & the like), but they are not representative of the HFT industry at all.
As for the comment on volatility, the idea behind HFT is actually to help reduce volatility. Outside of a few bugs (such as the Knight Capital debacle you mention) and the Global Financial Crisis, it seems like volatility is trending downwards in part due to HFT. In the same vein, liquidity has exploded thanks to HFT (which is why exchanges pay for HFT). It means that you no longer have to wait minutes or hours for your trades to execute and you can usually get the most up to date price very easily.
I don't know anyone who thought the Global Financial Crisis was a "bug." People went bankrupt, lost their retirement, lost their homes, lost their jobs. These are not small stakes.
> It means that you no longer have to wait minutes or hours for your trades to execute and you can usually get the most up to date price very easily.
Warren Buffet was able to build his empire before HFT, because he was investing, not trying to skim market moves. If the point is to have a stock market where people make investments, HFT literally serves no purpose. It's like getting a weather update every second. If you're investing in a company for its merits, those don't change by the second or even by the day.
1. The GFC which wasn't caused by HFT
2. Certain instances where HFT bugs led to "flash crashes"
>If the point is to have a stock market where people make investments, HFT literally serves no purpose
This is patently false. Imagine a situation where the price of AAPL suddenly plummets because of an influx of sellers. If you want to close a long position with AAPL to lock in your gains, there will simply not be enough buyers for your shares to sell most of the time. The purpose of market makers is to make this sale possible because they are still able to profit or nearly break even from buying your shares (and are often contractually obligated to do so). Without someone taking the risk that market makers take, the stock would just free fall indefinitely and nobody would be able to take a profit off their investment.
A similar situation is options market making. People/institutional investors like Warren Buffet hedge their bets by purchasing options to protect them when their expectations fail to pan out. In such a situation, it is imperative that options are able to be purchased for a fair price and due to volatility the price from 10 minutes ago is almost certainly wrong. Likewise, it's important that a contract even exists for your desired purposes.
The purpose of HFT isn't just to give the most recent update, but rather to provide the essential pricing of various derivative instruments which depend on immediate values like the price, volatility, etc. of the underlyings. The immediate ability for your order to execute also protects you from the market's volatility, since you can essentially lock in your price as soon as you see the market moving. The blatant truth is that many modern portfolios consist of more than just unhedged long positions in a variety of stocks and as such a lot more complexity is required than "investing by a company's merits," even in the world of "value investing."
I think HFT pretty much sets the bar in terms of cutthroat performance, so it's a incredibly interesting read.
It seems right on their wheelhouse: highly technical and makes a lot of money.
Does anyone which currently available CPUs are able to run at stably at 5.6GHz? (Presumably with 14 cores each)
I'm talking cash , not stock. From what I've seen above 200k is exceptionally rare.
Total cash compensation easily gets above 400-500k. People probably make less money at large firms, tho.
Would you consider your job high stress ?
Also working with some of the smartest people in the industry can be intimidating: if you're used to being the smartest guy in the room, it's very unlikely you will be anymore at a top notch firm. This is a form of stress in itself.
It's not stressful in the sense of shitty clueless bosses, long unpaid hours (people work long hours, but that work is often highly compensated for through pnl improvements).
Everyone I know feels extremely lucky that by just putting our heads together, we can mint money out of thin air. I don't think anyone I work with would rather be doing anything else. Trying to make money out of nothing certainly humbles you, but when something works, its the best feeling in the world.
Keep in mind that things might be different at a larger firm: lower compensation and upside, more hierarchy, etc.
Would it be easier to target bigger firms and then get into or start my own firm ?
Just apply for roles, if you're a good dev, I don't think you should have trouble landing interviews.
I do a lot of intern interviews and I haven’t seen a single non-target resume, or someone without FAANG internships, high placement in global hackathons, math olympiads, etc.
Experience is a way in but it depends what “fintech”.
e.g. more complex algos looking at the relationship between securities
> Anyone can trade frequently.
lol, fair
Oof this didn't need to be added to the end of the article.
I worked as an offshore employee for E*Trade back in 2010 - 2012. Their desktop app for HFT was written in Java Swing and their web client too was running on the JVM but the 'services' layer the one that interacted with the database was written in C++.
Four years ago heard that they were re-writing the desktop app in JavaFX.