Parsing Decimals four times faster
cantortrading.fi
cantortrading.fi
Now I’m more than a little shocked multiple someones are publishing real time data in JSON for algos?? I thought this was mostly a solved problem with OUCH, ITCH, FIX SBE and friends.
But great to see this writeup, getting these details right can make a big difference.
Some exchanges do offer colo.
HFT on a traditional exchange will be faster but that's not the competition. The competition in crypto faces the same problems so you just need to be faster than them.
Of course, if the whole process has too much uncontrollable noise (jitter) due to cloud specific reasons it probably doesn't matter. I hope they managed to control this before doing this optimization :)
These sorts of firms are super paranoid about the exfiltration of proprietary IP. When you add trading over the Internet in to the mix, you just have more ways to do that. You have highly dynamic cloud infrastructure to deal with (whereas traditionally everything is static), DNS security, tricky PKI/certificate management, the need for much tighter access controls, firewalling (both at the network and application protocol level) and the secure logging and monitoring of flow between networks. And everything needs to be tight enough so no single actor (external or internal) can compromise the system.
I beg to differ. FIX could almost be a case study on how not to design a protocol.
Session management and recovery is outside the login protocol? Check! Can’t reliably recover from a failed login or wrong password attempt? Check! [0] Message boundaries are nontrivial to detect? Check! Field values are binary but cannot represent the byte 0x01? Check!
But FIX has a brilliant extensible hierarchical message structure that is more type safe than JSON, you say! Sure, kind of, but the encoding of lists of nested objects is utter nonsense. Not only is it somewhat redundant (the number of repeats is sent first, but the number of repeats is also determinable after parsing the whole thing, resulting in some awkward questions about what to do if the numbers disagree), but it has worse problems. Specifically, if the sender and receiver don’t agree on the precise set of fields and their order in each nested object, then the list cannot be parsed. This makes the whole concept of extensibility someone dubious. (This is because there is no delimiter between repeated objects; instead you notice that the field you’re parsing can’t legally belong to the object you’re parsing, meaning it’s time to use a crystal ball to determine whether a new repeat has started, the whole list is done, or an error has occurred.)
On top of all this, the standard implementation, QuickFIX, is a work of art, and not in a good way.
FIX SBE is IMO not at all better except that it can be serialized and deserialized faster.
[0]. No, really. If you try to log in to someone else’s session and the server rejects the password, you have just fried their session. The protocol cannot recover. Sometimes I wonder if this was deliberate, because screwing this up at the protocol level takes some creativity. Imagine if trying to log in to someone else’s email account caused their next login attempt to fail! Oh wait, some password attempt policies do this, but FIX does it by corrupting the whole protocol state machine rather than merely having a misguided policy.
I think FIX makes the right tradeoffs.
I don't understand what you mean regarding killing someone else's session. You can reset sequence numbers in the login message.
As for sequence numbers: suppose your session drops. You’ve received the message with sequence number 5, but you’ve missed 6 and 7. You’ve sent 10 and 11 but the other end only got 10. Now someone (your own misconfigured software or someone else) connects and sends Logon to the server with your comp id, with sequence number 11 but the wrong authentication data. What is even supposed to happen? The session now contains two different message 11s to the server. It’s broken. You can (sometimes) recover by resetting sequence numbers, but that defeats the entire purpose of having sequence numbers in the first place: the server was supposed to recover the original message 11 on a reconnection attempt and correctly handle it as a possible duplicate. (As a practical matter merely experiencing network trouble during a reconnection also seems to break everything, mostly because common FIX implementations are terrible.)
I suppose it must be possible to be so deeply entrenched in the FIX ecosystem that this makes any sense at all, but the whole protocol is backwards. You should connect, authenticate that connection, and then run the session protocol inside the authenticated connection. The Logon should not be sequenced — that gets the layering backwards.
https://www.eurex.com/ex-en/support/initiatives/t7-release-1...
LSE's platform is much more modern: GTP being pretty much the perfect example of how to build a modern super-fast protocol
https://www.lseg.com/sites/default/files/content/documents/G...
Future optimisations using better on-chip features, or other new techniques: new instructions that reduce the task might get used by the compiler but it won't update your section in hand-coded assembler.
Maintainability: you don't want to force larger parts of your workforce/community to be familiar with assembly if it is not useful day-to-day, or lock your assembly capable people to maintaining those hand-written-assembly portions.
Every x86 computer, incidentally, still carries vestigial hardware support for BCD.
https://stackoverflow.com/questions/33182491/why-bcd-instruc...
I'm personally surprised that this firm isn't using "integer # of exchange ticks" as their storage format for prices...
Using base-10 decimals is a good compromise between the perfect solution and something that 'just works'. It also means we don't have to go rewrite everything and instead can just do some optimization work when needed.