Achieving super low latency for critical real world internet applications (2018)
netdevconf.org
netdevconf.org
There's problems with using the internet if you want low latency. Not just low latency, but also low jitter, ie consistent latency. You're sharing a bunch of infrastructure with other people, so sometimes your packet will have to wait. Try to ping some internet server a few times and you'll see it changes. I've done this with rented lines, and the number is the same every time.
Not inherent to the internet is the fact that REST seems to be a standard. Probably people will want a REST interface. So that means serializing and deserializing messages, and the messages are always going to be a lot longer than they need to be. It's not that you have to do it this way, though. There's also the fact that internet architectures tend to have a bunch of hops (load balancer / gateway / etc), which isn't going to help, but definitely something you could work around.
The figures he gives are interesting. If you really need under 20ms for AR to work, how is remote surgery going to work? 20ms is close to the time you get to render a frame in real time games at 60fps, seems a bit fast to me. The surgeon can't be far away if this is true.
For software, everything is done in house. As you said, it doesn't have to be done a way or another.
For the shortwave links in place between NY-Tokyo, for several I'm privy to details of, people have written their own custom dedicated purpose protocols that implement very low bitrate, ultra low latency data messaging with less overhead and ASIC packet processing time.
For what I do, Ethernet is acceptable as long as the latency is below 1 ms. Here're some quick stats for what I consider good enough this specific usecase:
rtt min/avg/max/std-dev = 0.069/0.095/0.121/0.016 ms
Generally it is better to keep things simple and use existing stuff whenever possible.
If we are talking about very low latency (say below 30 us) you may need many other things first, like a good timesource.
Then indeed, when you start rolling your own stratum1 NTP servers, rolling your own protocols can be a sensible answer to the problem you have.
The end customer (whoever actually wants to buy stocks) doesn't need the price to update every millisecond. But no one has yet figured out a better way to run an electronic exchange than "first come, first served".
Source: also worked in HFT for a long time.
The actual people working for HFT firms are probably not going to provide much insight into what's going on behind the scenes, due to very restrictive NDAs, security policies and litigious corporate overlords.
I see PTP as a good complement to NTP.
Ethernet is still viable in the LAN. With solarflare/mellanox nics + kernel bypass you can get a message from wire the memory in around 2-3 microseconds. The layer 3 10/40g ultra low latency switch offerings can pass a frame port to port in 200ns while layer 2 offering can do it in 5-50ns depending on how many features are stripped away.
HFT firms use FPGAs with integrated nics as well. Pci-e Latency to pass a message is around 800ns.
For the ULL crowd, when the money opportunity arises, everyone is stampeding to that order to get there first and causes a micro/nanosecond level spike of messaging/activity. FPGAs are great at staying deterministic during this burst.
But alas, i digress. Ethernet is still used and can be wicked fast.
For AR/VR, you need to detect the motion, calculate the scenery, and put it on screen. This alone eats up a big part of the 20 ms budget. And to avoid motion sickness, the delay should be even shorter.
First Person View / FPV quadcopter pilots still use analog TV transmission as it has a latency of about 26 ms. As the 'dots' are sent as they are sampled by the camera, expert pilots can judge from distortions in a picture how the quadcopter has moved between two full frames. When I looked it up, the best digital FPV video systems offered about 50/60 ms delay.
Take timestamping for example. NTP has unacceptable jitter.
That’ll require exploiting edge distribution to its limits, with an entire new level of internet infrastructure that doesn’t exist yet.
5ms reaction time mentioned on these slides are far cry for what computers can do. trading servers reaction time packet-in packet-out even without fpga can take as little as 20 microseconds.
- cpu thermal profiles and power governors
- cpu affinity
- non-evented io(if you prefer latency over the number of clients served)
- network interrupt coalescing(actually, disabling coalescing), network msi-x, hardware offloads, 10gb fiber or DAC connections(very cheap nowadays)
- lock-free and kinda-crdt data structures(between threads)
- predictable memory allocations(better - no allocations on critical path)
- specialized logging and tracing(no stdio)
can drive you quite far.
There is a big crossover in tech with shortwave band military over the horizon communications systems, and also serious ham radio enthusiasts who use big yagis and antenna rotators to talk to each other from opposite sides of the planet.
Things like chained 6 and 11 GHz microwave between NY-Chicago, which was the "hot new thing" in HFT 7-8 years ago to achieve better than fiber latency on that route, seem quaint now.
For those use cases it would be better to come up with a full 3d (not just stereoscopic) streaming video format that allows for some local camera movement.