Serving 6.8M requests per second at 9 Gbps from a single Azure VM
ageofascent.com
ageofascent.com
"HFT Quant Trading Servers" optimize things down to the micro and even nanosecond level from the hardware to the bios all the way up to the networking interconnects and applications + c libraries. Optimization to this level is simply impossible from the cloud, where even with VT and SRIOV, the performance penalty is simply far too high (and measurably so!).
Disclaimer: I'm a technologist who has worked in the "HFT Quant / Electronic Trading" Industry the past 8 or so years. I literally build infrastructure like this for a living.
If the OP is making analogies that aren't comparable, I'd rather know than not so I can dig in with a healthy bit of skepticism.
PacketDirect is the next step for Server 2016 which moves closer to the NIC https://www.youtube.com/watch?v=KaXfDjIhn0U
https://www.youtube.com/watch?v=CJeWIWkhVow&feature=youtu.be...
The advantage of Solarflare is that it's easy to use, but at the scale large cloud vendors operate at (millions of machines), it's much cheaper to just have a networking staff that can properly operate another vendor's NICs.
I saw a few hundred Solarflare cards purchased, the prices were not that absurd. Only 10-20% of the total cost of the server, for a pretty amazing perf gain. Mellanox etc. tested similar but had less helpful sales engineers IIRC.
That adds up to a significant number, even if it's "only" 10%-20% per server, and you're probably overestimating the cost of the rest of the server since large companies are able to get steep discounts on almost everything in the quantities they buy in.
> Mellanox etc. tested similar but had less helpful sales engineers
We've found Mellanox engineers to be more helpful than Solarflare folks. I don't think that's because one company has inherently more helpful engineers. We're just not in a market that Solarflare cares about and you're not in a market that Mellanox cares about.
> if you actually care about latency, you do not use windows nor use the public cloud
You might be surprised by how many HPC shops have moved from fancy in-house infiniband networks and custom fabric to cloud hosted HPC clusters. HFT isn't moving to the public cloud anytime soon, but a number of large customers who care about sub-microsecond network latency have found real value in moving to the public cloud.
Sorry, I do not know anyone who buys millions of high performance NIC cards. Maybe 1 or 2 supercomputer labs in the world? And with that in mind, I am not sure the price break a company will provide to someone buying a million $500 network cards is super important to me?
> That adds up to a significant number, even if it's "only" 10%-20% per server, and you're probably overestimating the cost of the rest of the server since large companies are able to get steep discounts on almost everything in the quantities they buy in.
Nope, I definitely knew what the whole server cost.
I think you're responding to something that's not what I'm saying here. I never claimed that our pricing is relevant to you. Just that your pricing isn't relevant to us and that you're making generalizations that aren't valid outside of your niche.
Disclaimer: Author of article
gamers "flip their shit" when ping is not steady, and are only happy at ~50ms response times. a 200ms response in video games is usually considered "unplayable".
See my reference to SubSpace farther down, it's a highly skilled game that can be easily played on a ~250ms latency. The large majority of the game is predicting where the other player will be in ~0.75s and setting things in motion ahead of time to intercept them.
Any game based on prediction and designed for smooth re-integration of the game state can be done on latent connections. Heck there's even been some really impressive stuff in the fighting game space involving re-winding gamestate to resolve hits ~0.25s in the past(the original Counter-Strike does this as well, although not as well which is why you'd sometimes see people rubber-band back around a corner when hit).
HFT is increasingly harder to make profits in: the real growth is in algos operating in the low milliseconds range capturing real market activity and non simple arbitrage.
For example in the interest management; it is working out what custom messages to send to what players about each other and with 50k players in the same combat it needs to do this without it ballooning to 2.5Bn messages for a round of position updates - of which there are many per second per player; then convert these individual compressed output streams that are unique per player; and move on to the next set - and even then as stated in the article we are still at 267M messages a second.
I read that and was hoping they were going to start talking about ring buffers and event sourcing, but no such luck.
Cool that you're getting a ton of throughput but you don't use UDP for throughput, you use it so that the one dropped packet doesn't back up all of your time-sensitive data. You want to drop packets that are out of order since you're using dead-reckoning to keep the game state psuedo-in-sync.
Seriously, there's a reason everyone uses UDP(or IPX!) since the days of doom.
> For our high throughput needs, prevailing wisdom would suggest you need to write your server in C++, use UDP rather than TCP and run on a very high spec bare metal box running linux. We are already running the client in javascript; so if we ran over TCP (websockets), wrote the server code in managed C#, ran on Windows on a VM in the cloud – is this just madness?
(Disclaimer: Author of article)
I totally understand the browser constraint part but you're going to hit some serious issues once you start having real-world clients that have >3% packet drop.
It gets even worse when you start adding different discrete layers on a single TCP connection(chat, larger game state not related to twitch parts, etc). Each one of those channels has the chance to drop a single packet and then your whole world comes to a stop until you can roundtrip that single packet that's unrelated to the realtime info you care about.
FWIW most stacks end up implemented channels over UDP with unreliable+unordered, reliable+unordered and reliable+ordered for the different latency needs of the specific channel.
I've done a fair bit of this stuff in a production environment and have to echo my original comment. TCP is great for lockstep but if you're doing dead-reckoning it will fall apart on anything other than a LAN.
Edit: I also don't mean this as a dig, just that I've seen many teams go through this and it's usually at a point where it's too late to make any changes.
The other layers (chat, non-space game state) are handled by discrete microservices on other connections to other hosts - to keep the space flight data slim and connection dedicated in use.
Can always give it a try, our next public playtest is on Saturday 6 Feb 8pm GMT/UK | 12pm Pacific | 3pm Eastern US at http://www.ageofascent.com/
Just don't try it on a phone; give us a chance with the latency ;-)
Games like SubSpace/Continuum were doing 300+ player twitch based on UDP based connections over 56k dialup back in '97. Any sort of packet loss(and subsequent TCP exp backoff) is going to weak havok on your latency.
http://hanselminutes.com/509/inside-age-of-ascent-with-ben-a...
I wonder what the intersection of "people that want to play MMOs" and "people that don't want to install anything" is. There are good game streaming technologies (e.g. Guild Wars) that minimize the patch time, so it feels like a bit of a straw-man.
I was looking up the release date since my mind was a bit fuzzy on the year and it turns out it's been released on steam as freeware last year: http://store.steampowered.com/app/352700/
It's taking all of my self-control to not install it and waste the rest of the day.
Assuming you are sending a packet per frame, this setup means that you are only sending a packet every few seconds per TCP connection, which means retransmissions should have already happened by the time you send another packet.
Their homepage has a testimonial/review-blurb of the game as the very first thing from a Microsoft "White paper".
Below that there are 2 videos, one of which is an "Interview Rob Fraser - Microsoft".
And of course the article itself focuses on Azure with another 12 mentions of Azure and how it runs Windows, .Net and C# so well (I'd hope so).
Also, the 7 mill requests/second aren't real world numbers but just a software benchmark.
Oh and the benchmark was run on an instance that has 16 cores, 3TB of SSD storage, and 224 GB of RAM. That's more like a dedicated server than a single cloud server instance (and waay more expensive than dedicated at $3061 per month per instance).
https://github.com/benaadams/benchmarks/blob/managed-rio-exp...
https://github.com/aspnet/KestrelHttpServer/graphs/contribut...
I love that they could have the resources and time to explore using safer languages for their use cases.