XDP for game programmers
mas-bandwidth.com
mas-bandwidth.com
I have a VFX(Houdini now, RSL shaders etc earlier) and openCL-dabbling and demoscene-lurking background, based on which I think I prefer 'shader' to 'kernel', that's what OpenCL calls them.. but that conflicts with the name of like, 'the OS kernel' at least somewhat..
I wish someone would take up this effort again. It'd be awesome to see VPP or someone target offload to GPUs again. It feels like there's a ton of optimization we could do today based around PCI-P2P, where the network card could DMA direct to the GPU and back out without having to transit main-memory/the CPU at all; lower latency & very efficient. It's a long leap & long hope, but I very much dream that CXL eventually brings us closer to that "disaggregated rack" model where a less host-based fabric starts disrupting architecture, create more deeply connected systems.
That said, just dropping down an fpga right on the nic is probably/definitely a smarter move. Seems like a bunch of hyperscaler do this. Unclear how much traction Marvell/Nvidia get from BlueField being on their boxes but it's there. Actually using the fpga is hard of course. Xilinx/AMD have a track record of kicking out some open source projects that seem interesting but don't seem to have any follow through. Nanotube being an XDP offload engine seemed brilliant, like a sure win. https://github.com/Xilinx/nanotube and https://github.com/Xilinx/open-nic .
It looks like they have some demo code doing something like that. https://docs.nvidia.com/doca/archive/doca-v2.2.1/gpu-packet-...
What kind of workloads do you think would benefit from GPU processing?
Actual traffic for a metaverse is mostly bulk content download. Highly interactive traffic over UDP is maybe 1MB/second, including voice. You're mostly sending positions and orientations for moving objects. Latency matters for that, but an extra few hundred microseconds won't hurt. The rest is large file transfers. Those may be from totally different servers than the ones that talk interactive UDP. There's probably a CDN involved, and you're talking to caches. Latency doesn't matter that much, but big-block bandwidth does.
Practical problems include data caps. If you go driving around a big metaverse, you can easily pull 200GB/hour from the asset servers. Don't try this on "AT&T Unlimited Extra® EL". Check your data plan.
The last thing you want is game-specific code in the kernel. That creates a whole new attack surface.
Typical bandwidth for multiplayer games like FPS (Counterstrike, Apex Legends) are around 512kbps-1mbit per-second down per-client, and this is old information, newer games almost certainly use more.
It's easy to see a more high fidelity gaming experience taking 10mbit - 100mbit traffic from server to client, just increase the size and fidelity of the world. next, increase player counts and you can easily fill 10gbit/sec for a future FPS/MMO hybrid.
God save us from the egress BW costs though :)
It will be interesting to see if Epic makes a streamed version of the Matrix Awakens demo. You can download that and build it. It's about 1TB after decompression. If they can make that work with their asset streaming system, that will settle the question of what you really need from the network.
The only thing in games that is bandwidth heavy is on-demand game asset delivery, which is highly cacheable and shardable. It will need no XDP-like networking tricks in either servers or clients.
This is just absolutely not true.
In GTA V, each region has a custom-built low-rez model, and that's what you're seeing when you're more than about 200-300m away from it. Watch closely and see where the cars appear and disappear in the distance. That's the edge of the real rendering area.
I'm looking at doing this for a metaverse. In the GTA V era, those impostors were a manual job done by game devs. That needs to be automated. Rather than doing mesh reduction on large areas, I want to take pictures of each area at high resolution from multiple angles, and feed the pictures through Open Drone Map to get a 3D mesh. The result looks like this.[1] For even more distant areas, those meshes can be consolidated into larger and lower-rez mesh tiles. It's the 3D equivalent of a slippy map. The amount of data you need to send is finite regardless of the world size, because the far-away stuff has lower resolution. The sum of that series is finite. This is similar to how Google Earth works when you get close enough to see 3D.
Handling a metaverse with user-created content is a big data-wrangling problem, but the compute and network loads are finite.
[1] https://content.invisioncic.com/Mseclife/monthly_2023_12/bas...
XDP is useful for applications that are network I/O bound. Gaming is not one of those.
The big-world high-detail user-created metaverse problem is being worked on.
> Why? Because otherwise, the overhead of processing each packet in the kernel and passing it down to user space and back up to the kernel and out to the NIC limits the throughput you can achieve. We're talking 10gbps and above here.
_Throughpout_ is not problematic at all for the Linux network stack, even at 100gbps. What is problematic is >10gbps line rate. In other words, unless you're receiving 10gbps unshaped UDP datagrams with no payloads at line rate, the problem is non existant. Considering internet is 99% fat TCP packets, this sentence is completely absurd.
> With other kernel bypass technologies like DPDK you needed to install a second NIC to run your program or basically implement (or license) an entire TCP/IP network stack to make sure that everything works correctly under the hood
That is just wrong on so many levels.
First, DPDK allows reinjecting packets in the Linux network stack. That is called queue splitting,is done by the NIC, and can be trivially achieved using e.g. the bifurcated driver.
Second, there are plenty of available performant network stacks out there, especially considering high end NICs implement 80% of the performance sensitive parts of the stack on chip.
Last, kernel bypassing is made on _trusted private networks_, you would have to be crazy or damn well know what you're doing to bypass on publicly addressable networks, otherwise you will have a bad reality check. There are decades of security checks and counter measures baked in the Linux network stack that a game would be irresponsible to ask his players to skip.
I'm not even mentioning the ridiculous latency gains to be achieved here. Wire tapping the packet "NIC in" to userspace buffer should be in the ballpark of 3us. If you think you can do better and this latency is too much for your application, you're either day dreaming or you're not working in the video game industry.
But games are not 99% fat TCP packets.
Games are typically networked with UDP small datagrams sent at high rates for most recent state or inputs, with custom protocols built on top of UDP to avoid TCP head of line blocking. Packet send rates per-client can often exceed 60HZ, especially when games tie client packet send rate to the display frequency, eg. Valve and Apex Legends network models.
Now imagine you have many thousands of players and you can see that the problem does indeed exist. If not for current games, for future games and metaverse applications when we start to scale up the player counts from the typical 16, 32 or 64 players per-server instance, and try to merge something like FPS techniques with the scale of MMOs, which is actively something I'm actually doing.
XDP/eBPF is a useful set of technologies for people who develop multiplayer games with custom UDP based protocols. You'll see a lot more usage of this in the future moving forward, as player counts increase for multiplayer games, and other metaverse-like experiences.
Best wishes
- Glenn
My apologies, the last time I looked at DPDK was in 2016, and I don't believe this was true back then. Either way, it seems that XDP/eBPF is much easier to use. YMMV!
As a side note though, I would encourage you to revisit DPDK. I've worked in the low latency space for a long time, and used pretty much every solution out there, open source or proprietary, and DPDK is one of my favorite.
This is because DPDK is not _just_ poll mode drivers, it's a a full featured SDK for low latency packet processing. You get top notch thread safe memory pools to store packets, thread safe queues, hand written intrinsics optimized hashing and CRC algos, a cooperative scheduler in case you isolate your CPUs, etc, etc.
EDIT: Might be some benefits in terms of latency through the stack and resource usage on the machine... But I don't think 10Gbps is where the pain is for these either.
Maybe we'd like to use some of that CPU to run the game instead of pushing packets.
Maybe if we can push the packets more efficiently, we save $$$.
Maybe game servers get DDoS'd and it's great to be able to quickly drop packets without any linux kernel overhead.
Maybe I would like my games to not bypass my firewall and VPN?
Also, eBPF is not a way to shave CPU cycles, you would use DPDK/Netmap/BSD BPF for that, it's a way to lower latency. And we're talking single digit microsecond latency, where a typical internet link has millisecond jitter.
If it's for servers then the article still doesn't make any sense. See my other comment https://news.ycombinator.com/item?id=39939876
https://www.cablelabs.com/10g https://www.nngroup.com/articles/law-of-bandwidth/
Most of the world doesn't even have 100M internet: https://www.speedtest.net/global-index
By the time your game can use even 1G per client, XDP won't even be helpful for 100G. (Which, depending on use case, it already isn't.)
According to Neilsen's law, we should see 10G internet become more common than 1G internet is today within 10 years.
That's a cool thing. Let's look forward to it, and the rising tide that lifts all boats. The world will be a much more connected place 10 years from now, and the 10G target from CableLabs.com as well as DSCP packet tagging, L4S and other technologies will make the internet a lot better than it is today.
Why the negativity. Embrace the future :)
You do realize the most random Python program, written by the most random programmer, running on the most random computer, could easily read 10k TCP packets per second per core?
Uh, "the internet" traffic shape is a terrible model for low-latency multiplayer games traffic shape. Surely you don't think that 99% of the packets being exchanged during a session of Counter-Strike 2 are fat TCP packets, right?
Hell, I don't even think a commodity gaming computer have enough cores to process line rate datagrams on a 1gbps link.
Kernel bypass is not going to be something that is necessary on the client (even though there is a trend towards symmetric 10G internet becoming available in the US today, and according to Neilsen's law, it should be widely available within 10 years).
To the experts in the field in question these comments are usually incorrect and often highly amusing, but they are stated with such gravitas and certainty by the poster, that a casual reader might mistake them for the truth.
Also: https://www.mobygames.com/person/68015/glenn-fiedler/
Note that MobyGames is wildly incomplete.
Still, I think bandwidth and latency should not be confused as it is on this thread.
From my experience, the Linux network stack is not bandwidth bound _at all_. I've used it extensively on 94gpbs infiniband networks without any issues. Problems arise only when you close up to line rate, meaning very small payload and low packet latency, which I think (though I'm not a game developer by any mean) is unlikely to happen in game use cases.
As for the technological side of things, I have my reserves that open source solutions like XDP, Netmap, or BSD BPF are able to compete with proprietary solutions like DPDK (which IMHO is more complete) or SF OnLoad - which is just an LD_PRELOAD on top of your existing BSD socket server.
> according to Neilsen's law
If were being honest, this is more of a fun fact than a law... We're talking about a single guy fitting a linear regression of a dozen points of his internet connection speed over time.
If you bypass the kernel stack on an publicly addressable network, then it is your responsibility to implement and calibrate backloging, handshake recycling, SYN cookies
He's fully focused on the backend.
I have no doubt that more bandwidth will change how games are made, but 5v5 shooters (and pretty much all existing multiplayer styles) are here to stay for a lot longer than that, in some form or another.
LOL, no way in America is that going to be true.
That opening blanket statement turned me off so hard I couldn't get past that paragraph to read the rest of the article.
I am struggling to justify 10Gb LAN in my house. Between purchase costs and energy requirements (some 10Gb arrangements seem to be crazy inefficient). And I like "more, more, faster, faster" in my tech.
How does this scale (and cost) across even a significant section of human society?!? According to the article's prediction things must get crazy soon.
Given it applies to "A high-end user's connection speed" I am dubious as to the applicability to everyone else - say the bottom 20% of society for example - and so I am most doubtful of the article's "everyone will have 10Gbit" claim.
Maybe "everyone" has a stricter definition than I think.
But as someone who has literally grown up alongside PC's - your link gives confidence to look forward to the crazy ahead. :-)
So these mini kernel programs are written in a subset of C?
Also event-based protocols with deterministic physics.
Last but not least, you need to use a language that can atomically share memory between threads; C (with Arrays of 64 byte Structs) or Java.