High Speed Networking: Open Sourcing our Kernel Bypass Work
bbc.co.uk
bbc.co.uk
The problem is that the linux kernel can't process many simultaneous small connections.
Just to be clear: - linux can easily transfer at 40 Gbps - linux chokes at around 1M packet per second per cpu socket.
Thats right.
So linux can easily transfer at 40 Gbps one or few simultaneous flows
But! It can't transfer several flows at 40 Gbps.
The bottlebeck is the number of packets per seconds it can inspect.
So if you are sending 100 big transfers, you will reach 40 Gbps.
If you send 5M small transfers, linux will die.
This is why netmap is handy.
It offloads the packets from linux directly to the app
Bbc is sending xM small requests per second, hence millions packets per seconds.
In conclusion, netmap is good if you need a lot of small simultaneous connections.
I reached 40M connections on chelsio 40 Gbps nic using netmap on FreeBSD 11.
The key point here is that a netmap driver bypass the kernel and therefore open the door to many millions packets per seconds or millions requests per seconds. Not many Gbps
which is why kernel bypass mechanisms f.e. dpdk/netmap/... etc. are getting some/lots of traction. fwiw, using dpdk (on x86-64 hw), i have had no problems with pushing minimal 64b sized packets for large (>250k) number of flows at line rate...
That is, sales will improve.
It is a very different application from what the BBC is describing. Rather than 1 80Gb/stream, we have tens of thousands of several Mb/s streams. And rather than kernel bypass, we use heavily optimized kernel path, via sendfile and in-kernel TLS.
One thing to note on topic, we developed a driver framework called 'iflib' in FreeBSD that is somewhat similar in scope to NDIS on Windows. If you write an iflib driver, you automatically get netmap support on FreeBSD. So we have support for all of Intel, and newer Broadcom as well, without potentially unfamiliar corporate devs having to understand the intricacies of soft ring management for userspace networking.
This is a huge myth.
1) even if you open-source, the company retains control over the project, and push their priorities first.
2) codebase can be huge and intricate, very few external contributors would actually put the time to learn it, leading to poor contribution (ie. might fix a given problem, but breaks other subsystems)
3) even if there is external contributors, is the company willing to put the manpower to process the changes and handle community interaction ? Ie. do you really want to handle bikeshed'ing style patches, potentially costing time to actual paid devs ?
Most of the time, niche opensource contributions are merely are more of a marketing tool (or HR tool ala. "your profile looks great however we're too greedy to hire you, please contribute to our software for free").
#1 is trivially true for all products.
#2 is essentially saying "if your code is crap, you will get crappy contributions". Well, if your code is crap, maybe you want to fix that?
#3 is valid. You don't necessarily get great results without putting in a little effort.
#2 does not mean that code is crappy at all, unless you consider that not writing down every single assertion made by a developer who wrote the whole subsystem as "crappy". In which case, I can assure you that all software is crappy. This might sounds as a huge generalization, but unless infinite amount of time and money, everything is a compromise, especially when deliveries are constrained.
Devs who are working daily on the codebase tends to know these compromises way better than newcomers. That's my point.
Can you elaborate on what you mean here? Is this reference to "the cathedral and the bazar?" What is meant by "rather few inertia"? Thanks.
Imagine you are google and you have created your own custom build system. You now want to hire more people that are familiar with it. Unfortunately the potential applicant pool is exactly 0 because nobody outside of google can even start to learn how to use the google specific tools because they simply don't exist outside of google.
1. People knowing about it.
2. People needing it.
3. People trusting it.
Like a fire needs heat, oxygen and fuel, new products survive only when they get all three corners of the triangle. If your product is cheap enough, some people will try it when they think that they have a need for something like that; trust comes eventually. Open source builds trust that in the worst case, people using your product will have a glide path out instead of a sudden brick wall.
In the instant case, Mellanox gets a highly visible, reputable customer telling everyone in their industry that Mellanox NICs are high-performance, are trusted, and can be adapted to their needs. Anyone who reads this article and thinks about high-performance NICs will have a bit more trust in the Mellanox brand as a flexible system.
https://www.bbc.co.uk/rd/projects/nearly-live-production
The idea is explicitly that you would run the equivalent of a TV control room on a web app that knows how to switch between video feeds (potentially from fixed cameras without human operators) and then can output video to... whoever is your audience.
The "moving window" of near-live editorial decisions is pretty interesting as well, and seems geared towards giving non-professional editors a chance to fix mistakes.
"Professional, live, multi-camera coverage isn’t practical for all events or venues at a large festival. For our research at Edinburgh Fringe Festival 2015, we experimented with placing three unmanned, static, ultra high definition cameras and two unmanned, static, high definition GoPro cameras around the circumference of the BBC venue. A lightweight video capture rig of this kind, delivering images to a cloud system, could allow a director to crop and cut between these shots in software, over the web and produce good quality coverage ‘nearly live’ at reduced cost."
Maybe we're talking about different kinds of events, but there's actually a lot of creativity that goes into camera framing, movement, shot selection, timing of cuts, etc. to end up with something that's actually interesting to watch...
"The user can pause the action at any point in time and seek back through the session to fine tune edit decisions using a visual representation of the programme timeline. On resuming, the play-head seeks forward to the end of the timeline. This functionality is made possible by ensuring the time shift window (DVR window) of the camera feed is infinite so that all live footage is recorded and can be randomly accessed by seeking.
Once the operator has established a sufficient buffer of edit decisions, they can begin broadcasting them. [...]
A research goal is to investigate how big the window of time should be between the broadcast play head and edit play head to ensure the operator has enough time to perform edits without feeling rushed and whilst keeping the programme as close to a live broadcast as possible. We call this the ‘near-live window’"
1 - https://www.bbc.co.uk/rd/blog/2017-07-compositing-mixing-vid... 2 - https://www.bbc.co.uk/rd/projects/ai-production
I think their proposal was that you could give up on being able to do camera framing and movement, but you would still maintain the ability to cut between shots at aesthetically pleasing moments. I could just about imagine that producing good results, with a relatively predictable performance and a decent technical director.
But I think a system like this doesn't compete with professional video production. It competes with someone making a shaky cellphone video in a small venue -- a form of video which a lot of people seem to accept, these days. I could see the BBC's proposal being a step up from that (and a lot cheaper than a mobile TV studio setup).
BBC R&D also has a project that's using AI and machine learning to investigate automating that work: https://www.bbc.co.uk/rd/projects/ai-production
With IP, it's usually a software defined network model, along with some IGMP components for controlling video flow.
So the clients (software or control panels) would send a request to the routing orchestration system to request a route to send multicast flows (SMPTE 2110, separate multicast for video, audio streams, ancillary data) from a source to a destination, same as with a baseband router. With IP, the orchestration layer also drives the receiving device to join the multicast explicitly if it isn't IGMP based.
This is our SDN orchestration; https://evertz.com/solutions/magnum
(we're hiring!)
In 40 years time we'll be still having to comply with this nonsense because some manufacturers didn't want to update their FPGA designs. It's like fractional frame rates all over again.
This is really where the broadcast industry lost its collective mind.
There are valid use cases for both hardware solutions where appropriate, and software solutions for other use cases. There are many low latency use cases where crazy requirements are still actual requirements.
This reminded me of something. At this point it seems like Jumbo frames are never going to be widely adopted, are they? Otherwise this seems like the perfect application - massive datarates, controlled hardware/software, high-quality wiring...
We got a couple of Connect-X5 cards, which allow switchless connections, akin to a ring topology. A lot of neat things, at just stupid line speeds, and latency levels I haven't seen in software, ever.
Everything else was consumer level though. I did ensure that everything was flipped to jumbo frames. I/O shouldn't have been the factor since it was SSD on both sides.
Mellanox has some interesting stuff, I used to work with it back in my HFT days.
There are different pass-through/fastpath patch for different chips to avoid memcpy, or do zero-copys, but they are all kernel patches, kernel-only.
An alternative method will be BPF/XDP and DPDK for which you will need modify the kernel drivers somehow for good performance. Wondering if Netmap does that already or is has nothing to do with them.
All of them are addressing Dataplane packet move, hope I can have an environment to experience these close-to-100bps network in my next projects.
In the meantime, I am wondering, why do you pass 4K uncompressed video using IP packets...
One nice thing about Netmap is it can fall-back to emulation mode and work with any NIC. This can be really helpful if you want a single codepath but don’t always need the performance. DPDK might support this too now, I haven’t looked at it for quite a while.
DPDK can fall back to emulation mode as well. See the AF_PACKET driver.
Also, I can't find a hello world like "send a packet" application, and it's complement of "receive a packet", only forwarding apps. That's not great when I'm trying to test if my config works.
I still don't know how many huge pages I need to reserve. The quick start guide says 64: http://dpdk.org/doc/quick-start
That didn't work. The Mellanox rep said 8192, which did work. That's a big difference. I can't find an explanation.
I invested about 1% as much time in Netmap to get to the same level of "my app nearly works".
Hello world apps are testpmd or dpdk-pktgen.
Hugepages depends on how much memory your program is going to allocate. They also come in different sizes, so that example you gave is only 64 2MB pages, which is pretty small. I usually go for the largest hugepages possible in the processor I'm on (2GB) and allocate a small number of those. The reason being that most of my DPDK applications are the only things running on that machine, so I don't need to worry about sharing memory.
I agree it's tougher to learn, but it's much more powerful. They likely need some kind of beginners' guide, since once you get past the tough parts, it's great.
FWIW, as far as I can tell, if you have no significant CPU work to do per packet, then you never need more than one core. A single core on a 5 year old laptop CPU can memcpy at a few hundred Gbit/s. In my tests I could even do the UDP checksum in software at 40 Gbit/s on one core. Mostly for me the bottleneck is the DDR RAM bandwidth, and using more cores doesn't increase that.
The docs should explain why/how the defaults values were chosen in the example apps.
I still don't know if there is any advantage to using more than 1 queue (aka ring) when using a single core. When I tried it, I found I could send nearly twice as many packets, but got lots of packet loss. Weird. I'm working on a public cloud, so the rules of what "line rate" really is are not available. There's some kind of fairness rules that either the hypervisor or the TOR enforces (I think). It's not at all clear if/how back-pressure is asserted to the NICs.
> Hello world apps are testpmd or dpdk-pktgen.
testpmd is 34,000 lines of code. dpdk-pktgen is 59,000. These are not hello world apps. I cannot read them to learn how to write a DPDK app. To understand if I had the basics right, I wrote a simple UDP sender and UDP echoer. They were 350 lines each. It seems like something similar should be the first apps everyone beginner sees.
> Hugepages depends on how much memory your program is going to allocate
Yeah, fair enough. I didn't explain my actual problem. My hello world programs worked fine with 64 huge pages on Intel cards. But I found no matter how small I made the ring buffers, 64 huge pages wasn't enough with Mellanox cards. And the error messages I got were gibberish from the bowels of the Mellanox driver (which isn't DPDK's fault). The Mellanox docs should have said how many hugepages their driver needs as an overhead (it appears to be more than 64).
I'd love to see a quick writeup (the code, the setup steps required to run the code).
Is this simply due to the overhead of IP transport supporting bidirectional communication? That is, a TV broadcast only needs to support a fixed set of N unidirectional flows (channels), but IP needs to support a dynamic set of N bidirectional flows?
I was amazed that a bit of software that is basically in all home routers is so old, unloved and misunderstood.
Multicast is a funny thing. For years people thought it was the future until they realised 1) it requires the WAN end to support it and 2) people to watch stuff at the same time.
Anyway that is the story I heard from one of their router vendors. Technically it is all correct.
From the viewer's standpoint, the bandwidth usage is roughly the same, but the broadcaster and CDNs have a lot more data to deliver.
There are a lot of scenarios with massive scale on the back-end, with lots of streams managed as part of the production process ahead of delivery to consumers... We have an IP video router (https://evertz.com/products/EXE-VSR) that has 2,304 10Gbps ports (moving to 25Gbps per port), each of which can do 6x fully uncompressed SMPTE 2110 video (1080i 29.97fps) flows in each direction, with a 46Tbps non-blocking back-plane so the scale can get a bit crazy.
The full scale back-end stuff is pretty invisible from the consumer side. Eventually that internal infrastructure feeds into distribution encoders that produce lower bitrate streams for cable/sat/web distribution (for real-time events, with separate file delivery for VOD platforms like iTunes & Netflix)
Can my computer monitor become just another network device plugged into Ethernet like my printer already is?
In my living room, can I ditch the video routing part of my AV receiver and just plug everything (streaming devices, video games, cable box) into a very fast ethernet switch?
Hard to see uncompressed consumer video catching on though since compression works so well.
Remember too that the feed will transit over many different IP-based systems. A seemingly acceptable delay on one system, a few microseconds here and there, adds up or multiplies as data streams between equipment.
Traditionally within a broadcast house, video+audio was sent between machines using HD-SDI using BNC coax[1]. For the first generation of HD, this ran at 1.5Gb/s. For 1080p/59.94Hz, a new standard was developed to run at 3Gb/s. For 4K at 59.94Hz, there is a 12Gb/s standard.
HD-SDI routers are extremely expensive. Each input and output from a device has to be individually cabled to a router and some devices can have many dozens of I/Os. There has been a push to switch from using ridiculous numbers of cables to using IP solutions and off-the-shelf IP routers. SMPTE 2110 provides a framework for doing this using RTP [2].
4K needs roughly 12Gb/s, so to carry a single video, you need at least 25Gb/s networking. If you want to carry many signals, the network bandwidth goes up fast -- that's why BBC is interested in 100Gb/s links.
Also keep in mind the traffic is not "bursty". These links may be nearly saturated 24/7. Most off-the-shelf routers are not designed with enough buffering to handle fully saturated links on all connectors.
[1] https://en.wikipedia.org/wiki/Serial_digital_interface
[2] https://en.wikipedia.org/wiki/SMPTE_2022 (older IP standard)
I have never heard of this (I am in the UK). Yes, they'll send you letters from time to time (yearly?) but "coming round every two months"?
the licence system is too heavy handed, especially for those that do not know their rights.
The license collectors remain a favourite grumble of anti-government or anti-tax people, most people aren't affected. (Most people buy the license, most people who don't need to instead tell the BBC roughly every 3 years.)
people can watch netflix and still be bothered by those people.
My view is that of all my susbscriptions for entertainment - sky, disney, netflix, and time wasted listening to ads, the one that gives easily the best value for money is 12 quid a month to the BBC.
(I mean sky is nearly 60 quid a month and i cannot persuade my other half to go broadcast-less)
There's something about broadcast stuff i'd love to work with again. Even just a broadcast mixing desk jogwheel was a league above anything i've ever DJd with, the engineering to make it feel correct was amazing. I'd imagine these days moving to IP it's more of a challenge I can step up to. Are the BBC hiring devops people...?
I have my doubts a micro kernel can come close to the Linux kernel efficiency.
Zircon is the new micro kernel that is part of Fuschia.
I agree that it'll be interesting to see how fuchsia turns out.
But we have to see it actually happen.