Streaming video on 10 Gigabit Ethernet and beyond
bbc.co.uk
bbc.co.uk
I know it is overkill, its just that it has been about ten years already, isn't it cheap enough yet? Can't a modern ssd keep up with it?
I can't think of anything except connecting to a _fast_ SAN that would require a 10GBE port in a laptop. Maybe something specialized for a network engineer, but even then it's probably easier to buy dedicated equipment for line rate port monitoring
I found this:
http://www.fastestssd.com/featured/ssd-rankings-the-fastest-...
Pushing 3000MB/s, which is 3GB/s which should be times 8 for gigabits, I think it should be viable now, no?
An off-the-shelf consumer raid-5 nas would require more than a gigabit port (so a 10 gig port would be needed).
At best I can get 60m/s out of it. Each drive can do about 100m/s sequential but that is rare.
Putting an ssd on for caching read/writes though really changes the calculus of this.
The M.2 interface has finally shrugged off the SATA bottleneck for commodity hardware. It's common on new motherboards, new laptops, and I recently read there's similar circuitry in the new iphone 6s.
There are Thunderbolt adapters. USB3.0 only has 4Gbit/s available, ExpressCard only 2GBit/s, so both aren't really good options.
There's a lot of other bandwidth issues that aren't sorted out yet, too. If the adapter and drivers don't have TCP offloading, you'll be very hard pressed to get more than 3Gbit for anything other than a UDP dump.
Most of the OS network stacks aren't properly tuned for 10Gbit either.
I'd say give it a few years; as it becomes more mainstream, things should improve.
Any modern adapters you can think of that don't support TCP & UDP offloading (+ARP, etc.)? As far as I know, all of them support it.
I didn't mean to suggest that it's hard to find, just that there are a host of things that need to be in place before you'll get the expected speeds. This isn't a knock on the tech or any manufacturer, it's just that it's not mature enough to to be like 1Gbit where you plug it in an almost everything starts running at 100Mbyte/s.
Here's an example of what to expect (I have no affiliation with this thread): https://forums.creativecow.net/thread/197/860183
Of course it's a completely another story without SIMD. A naive traditional checksum loop with a register dependency stall is just not going to be fast.
The L3 layer checksum is useless because IP packet is small and the kernel has to read/write all the fields anyway.
The L4 checksum covers TCP/UDP packet data, which the kernel can avoid touching if necessary.
When a TCP sender uses sendfile(), the kernel does a DMA read from storage to a page if the data is not already in memory (in the so called page cache), and just ask the network card to send this page, prepended with a ETH/IP/TCP header. That only works if the NIC can checksum the TCP packet content and update the header.
If the network card can do TCP segmentation offload, the kernel does not have to repeat this operation for each 1500 bytes packets, it can fetch a large amount of data from disk, and the NIC will split the data in smaller packets by itself.
the other kind of offloading that the kernel can use is TCP Segmentation Offload (TSO), which is much more complex to implement in hardware, and you won't find it on cheap NIC (like Realtek)
> Each core needs to generate a few thousand data packets per second, because Ethernet packets typically contain up to 1500 bytes. This gives the CPU around 100 microseconds to process each packet.
No it doesn't, not when using TCP Segmentation Offload (TSO)
This only works for a particular use-case: sending static data using TCP, but this is the most common use-case since a typical "video streaming server" is actually a simple HTTP server that serves static MP4/MPEG-TS data.
for each connected client this is what happens - nginx/apache does sendfile(file, sock, off, <large_number>) - kernel issue large (> 10kB) DMA read to the file storage backend into a set of memory pages and wait for completion - kernel allocates/clone a small IP/TCP header (40 bytes) - kernel gives that small header + set of memory pages to network card, which will segment and create those 1500 bytes packets and send them on wire
if you have a lot of RAM, the read from storage could even be skipped because the previously read data pages are kept in the page-cache with a LRU approach. (help if clients are requesting the same file).
you can easily saturate a 10G link with spare CPU cycles on cheap hardware with that approach, no need to bypass anything.
So right now, i am missing the information on what kind of NIC they were using. Any thoughts or comments on that HN-community?
What vendor and product model would be a reasonable entry point for such endeavours? Answers very much appreciated.
Cloudflare wrote a blog post recently about accelerated packet IO and their post mentions the 82599 NIC: https://blog.cloudflare.com/kernel-bypass/.
His project, Snabb Switch, also utilize the Intel 82599 10G NIC that 'xtacy mentioned.
Use the right tool for the job and don't funnel network data through your instruction pipeline. When they realized that for memory, they called it "DMA", and when graphics was scaling up, we created the GPU.
other than that they are just bypassing tcp, arp, ethernet, etc. basically one pc has a kennel that says "every one in this file descriptor turns the voltage up on this cable" and the other pc has a driver that "every voltage up on this cable writes a one on this file" then they add some rudimentary sync logic for the timings. maybe just a known initial handshake that both expect and know... like a modem have shake.
i wonder if there is already a well known project/Linux kernel driver for this dumbed down network-as-fast-interface around or if they are writing it. the article is really lame on any detail
What they're doing is bypassing the kernel overhead of header parsing, demultiplexing and copying to user space. Instead, the network card's ring buffers are mapped directly into the user processes' address space.
The application is almost certainly still talking TCP/IP (or maybe UDP), and there are no changes at the physical layer at all. It isn't a case of a file descriptor being hooked up to generate voltages on a cable at all -- in fact, the overhead of read()/write() calls on a file descriptor is one of the things they're trying specifically to avoid!
http://dpdk.org/ is an open-source library to implement this kind of thing, mostly aimed at Intel NICs.
(When you have a hammer (web technologies) everything looks like a nail I guess)