Snabb: 100 Gbit/s pure software switching using Lua (2019)
github.com
github.com
The NAS is TrueNAS with SSD and all the ZFS goodness, ZIL, etc. I can built hand crafted iperf streams between it and another 100g server that hit line rate (including overhead).
The reason for the 100g is there are racks and racks of BS that pull down their images so aggregate we can hit 100g with a fully non-blocking Clos fabric.
It may not pan out in practice, but in theory isn't that the perfect usecase for p2p image distribution?
[1]: https://sourceforge.net/projects/systemimager/files/flamethr...
Snabb is Swedish for fast. Switzerland is often confused with Sweden and vice-versa, stereotypically by the under-educated kind of americans.
Grr. :]
edit: Looking at the older threads linked above I found this:
https://news.ycombinator.com/item?id=8009309
Oh well I still choose to believe the Swiss are trolling the Swedes. :)
[0]: https://media.ccc.de/v/35c3-9670-safe_and_secure_drivers_in_...
It is hard to compare w/o knowing packet sizes (64-bytes? IMIX? 1500-bytes?), CPU type, #cores, workload (L2 patch? L2 switching with MAC-learning and 1M entries? IPv6 forwarding with 500K routes? etc.)
This is why we have the FD.io CSIT project hosted by LinuxFoundation where we can compare lots of interesting, automated, reproducible benchmarking scenario in a controlled environment on various HW platforms: https://docs.fd.io/csit/master/report/
Essentially, even if you write in C, reaching higher speeds will involve using "assembly masquerading as C" rather than depending on compiler optimizations.
Also, Snabb uses LuaJIT, which already generated quite tight code, so the performance gap that I suspect some imagine just isn't that wide.
==
You can write great C-based systems and avoid assembly if you a) know, always, what your compiler is doing and b) know, always, what your CPU is doing...
Is it possible to get better? yes. Would it count as "normal C"? I would say not really (if we say yes, then CL code on SBCL with huge custom VOPs counts)
Of course it does not apply to everything, you need a few hotspots, but it is quite common: audio/video codecs, scientific computation, games, crypto... And even networking.
I think if you are going to claim that C or C derivatives aren't actually fast and the idea that they are is due to "propaganda" then you should back that up with something concrete, because it goes against a lot of established expertise.
it seems that Snabb uses pflua at its core - which claims to be faster than BPF ( https://github.com/snabbco/snabb/blob/b9da7caa1928256a3b4908... )
Snabb side-steps the kernel by providing the performance aspects of kernel-like access while staying securely in user space.
So it's great because the eBPF verification makes sure that whatever the user loads doesn't blow up the kernel immediately, it's very restricted (no infinite loops, so not Turing complete, can only call predefined API functions, etc). But it still gets JITed into machine code, and still has access to a vast trove of kernel gadgets and data. So it's like a power tool that was designed to be almost impossible to cut your own limbs off with.
But it's still a power tool, it still not something you give to any random passer by.
Original BPF was safe by design (no verifier needed) but slow and far more restrictive, and even then there were exploits over the years. When eBPF and then eBPF JIT became a thing it was entirely predictable what would happen.