A high-speed network driver written in C, Rust, Go, C#, Java
github.com
github.com
Java: https://github.com/ixy-languages/ixy.java/blob/master/ixy/sr...
C#: https://github.com/ixy-languages/ixy.cs/blob/master/src/ixy_...
Java needs a bit more C to make it work. C# only seems to need it for DMA access. But when you look at the Go code, they got away with being pure Go and using the syscall and unsafe package. So that's at least one plus for Go.
(the main readme calls this out, but at least one thing worth mentioning here too).
As a Java coder for my day-job, I do like the breakdown they have of the performance of the different GCs for their Java implementation. https://github.com/ixy-languages/ixy-languages/blob/master/J...
We do it all the time in Windows low level coding within .NET, the reason why they didn't beats me, most likely not knowledgeable enough of .NET capabilities.
As for Java, hopefully with projects Valhala, Panama and Metropolis Java will finally have the performance language features that should have been part of Java 1.0.
>"As C# cannot call mlock or get a raw pointer from a memory mapped file, DMA memory allocation is performed in C and called with the C# P/Invoke mechanism. Fortunately, this is the only instance of the driver calling a C function and the total amount of C code is only around 30 lines."
On Windows even for really obscure FNs you can almost always PInvoke if you know the offsets, and if you really want to be evil you can traverse the PEB. There isn't much in low level terms that is beyond the reach of C# since you can manipulate memory directly. I've also accessed hidden COM interfaces by traversing V-Tables using similar direct memory techniques as you would in C.
[DllImport(..., EntryPoint = "mlock")]
extern int MLock(UIntPtr addr, uint len);
And then getting the pointer either from MemoryMappedViewAccessor or AllocHGlobal?Why not? What am I missing? https://gist.github.com/Const-me/49f3da0ae744194fbf5be535527...
To measure how secure the resulting code is, have a test suite of malformed packets or other input and see how many of them the code in each language handles.
That measures how well a particular developer takes, with some noise induced by the working environment (say there's construction noise during python week, not during java week, or more meetings one week than another, or the developer's has relationship troubles at home). Randomness happens.
Deducing a number that's more generally valid requires having n workers doing the same work and then doing statistical analysis. Happily that also takes care of the single-worker problem. Still, the cost of the experiment easily rises by a factor of twenty or a hundred, depending on how well the noise can be controlled and how much accuracy is needed.
Asking for improvements that would increase the cost of an experiment by many thousand per cent is a %$#@!%#@! $%#@!%@#$ thing to do. IMO.
Objective-C was often called slow because iteration NSArray was much slower than doing it in C. Well, if you needed to do it fast in Objective-C you wouldn't do it using the user friendly and safe (for 1984) higher level objects.
I think only Rust really allows you to write really safe and still really fast code though.
I clicked through to a performance analysis showing ARC taking about 3/4 the time (in the release build).
You don't really need to be doing a lot of ARC in the inner loops if you don't want to.
Pull requests proving otherwise are welcome
One fundamental aspect of swift is the distinction between reference types -- which are reference counted -- and value types -- which are not. Generally in Swift you'd use a value type over a reference type unless you have reasons not to. E.g.: https://developer.apple.com/documentation/swift/choosing_bet...
I mean, I don't know what the right approach for this library is. The authors are going to have to fix their own code. IMO, coming up with a demonstrably poor solution and trying to defend it as "idiomatic" is pretty weak.
Snabb (https://github.com/snabbco/snabb/) is written in LuaJIT . I assume an equivalent project in Rust would be a lot more expensive in implementation time and also lines of code.
There is also Common Lisp and SBCL in particular, which can produce extremely fast code without compromising on safety.
Bidirectional forwarding, Packets per second: Here, the batch size matters; small batches have a lower packet rate across the board. Each language has increasing throughput with increasing batch size up to some point, and then the chart goes flat. Python is by far the slowest, not even diverging from the zero line. C is consistently the fastest, but flattens out at 16-packet batch at 27Mpps. Rust is consistently about 10% slower than C until C flattens out, then Rust catches up at the 32-packet batch size, and both are flat at 27Mpps. Go is every so slightly faster than C# until the 16-packet batch size where they cross (at 19Mpps), then C# is consistently about 2Mpps faster than Go. At the 256-packet batch size, C# reaches 27Mpps, and Go 25Mpps. Java is faster than C# and Go at very low batch sizes, but at 4 packets per batch Java slows down (10Mpps), and quickly reaches its peak of 11 to 12 Mpps. OCaml and Haskell follow a similar curve, with Haskell consistently about 15% slower than Java, and Ocaml somewhere between the two. Finally, Swift and Javascript are indistinguishable from each other, both about half the speed of Haskell across the board.
Latency, at 90, 99, 99.9, 99.99.. etc., percentile. 1Mpps: All have zero-ish latency at the 90 percentile point, then Javascript latency quickly jumps to 150us, then again at 99.99%ile jumps again to 300us. C# is the next to increase: at the 99%ile mark there's a steady increase till it hits 40us at 99.99%ile. Then a steady increase to about 60us. Haskell keeps it at about 10us until 99.99%ile, then a steady increase to about 60us, and a sudden spike at the end to 250us. Java latency remains low until 99.95%ile, then it quickly spikes up reaching a max of 325us. Next OCaml spikes at around 99.99%ile, reaching a max of about 170us. Next comes Swift, with a maximum of about 70us. Finally, C, Rust, and Go have the lowest latency. Rust and C are indistinguishable, and Go latency diverges to about 20% higher than the other two at the 99.999%ile mark, where it sways, eventually hitting around 25us while C and rust hit about 22us.
It would be a bit of a cheat, as it isn't portable, but it would be nice to see prefetching in the C implementation for the sake of comparison.
I've written NIC drivers for some older chipsets, and IMHO it's not something that's particularly "algorithmic" in computation or could necessarily show off/exercise a programming language well; what's really measured here is probably an approximation to how fast these languages can copy memory, because that's ultimately what a NIC driver mostly does (besides waiting.) To send, you put the data in a buffer and tell the NIC to send it. To receive, the NIC tells you when it has received something, and you copy the data out. Nonetheless, the astonishingly bad performance of the Python version is surprising.
Although I haven't looked at the source in any detail, I know that newer NICs do a lot more of the processing (e.g. checksums) that would've been done in the host software, so that would be another way in which the performance of the host software wouldn't be evident.
One other thing I'd like to see is a chart of the binary sizes too (with and without all the runtime dependencies).
In the paper, they point out that the Python version is the only one they didn't bother to optimize.
However, my takeaway is that practically everybody can handle north of 1 Gigabits per second (2 Million packets per second x 64 bytes per packet) even on a 1.6GHz core. I find THAT quite a bit more astonishing actually.
That's not shabby for a language like python.
Lack of necessity.
Since the telcos are a gigantic bottleneck to everything in the cloud, and now that everything is in the cloud, there is no need for >1Gbps home networking.
I run 10gbit inside my home and it didn't even cost me that much (if you go with 10Gbit fiber instead of copper) with the sole reasons of getting quicker transfers between my PC and NAS. My NAS has 4 SFP+ ports and functions as a switch. I bought second hand PCIe SFP+ NICs for $40 each and matching transceivers for $15 each. 10M of fiber costs less than $10.
There's no point in going higher, because 10Gbit is already way past the sequential writing speed of the drive array in my NAS, and it's pretty much saturating the NVMe cache drive in the NAS or the NVMe storage in my PC.
That's not to say you can't go faster, because 100, 200 and 400Gbit are very much possible and in use in datacenters and the like.
That hasn't been true for a long time. Even one single spinning rust hard drive made in the last decade can do sequential reads at ~120-150MiB/sec, which is easily enough to saturate a 1 Gbit/s link.
SSDs have way, way higher throughput for sequential read and write. Good SSDs will also beat that number handily for random read/writes.
And of course, any machine with more than 1 hard drive can easily saturate a 1Gbit/s network.
I also find it surprising that wired networking has been 'stuck' on 1Gbit/s for decades.
If your driver copies memory you are doing something wrong.
I looked at the C# thesis. I think with care the programmer was able to reduce the amount of heap allocations and memory copying enough that it's similar to the mix in the go driver. I think also the modern processors ability to execute multiple instructions/code paths in parallel tend to negate the advantage efficiently compiled languages like C and go. So cache misses, heap allocation, and garbage collection tend to dominate over raw numbers of instructions executed.
The authors specifically call out the issue of avoiding heap allocations when asked about Java v C# (as they're pretty similar languages), noting that they couldn't get under ~20 bytes allocated per forwarded packet in Java. C# (and Go) would have much better facilities to work entirely out of the stack, avoid memory copies and reuse allocations in the main loop.
I expect Haskell and OCaml have similar issues.
Only for throughput. The latency difference is enormous.
There's some more evlauation for Swift here: https://github.com/ixy-languages/ixy.swift/tree/master/perfo...
It's just a coincidence that JavaScript and Swift end up with almost the same performance; there is nothing similar between these two runtimes and implementations.
One thing that I notice about the Swift version is that he's making heavy use of classes, and in the performance section it mentions that there is quite a lot of time spent on retain-release. There's probably a lot of room for performance optimization in this implementation.
A search for 'raspberry pi mmap' will yield a lot of good starting points.
I haven't used C# much over the last year due to job change but always felt like one of the most mature languages out there. Now working in Go and it's a bit frustrating in comparison.
I bet a G2EE variant isn't too far away.
Yep, they aren't pure C#, still way more relevant to the world IT infrastructure than anything Go.
The core sql server rdbms engine is hosting the .net runtime for a few things - but there its maginal compared to C++.
Modern stacks are seldom pure blood language X, thus if 5% of it is written in Y, the product is written in a mix of X and Y.
Which I also mentioned on my comment, "Yep, they aren't pure C#", naturally overseen when one intends to champion its language as "Year of Desktop Linux" on IT infrastructures.
The Linux kernel is 0.2% of shell scripts, yet no one would say it's written in shell. Same for windows or sql server. .Net is marginal there.
Btw I do code in Go, but I mostly use go apps and enjoys their small memory footprint and ease of deployment.
I must not be the only one as I see go apps pretty much everywhere in ops teams.
Fortune 500 prefer to care about actual delivered business value.
Go talk to some SRE team in F100 and F500 and ask them what they think about C# infra side lol.
The fact that C# was running on Windows only until 2 years ago explains why.
Plenty them do run production servers on Windows.
You forgot there are 498 left to check.
What infrastructure, those riding the consulting and conference Docker and K8s 2019 wave fad?!?
All of this in a conservative big bank. My friends in the banking sector tell the same story.
True there are lots of c# enterprisey web apps.
However given the amount of boilerplate you describe, I cannot understand how such useful and reliable tools can be delivered in Go.
A hint:just because a language is not to your liking, it does not mean that it is not useful, performant and reliable.
AWS, Azure, actual hardware racks, plain old VMs, JEE containers, .NET packages, Ansible, Puppet, Chef, whatever scripting stack, but surely not one line of Go related code.
I don't follow your logic. Multitudes of useful and reliable tools were built with assembly languages - is that evidence that assembly code doesn't have a lot of boilerplate (relative to modern languages)?
Golang in 15 years will likely converge and embrace many of the missing features of mature languages (its happening now already), especially if it wants to reach broader adoption.
A reference like "Go is a get shit done language." very much reflects overall immaturity of the language that I see day to day.
Go is 10 years old already and picking new features at an extremely low rate, with no hints at a pace change.
I think error management and generics should be the only major changes to expect within the next 5 years. C# is more complex by an order of magnitude... And thus its evolution was and is still way faster.
Honestly I think that's pretty wrong. Go is designed to let you get coding quickly. It's not designed to make your ultimate solution well designed are easily refactorable. In a few years there's going to be a ton of Go code that becomes almost as bad as C where it becomes untouchable because people quickly threw something together and didn't think about long term design.
https://news.ycombinator.com/item?id=20944403
Rust was found to be slightly slower than C because of bounds checking, which the compiler keeps even in production builds.
Does `unsafe` actually impede optimization in this way? I thought it just disabled certain type checks and error messages but didn't affect anything on the LLVM level.
let queue = &mut self.rx_queues[queue_id as usize];
rx_index = queue.rx_index;
last_rx_index = queue.rx_index;
for i in 0..num_packets {
let desc = unsafe { queue.descriptors.add(rx_index) as *mut ixgbe_adv_rx_desc };
let status =
unsafe { ptr::read_volatile(&mut (*desc).wb.upper.status_error as *mut u32) };[1]: https://github.com/ixy-languages/ixy.rs/blob/master/src/ixgb...
queue.bufs_in_use[rx_index]
If so the bounds check could possibly be safely eliminated by the programmer because I think `wrap_ring` ensures that rx_index will always be in bounds?It would be really nice if this wasn't needed, but it's a valid use of unsafe code.
That makes sense. Their table of CPU performance counters [1] shows way more instructions per cycle for Rust and way fewer L1 cache hits for C.
https://github.com/ixy-languages/ixy-languages/blob/master/R...
You have some the results not quite following the conventional expectation. For example the Swift implementation is as slow as JavaScript. JavaScript is a lot faster than Python. Java is considerable slower than the usually very similar C#.
The implementation is fairly complex; so it is a bit hard to see what is going on. But it must be possible to pin the big performance differences implied by the two graphs to something?
> Unfortunately, the performance analysis shows, that the retain and release introduce a big performance hit, with little information on why exactly so many calls are necessary.
https://github.com/ixy-languages/ixy.swift/blob/master/perfo...
Now, a just-in-time (JIT) compiler transforms the code into machine code at runtime. Usually from bytecode. Java, C# JavaScript all use this model predominantly these days. This takes a bit of work during runtime and you cannot afford too complicated optimizations that a C or C++ compiler would do, but it comes close (and for certain reasons is even better sometimes). So that's the main reason why JavaScript is faster than Python. Theres a Python JIT compiler, PyPy, that might close the gap, though. And for Python in particular there are also other options to improve speed somewhat, one of them involves converting the Python code to C. Not too idiomatic, usually, though.
As for Java and C#, that's a point where it can sometimes show that C# has been designed to be a high-level language that can drop down to low levels if needed. C# has pointers and the ability to control memory layout of your data, if you need it. This turns off a lot of niceties and safeties that the language usually offers (you also need the unsafe keyword, which has that name for a reason), but can improve speed. Newer versions of C# increasingly added other features that allow you to safely write code that performs predictably fast. But even value types and reified generics go a long way of making things faster by default than being required to always use classes and the heap.
Java on the other hand has few of those features where the develop is offered low-level control. It has one major advantage, though, in that its own JIT compiler is a lot more advanced and can do some crazy transformations and optimizations. One might argue that Java needs that much magic because you don't have much control at the language level to make things fast, so as far as performance goes between C# and Java this may be pretty much the tradeoff between complicated language and complicated JIT compiler.
As for which benchmark shows Java being faster than C# depends a bit on how the code was written, but recently .NET has become a lot better as well and popular multi-language benchmarks show C# often faster than Java.
I wonder how PyPy would do on this benchmark...
https://github.com/ixy-languages/ixy.ml/blob/master/app/dune...
1 - https://en.wikipedia.org/wiki/Singularity_(operating_system)
Given that C# and "Rust" are neck and neck, I'd rather have a nice GC language to work with.
You can literally develop entire applications in unsafe mode, with C's level of unsafety. Nobody does, but you could.
.NET always had very good performance and the new .NET Core cross-platform framework and runtime is now consistently among the fastest in various performance benchmarks for all kinds of applications.
I'm not sure that's the conclusion to draw from a niche benchmark.
1. https://devblogs.microsoft.com/dotnet/performance-improvemen...
2. https://devblogs.microsoft.com/dotnet/hardware-intrinsics-in...
I think it's mostly a lack of insightfulness and thought that leads to cherry-picking bad guys when the system is structured in a way where unscrupulous behavior is worth risking. I could be wrong, maybe Microsoft is evil incarnate but at the level I operate on, I don't see it.
But yes, no reason to not embrace the good parts from Microsoft, or anyone else. That's how I do it. I chose C# to stick with because I believe in what they're doing around it. Huge fan of Blazor, appreciate their focus on long-term support, their product integration, and their excellent tooling. The language is good, and I can get a job doing it anywhere in the country without being in a major metropolitan area. In some of these metrics, I personally don't think you can beat the C# ecosystem.
C# recently getting new low level memory types definitely gave it the edge there, it does not reflect real world scenarios very accurately.
In my experience, C# is by a large margin the most performant "managed" language vs. Java, Python.
I ported a C hash function to C# recently, and using the low-level feature that C# offers, performance was very close to the C version.
C# is really nice to work with for this kind of thing.
And - .NET Core 3.0 introduced support for hardware intrinsics, so I could probably bridge the gap if I spent some time on vectorisingthe C# code.
"full bidirectional load at 20 Gbit/s with 64 byte packets (29.76 Mpps)". sounds like 20Gb/s should be closer to 40Mpps than to 30Mpps. Did you hit CPU limits on the packet generator, or am I missing some packet header overhead ?
Did you try bigger that 64-byte packets ? I'm curious how various runtimes would handle that.
And how long did you run the benchmarks ? I couldn't really figure it out from the github or the paper. Mostly wondering if java and other Gc'd language showed improvement or degradation over time. I could see the JITs kicking in, but I could also see the GCs causing latency spikes.
Yes: Ethernet adds 20 bytes: 8 byte preamble/start of frame delimiter + 12 byte interframe gap
=> the "on-the-wire" size is actually 84-bytes
=> 20Gbps/84-bytes = 29.76Mpps
> Did you try bigger that 64-byte packets ? I'm curious how various runtimes would handle that.
In typical forwarding, packet size does not impact forwaring that much until you hit some bandwidth limit (PCIe, DDR and/or L3 cache) because you only touch the packet header (typically the 1st 64-bytes cacheline in the packet). The data transfer itself will be done by NIC DMA.
Both C# and Java are faster than golang.
If you do have links to share, however, please do.
A lot of the projects I work on, for instance, heavily utilize dependency injection for no gain. There's only one implementation, theres no test mocks. Its just overengineered and obsfuscated for no reason.
Coming from a predominantly C++ background, we eschew virtual wherever possible, favoring compile time polymorphism to runtime whenever possible, because we're cognizant of the overhead of the indirect dispact and likely loss of optimized opportunities to inline trivial calls.
For sure, one can write C# or Java that can keep up, or even outperform C++ in some circumstances, but youre not going to do it with "enterprise" patterns hiding behind interfaces and factories and dependency injection.
There are the CppCon talks, the Modern C++ advocacy, and then there is the code that everyone at most corporations actually write.
Then there are ORM like the ill fated POET.
Yeah just like not everyone is writing code that "eschew virtual wherever possible, favoring compile time polymorphism to runtime whenever possible", specially on large corporations with mixed language teams.
Beyond C++ conference talks, I am yet to see stuff like SFINAE and tag dispatching in the C++ code I occasionally deal with. Grated those are libraries that get called from Java/.NET projects.
You don't usually see these sorts of types wrapped for Java or .Net, and if they are, you usually have some sort of proxy in between to hide the templates.
In this case Golang outperforms (in terms of throughput) Java on batch sizes > 4 and does so by nearly 2x at batch size of 256.
Now for some "popular" benchmarks:
https://www.techempower.com/benchmarks/
https://benchmarksgame-team.pages.debian.net/benchmarksgame/...
You see that Go is sometimes faster sometimes slower but the memory usage and latency is way bellow both of those languages.
You end up with a highly customized implementation not suited for wide use to get the higher performance benefits (and it still doesn't beat java on benchmarks like single/multiple queries and JSON serialization).
All these benchmarks should be taken with a large grain of salt. The golang compiler doesn't even pass function variables in registers (they're all stack allocated as far as I know), let alone do any of the advanced inlining and optimizations the JVM does.
[1] https://docs.google.com/presentation/d/e/2PACX-1vTxoBN41dYFB...
> Drivers are written in C or restricted subsets of C++ on all production-grade server, desktop, and mobile operating systems. They account for 66% of the code in Linux, but 39 out of 40 security bugs related to memory safety found in Linux in 2017 are located in drivers. These bugs could have been prevented by using high-level languages for drivers.
https://developer.apple.com/videos/play/wwdc2019/702/
https://source.android.com/devices/architecture/hal-types
https://docs.microsoft.com/en-us/windows-hardware/drivers/de...
Someone add D lang to this test! I want to know!
The author asserts, that "it's virtually impossible to write allocation-free idiomatic Java code, so we still allocate... 20 bytes on average per forwarded packet". This sounds questionable, — does that mean, that he actually performs a JVM memory allocation for _every_ packet?! Furthermore, the specifics of memory management look murky. One implementation uses "volatile" C writes [1] (simply storing data to memory). Another implementations of the same thing uses a full CPU memory barrier [2]. Which one is right?
In my opinion, significant inconsistencies between implementations render any comparison between them invalid. And when a whole cross-language test suite is written by one person, you can be sure, that they don't really excel in many of those languages.
This is why I like Benchmark Game — all benchmarks are submitted by users, so they are a lot closer to how a real-world decent programmers can solve the problem. Still not perfect, but at least that counts as an attempt.
1: https://github.com/ixy-languages/ixy.java/blob/fcad50339e537...
2: https://github.com/ixy-languages/ixy.java/blob/fcad50339e537...
The code is public, I'm sure they'd be happy to have your insight and fix this issue, it doesn't seem like they were happy about it.
A full memory barrier is not required, but some languages only offer that. For example, go had the same problem. It's not a bottleneck because it goes to MMIO PCIe space which is super slow anyways (awaits a whole PCIe roundtrip).
And no, it obviously wasn't written by only one person but a team of 10.
No, we are not saying that we allocate for every packet. We say that we allocate 20 bytes on average per packet.
https://docs.oracle.com/javase/9/docs/api/java/lang/invoke/M...