Intel I219-LM Running at ~60% of Maximum Speed Due to Linux Driver Bug
phoronix.com
phoronix.com
All my problems gone, life is good.
[1] Ironically, in this case the Intel has to have a fancy feature, intended for increasing performance, disabled in order to increase performance.
For 1G realteks, I've had ok luck with them, but I do see some issues. I recall having some problem with them while running the Linux tree drivers, but I've since switched to all FreeBSD. With my current realtek NICs and FreeBSD 13.1, I have the choice of the kernel driver where sometimes the NIC will stop processing packets, and the NIC acknowledges reset, but doesn't actually immediately reset and sometimes processes old packets after resetting, resulting in wild writes and bad behavior. There's a vendor driver, which doesn't seem to get into that bad state, but it does enable ethernet PAUSE frames which are an abomination. The vendor driver is full of undocumented magic values and there's no public documentation on the NICs at all, so figuring out what to frob to make the NIC not get stuck, but not send out PAUSE frames would be an exercise in frustration, that I'm not willing to do.
Either way, the interrupt design is deficient: there's a shared interrupt for rx, tx and administration, and the status register doesn't really work right either --- it's possible for the host and device to disagree about what irqs were acknowleged and then things will get stuck (that doesn't seem to be the FreeBSD driver issue though; you can work around the stuck status communications by just assuming something probably happened aftet a few seconds, or checking for descriptor progess on rx and tx on any interrupt, etc). Having only a single interrupt means you can't meaningfully process incomming packets and finished outgoing descriptors in parallel which makes it hard to get full throughput in both directions simultaneously.
Anyway --- if they work for you, great. I'm going to avoid them where practical, and be careful with them elsewhere.
There is plenty if you search around, and they're one of the recommended NICs if you want to write an OS/driver because of that; they also show up in virtual form in various VM hosting software due to their simplicity:
https://wiki.osdev.org/RTL8139
It's an OK NIC to work with hobby wise, although I quickly ran into issues with the status register, and I can only hope one day I'll get enough throughput to overwelm the nic. And then I'll move to an Intel nic of which I have several.
At the end, Debian reverted Intel's patches and shipped a modified version which probably didn't enable some features on the newer NICs, but didn't break the older ones in the process.
It's ironic that the underdog Realtek has much more reliable cards which doesn't shudder under constant load and/or very long uptime scenarios. Realtek won't cut in server scenarios, but Intel's and Broadcom's server class NICs are completely different beasts when compared to Intel's consumer NICs, too.
Wish they had given some more attention to them while developing their drivers.
The TP-Link and Netgear switches had the least issues for me in that I only had to force a particular speed to avoid sometimes negotiating at 100mb. I am not sure people will notice the negotiating bugs unless their family are heavy bandwidth users. I noticed it because of streaming and gaming at the same time plus I shut down my machines at night so I have a coin-flip chance of negotiating incorrectly every day.
Damn. That's pretty much a requirement for any upcoming home lab gear here, so I'd better be careful about this aspect when researching.
Thanks. :)
consensus is Intel went down the shitter quality wise, just one example https://www.youtube.com/watch?v=DXNyHFOWx_k
If anything, that message seems to be implying to not upgrade if everything is already working well.
A bit tangential, but ever since they came out with the '217 I've thought this is one of the worst-named products and could never remember what the actual letter is. Here it's an uppercase I, but even Intel seems to think it's a lowercase L (for LAN?) sometimes:
https://www.intel.com/content/www/us/en/support/articles/000...
I'd say a lot. Not seldom when I go over to a colleague I notice their laptop is breathing hard. "Oh yeah, so annoying! It's been doing that for a while now". Check Task Manager and sure enough, some process is stuck sucking 100% of a core... 5 CPU days worth.
Even I find it difficult to spot this, with my 16 cores. I got an efficient cooler, so no fan spins up noticeably. Randomly check Task Manager, oh explorer.exe is pegging a core, and has been using 15 days worth of CPU... Gee thanks!
But to be fair, this is an old NIC with a hardware bug that requires an odd workaround disabling an otherwise very useful performance feature. Missing that edge case is understandable.
Saying that because at some point, that NIC was a new thing, not an old one. So extensive regression testing should have picked it up then, yeah?
It is not a new thing, but no test is exhaustive and they must prioritize hardware tested for practical reasons. This particular test would need to check if the working NIC could reach full gigabit speeds in TX, as everything else checked out.
Regardless, it cannot be said that they do not have regression testing or are irresponsible in that sense, as they provide important services to the community in exactly this area.
To me, that sounds weird.
My expectation is that every non-dodgy manufacturer of a network device would test the features they're developing.
So, when (say) offloading support of some variety is newly added to a chip set, whoever is adding that support would be testing it at the full speed of the device.
If the developers aren't verifying the things they're adding work (somehow), there's something wrong. :(
But generally speaking you want all the offload you can get. For 1G, it allows pushing power efficiency much further. For 100G, it allows you to spend time on something other than parsing and writing packets.
Reading the comments here, we aren't the only ones taken by surprise and/or aback.
(1) HW GRO, most often LRO (TCP)
(2) SW GRO aka GRO
Hardware GRO was typically called LRO for TCP and is often disabled in practice. One reason is XDP, with XDP some offloads need to stay off. Another is device quirks: while idea of waiting for more data, and passing larger packets into kernel soiunds nice, it also means that device must do packet concatenation. Think about ECN or selective-acks. The concatentation is lossy and unless the flow is perfect, it might not be possible to concatenate packets. Furthermore devices often had issues with being too aggrressive and loosing important data.
In Linux there is a software approach that does kindof the same thing, but in software, called GRO, and this is often sufficient. It also supports more than just TCP.
In other words - having LRO disabled is common.
If you want to recover, set a ifdown script to `rmmod igc && modprobe igc` and you will never panic or hang, just have a longer interface bounce time. I'm running this on a 6 nic system which is completely egregious, but it hasn't crashed / hung / panic'd in the field since doing that.
for those 3 nics they use.
I doubt some random 2.5Gbit nic is on their list of stuff to fix
and then we look at the wifi stack in BSD(s)...