Nyxpsi – A Next-Gen Network Protocol for Extreme Packet Loss
github.com
github.com
What it does _not_ do is anything resembling intelligent determination of appropriate bandwidth, let alone real congestion control. And it does not obviously handle the part that fountain codes don’t give for free: a way to stream out data that the sender wasn’t aware of from the very beginning.
Network diodes have very low packet loss rates, for the simplest ones they are literally fibre to ethernet transducers with a disconnected channel and a few centimetres of fibre. There is no external disturbance that could generate significant bit flip. The more complex diodes use multiple transducers, power supply, fibres for redundancy and protocols to detect when one of the channels is out of order.
The diode manufacturers don't give details but we're probably looking at rather simple error detection and parity codes that can be easily implemented on an ASIC or FPGA to obtain high data rates.
On top of this you can use unidirectional protocols in UDP. But the diode works mainly at the physical and data layers of the OSI model.
Some diode manufacturers include proxies to simulate an FTP server and make it easier for users who just want to retrieve logs from critical infrastructure.
EDIT: It probably also needs to be clarified which TCP algorithm they are using. TCP standards just dictate framing/windowing, etc. but algorithms are free to use their own strategies for retransmissions and bursting, and the algorithm used makes a big difference in differing loss scenarios.
EDIT 2: I just noticed the number in parens is transfer success rate. Seeing 0% for 10 and 50% loss for TCP sounds about right. I'm not sure I still understand their UDP #'s as UDP isn't a stream protocol, so raw transferred data would be 100% minus loss %, unless they are using some protocol on top of it.
A TCP connection with 10% loss will work and transfer (it's gonna suck) and be very very very slow but their TCP 10% loss example is faster somehow?
TCP being reliable will just get slower with loss until it eventually can't transmit any successful packets within some TCP timeout (or some application-level timeout).
Even a 50% packet loss connection will work, within some definitions of the word "work" and also this brings up the biggest missing point in that chart: this all depends heavily on latency.
50% loss on a connection with 1ms latency is much more tolerable than 1% loss on a connection with 1000ms latency and will transfer faster (caveats around algorithms and other things apply, but this is directionally correct).
A real chart for this would be a graph where X is the % of packet loss and Y is the latency amount with distinct lines per protocol (really one line per defined protocol configuration, eg tcp cubic w/nagle vs without and with/without some device doing RED in the middle or different RED configurations, etc, many parameters to test here).
If this sounds negative, it's not, I think the research around effective high-latency protocols is very interesting and important. I was thinking recently (probably due to all the SpaceX news) about what the internet will look like for people on the moon or on mars. The current internet will just not work for them at all. We will require very creative solutions to make a useful open internet connections which isn't locked down to Apple/Facebook/Google/X/Netflix/etc.
It has led to some weird things whereby the L2 protocols (thinking wifi and LTE/cellular) have their own reliability layer to combat the problem you're describing. I'm not sure if things would be better or worse if they didn't do this and TCP was responsible, the iteration of solutions for it would be much slower and probably could never be as good as the current situation where the network presents a less-lossy layer to TCP.
We have to completely rethink things for interplanetary networking.
I thought I recognized your username, I remember jailbreaking the original iPhone on IRC with you helping :)
Yep, same author of Cydia [1].
https://egbert.net/images/tcp-evolution.png
https://en.m.wikipedia.org/wiki/Space_Communications_Protoco....
Like, if they're looking for publicity by sharing their github page, I'd expect the readme to have a basic elevator pitch, but their benchmarking section is a giant category error and it's missing even the most high level of summary as to what it is doing to achieve good throughput at high packet loss rates.
> https://github.com/nyxpsi/nyxpsi/blob/bbe84472aa2f92e1e82103...
This is not how you "simulate packet loss". You are not "dropping TCP packets". You are never giving your data to the TCP stack in the first place.
UDP is incomparable to TCP. Your protocol is incomparable to TCP. Your entire benchmark is misguided and quite frankly irrelevant.
As far as I can tell, absolutely no attempt is made whatsoever to retransmit lost packets. Any sporadic failure (for example, wifi dropout for 5 seconds) will result in catastrophic data loss. I do not see any connection logic, so your protocol cannot distinguish between connections other than hoping that ports are never reused.
Have you considered spending less time on branding and more time on technical matters? Or was this supposed to be a troll?
edit: There's no congestion control nor pacing. Every packet is sent as fast as possible. The "protocol" is entirely incapable of streaming operation, and hence message order is not even considered. The entire project is just a thin wrapper over the raptorq crate. Why even bother comparing this to TCP?
> For more information or to contact us open a PR or email us at nyxpsi@skill-issue.dev
This sounds to me like a troll. In any case, I don't think blatantly advertising your lack of expertise in networking is a good way of improving one's portfolio.
Who’s reading code these days?
[x] Sounds cool [x] Good number of stars
Hired
The reason why TCP will usually fall over at high loss rates is because many commonly-used congestion controllers assume that loss is likely due to congestion. If you were to replace the congestion control algorithm with a fixed send window, it'd do just fine under these conditions, with the caveat that you'd either end up underutilizing your links, or you'd run into congestive collapse.
I'm also not at all sure that the benchmark is even measuring what you'd want to. I cannot see any indications of attempting to retransmit any packets under TCP -- we're just sometimes writing the bytes to a socket (and simulating the delay), and declaring that the transfer is incomplete when not all the bytes showed up at the other side? You can see that there's something especially fishy in the benchmark results -- the TCP benchmark at 50% loss finishes running in 50% of the time... because you're skipping all the logic in those cases.
https://github.com/nyxpsi/nyxpsi/blob/main/benches/network_b...
Mathis equation suggests that even with 50% packet loss at 1ms RTT, the max achievable TCP [edit: Reno, IIRC] throughput is about 16.5 Mbps.
I can imagine noisy RF, industrial, congested links, new queueing at the extremes in densely loaded switches, but the thing is: usually out there are strategies to reduce the congestion. External noise, factory/industrial/adversarial, sure. This is going to exist.
Your wifi and home network probably are closer to 1% than 0%.
at 1-10%, you're probably on some kind of shared connection (CMTS or GPON with TDMA) and your provider has an overloaded network design (and you're at the end of the upgrade/split queue).
at anything above 10% you're in the realm of weird broken stuff. Congestion on the internet between major providers is a thing but is far less common than it was a decade ago, it does still happen though. Major providers who have NECMP or NLAG links inside their backbones or between providers where one of the links breaks in a way which doesn't remove it from the ECMP/LAG and suddenly you drop 1/N of the traffic (where N might be 2 == 50% loss). More common was finding LAG/ECMP where 1/N of the links was oversubscribed but the others were fine due to unequal traffic distribution.
> noisy RF, industrial
Pretty uncommon in my experience as most of these environments are not as cost optimized and know already that they will be in weird RF/electrical environments so they just use wired ethernet (and even shielded cable for really janky situations).
While there are strategies to reduce congestion on backbone links, they are not commonly implemented and even less commonly implemented well and sometimes implemented intentionally poorly.
Remember, 75% of content now comes from a DC within 1 AS hop of your ISP.
For non-western economies reliant on mobile IP, there is a level of loss and congestion being seen in the radio packet layer but that layer is usually dealing with it.
Starlink has worked out how to manage it's TCP loss issues despite steering dishy between sats in orbit in a 15 second timeslot model.
I'm not trying to dismiss this work: I don't understand the use-case because from my perspective (admittedly in the western economy, but in asia and exposed to the other kind of internet for less developed places) this level of sustained packetloss isn't usual, And when it is, the link layer is usually doing something like FEC to take care of it.
Error correction codes on L4 level are generally only useful for very low latency situations, since if you can wait for one RTT, you can just have the original sender retransmit the exact packets that got lost, which is inherently more efficient than any ECC.
All I can see are hardcoded ping/pong “meow” messages going over a hardcoded client and server.
But maybe the ping/pong is part of the protocol?
It’s not clear.
Anyway, this redundancy-based protocol doesn’t seem to take into account that too many packets over the network can be a cause of bad, “overloaded” network conditions.
Raptorq is a nice addition, though.
The patents are hopefully expiring soon.
Generic tornado codes are likely patent free, having been expired for a few years now: https://en.wikipedia.org/wiki/Tornado_code
EDIT: looking a bit deeper into this repo, it's really just a wrapper over the raptorq crate with a basic UDP layer on top. Nothing really novel here that I can see over `raptorq` itself.
One killer feature was multicast streaming of data. The streaming could do an extra 5% of broadcasting packets instead of several round trips and retransmissions.
Now that I think about it, I wish we actually used multicast.
SGI had good support for it in IRIX and I am pretty sure I saw one video streaming solution using it on a LAN to maximize the throughput possible.
Seems like a good tech that fell into disuse.
It has found good success in IPTV delivery inside provider's own networks though. Cameras at casinos/hotels/Transit systems and the like too.
There's 2^112 possible global multicast addresses with ipv6 as well (1). Though yeah, you'll still have queuing overloads as well and other issues.
1. https://learningnetwork.cisco.com/s/question/0D53i00000Kt0EK...
That doesn't help if all routers have it turned off.
Effectively, the only real "improvement" for the routed multicast case is you have more private multicast addresses to pick from.
Though I’d argue that 112 bits of random addresses makes the need for global registration largely unnecessary. Similarly to the rest of IPv6, the address is intentionally so large that it allows random IP generation with very low collision probability.
I wonder if it opens up new opportunities erasure coding replacements, e.g. RAID for disks, PARQ2 for data recovery, or as Reed-Solomon replacements for comms?
For adding redundant blocks to a read-mostly/only file (e.g. to correct for sector errors) they could be useful indeed. as that's a case where you might have a few dozen correction packets protecting millions. I'd really like to see some FS develop support for file protection because on SSDs I'm seeing a LOT more random sector failures that disk failures.
RS codes are optimal for erasures so you really only want to use something else where there are so many packets in the group that RS code performance would be poor... or where a rateless code would be useful.
With raptorq, the client needs to receive 102 packets (I believe). And this can be any combination of original 100 packets and the large number of potential recovery packets.
Is 10% loss common for backbone networks? Maybe if you’re dealing with horribly underprovisioned routers that don’t have enough RAM for how much traffic they’re processing? Not sure otherwise the use case…
The challenge of course is that this doesn’t work so well:
1. Packets come in discrete chunks but FEC works best on a stream of data at the bit level (at least in a transmission context)
2. Lossy links will have their own understanding of the PHY that can’t be explained to applications / needs to happen much more quickly.
3. There’s information that can’t be corrupted in the packet full stop (ie the IP and maybe TCP framing).
2 is really the killer - applications can’t respond to noise issues like the MAC can to issues with the WiFi link for example. And higher layers can’t know about all the intermediary hops that might exist and how to characterize that to tune the FEC parameters.
3 is also a killer because it means you still need to apply FEC to the PHY to guarantee the framing packet and TCP requires checksums to pass on the packet. It’s really really difficult to use FEC at the application layer in an end to end way. It’s really something that works best at the PHY level for transmissions, at least from the brief time I’ve spent thinking on it. Possible I messed something up in my analysis of course.
I do also think that exposing some knobs for the quantity of FEC to apply to a higher level might make sense; some packets are 100% fine to drop, others are very much not. Standards trying to do this exist, but don't seem to be widely supported.
Middle boxes only need to speak IP, which they do already, and not block nyxpsi packets (which they probably do).
Edit: It does.
> And all it takes to win the game is to transmit classical bits with digital error correction using hidden variables?
Deep Space Communications, Satellite Network Reliability, First Responder Communication in Disaster Zones, Search and Rescue Operations Communication, Disaster Relief Network Communication, Underwater Sensors