Timeliness without datagrams using QUIC
quic.video
quic.video
- TCP will waste roundtrips on handshakes. And then some extra on MTU discovery.
- TCP will keep trying to transmit data even if it's no longer useful (same issue as with real-time multimedia).
- If you move into a location with worse coverage, your latency increases, but TCP will assume packet loss due to congestion and reduce bandwidth. And in general, loss-based congestion control just doesn't work at this point.
- Load balancers and middleboxes (and HTTP servers, but that's another story) may disconnect you randomly because hey, you haven't responded for four seconds, you are probably no longer there anyway, right?
- You can't interpret the data you've got until you have all of it - because TCP will split packets with no regards to data structure. Which is twice as sad when all of your data would actually fit in 1200 bytes.
That's useful for some things, but often you can make use of package n+1, even when package n hasn't arrived, yet.
For example, when you are transferring a large file, you could use erasure encoding to just automatically deal with 5% package loss. (Or you could use a fountain code to deal with variable packet loss, and the sender just keeps sending until the receiver says "I'm done" for the whole file, instead of ack-ing individual packages.
Fountain codes are how deep space probes send their data back. Latency is pretty terrible out to Jupiter or Mars.)
That’s one of the reasons people use TCP.
> For example, when you are transferring a large file, you could use erasure encoding to just automatically deal with 5% package loss.
It’s never that simple. You can’t just add some erasure coding and have it automatically solve your problems. You now have to build an entire protocol around those packets to determine and track their order. You also need mechanisms to handle the case where packet loss exceeds what you can recover, which involves either restarting the transfer or a retransmission mechanism.
The number of little details you have to handle quickly explodes in complexity. Even in the best case scenario, you’d be paying a price to handle erasure coding on one end, the extra bandwidth of the overhead, and then decoding on the receiving end.
That’s a lot of complexity, engineering, debugging, and opportunities to introduce bugs, and for what gain? In the file transfer example, what would you actually gain by rolling your own entire transmission protocol with all this overhead?
> Fountain codes are how deep space probes send their data back. Latency is pretty terrible out to Jupiter or Mars.)
Fountain codes and even general erasure codes are not the right tool for this job. The loss of a digital packetized channel across the internet is very different than a noisy analog channel sent through space.
Yes, and you'd want to mostly build this extra protocol only once, and stick it into a library.
> You also need mechanisms to handle the case where packet loss exceeds what you can recover, which involves either restarting the transfer or a retransmission mechanism.
Or you can use a fountain code, and just keep transmitting.
> The number of little details you have to handle quickly explodes in complexity. Even in the best case scenario, you’d be paying a price to handle erasure coding on one end, the extra bandwidth of the overhead, and then decoding on the receiving end.
Well, you can also use a simpler mechanism: you transmit all the packages from A to B once, then at the end B tells A which packages get lost, and A sends those again. Repeat until you have everything.
That way needs more back-and-forth communication, but doesn't need any fancy error correcting code.
> Fountain codes and even general erasure codes are not the right tool for this job. The loss of a digital packetized channel across the internet is very different than a noisy analog channel sent through space.
Yes, the loss model is different. However you can eg interleave your bits to get something that close enough to work.
The main reason you don't need to use error correcting codes, is that transmitting feedback to the sender is typically a lot cheaper on the internet than in outer space. Not least because even bad latencies are typically measured in seconds at most, not minutes or hours.
(You could however use these codes when for some reason you have a very asymmetrical link. Eg if you have internet via geo-stationary satellite, or if one sender is broadcast to lots and lots of different receivers, and doesn't want to deal with each of them individually.)
This is actually very nice for an application developer. Probably a tiny bit less nice if you are running popular services taking 100k req/s.
That's why I would suggest that were possible the kernel should only expose something at most as high level as UDP, and we build the TCP abstractions on top of that in user level. Of course, you wouldn't want to re-write everything from scratch all the time: you'd use libraries.
The nice thing about user level libraries is that you can swap them out for better versions or for different abstractions without any privileged access.
But for the "simple well-designed" - the problem is, the whole thing is a black box for application developer, and not a bugless one. So much of the "well-designed" part is to figure out when is the right time to pull modem's proverbial power cord. If you don't, it'll gladly eat up your battery in 48 hours after a cell handover goes wrong with no warning.
The carriers require you to allow them to force a firmware upgrade on you. This kills your battery dead long before anything you would do on the device will.
Technically, however, 10 year battery life is generally unreasonable. Most batteries that exist self-discharge over 10 years. 5 years is doable, but requires quite solid engineering. 3 years is pretty reasonable for most things.
If you want 10 year lifetimes, I think you specifically need lithium thionyl chloride batteries.
If your market is large enough that managing your own infrastructure is attractive (vs cellular), you were also able to afford a few EEs and RF guys 30 years ago, and your needs are already met by a competitive ecosystem of players who have been walking this path for a long time. It doesn't leave much to compete on except price, and you don't want to compete on price.
> If your market is large enough that managing your own infrastructure is attractive
This. I imagine proprietary is perfect when you are deploying in a single country with single RF plan. And your customer utility company that is willing to maintain a bit of extra infrastructure. But if you want to deploy on several sites scattered across all ITU regions, and each site is a huge plant with devices from maybe fifty different suppliers, it's just a no-go.
1) Managerial--Just like back in the old days when the carriers stupidly fought the change from voice to data they are stupidly fighting the change from humans using the network to machines using the network. And it's just as stupid this time around and it will be just as stupidly profitable for them after the fact.
2) The carriers refuse to support older firmware so you have to be able to do multi-megabyte firmware updates every 3 months. This is anathema to anything running on a battery as your firmware updates are like 2 to 4 orders of magnitude(not exaggerating) larger than the total actual data your device will transmit over its lifetime.
The carrier cranio-rectal impaction is why LoRA still continues to exist and expand.
* the telecom can and does blacklist IMEI numbers for unsigned firmware
* it is illegal to fix bugs that cause eSIM hardware lock-ups
* You can catch g05t5 on the carrier VPN network sandbox hitting each client
Sometimes it is better to wipe the slate clean, and start again... LoRAWAN certainly does make sense in some use-case applications.
What does that mean?
Have a great day, =3
Don't lower levels of wifi below TCP now do often their own retransmit, even if it is just from weak/noisy signal and not a sensed collision?
This is especially true when using TLS (which is at least a major use case nowadays if not the default assumption); the TCP handshake has to complete before the TLS handshake can occur, whereas QUIC has protocol-level support for doing TLS negotiation as part of the initial handshake. Having the network protocol and the encryption be composable rather than tightly coupled feels more elegant, it's hard for me to convince myself that being able to save a round trip per connection isn't a more practical benefit in a world where we now expect pretty much everything to be encrypted in transit.
Yet, it's not.
That kind of elegance ends up in application level implementations like STARTTLS and socket data becoming unusable on the middle of a connection while both parties scramble to decide what to throw away and what to interpret. It does really look like the real thing, but it's decoupling things that are very coupled by definition.
Naive UDP protocols have reachability problems in numerous scenarios, variable costs can balloon (concurrency with polling links is dumb), and can open a hole in your infrastructure at scale without special equipment mods. Only if you have a tier 3 or better WAN trunk or cloud-center should you even consider something silly like QUIC.
YMMV, and I love the AstroTurf on YC... lol ;)
Thanks, I'll see myself out... lol =)
Indeed, until the problems it creates begin to feature.
Have a great day =3
—G. Marx
I recall my computer networking professor describing this as "Best-effort means never having to say you're sorry". In practice, best-effort does not mean you try your hardest to make sure the message gets from A to B, it means that you made an effort. Router in the path was congested? Link flap leading to blackholing on the order of 50ms before fast reroute kicks in? Oh well, we tried.
Meanwhile, TCP's reliable delivery will retry several times and will present an in-order data stream to the application.
Reliable vs unreliable might be bad terminology, but I don't think best-effort is any better.
My experience with unreliable systems is that they're great something like 95% of the time, and they're great for raw throughput, but there are many cases where that last 5% makes a huge difference.
You can build a 'reliable' protocol on top of UDP, and still not get TCP.
Eg if you want to transfer a large file that you know up front, then TCP's streaming mechanism doesn't make too much sense. You could use something like UDP to send the whole file from A to B in little chunks once, and at the end B can tell A what (numbered) chunks she's missing.
There's no reason to hold off on sending chunk n+1 of the file, just because chunk n hasn't arrived yet.
Congestion control comes to mind -- you don't necessarily know what rate the network supports if you don't have a feedback mechanism to let you know when you're sending too fast. Congestion control is one of those things where sure, you can individually cheat and possibly achieve better performance at everyone else's expense, but if everyone does it, then you'll run into congestive collapse.
> You can build a 'reliable' protocol on top of UDP, and still not get TCP.
I agree -- there are reliable protocols running on top of UDP (e.g. QUIC, SCTP) that do not behave exactly like TCP. You don't need an in-order stream in the described use case of bulk file transfer. You certainly don't need head-of-line blocking.
But there are many details and interactions that you and I wouldn't realize or get right on the first try. I would rather not relearn all of those lessons from the past 50+ years.
Oh, the model I had in mind was not that everyone should write their network code from scratch all the time, but rather that everything that's higher level than datagrams should be handled by unprivileged library code instead of privileged kernel level code.
If speed is an issue, modern Linux can do wonders with eBPF and io_uring, I guess? I'm taking my inspiration from the exokernel folks who believed that abstractions have no place in the operating system kernel.
the unix model is different: it's basically premised on a minicomputer server which would be running a diverse set of independent services where isolation is desired, and where it makes sense for a privileged entity to provide standardized services. services whose API has been both stable and efficient for more than a couple decades.
I think it's kind of like cloud: outsourcing that makes sense at the lower-scale of hosting alternatives. but once you get to a particular scale, you can and should take everything into your own hands, and can expect to obtain some greater efficiency, agility, autonomy.
Exokernels provide isolation. Secure multiplexing is actually the only thing they do.
> and where it makes sense for a privileged entity to provide standardized services. services whose API has been both stable and efficient for more than a couple decades.
Yes, standardisation is great. Libraries can do that standardisation. Why do you need standardisation at the kernel level?
That kind of standardisation is eg what we are doing with libc: memcpy has a stable interface, but how it's implemented depends on the underlying hardware; the kernel does not impose an abstraction.
That said, the typical reason why TCP doesn't send packet N+1 is because its congestion window is full.
There is a related problem known as head-of-line blocking where the application won't receive packet N+1 from the kernel until packet N has been received, as a consequence of TCP delivering that in-order stream of bytes.
Basically, when transmitting a file, in principle you could just keep sending n+k, even if the n-th package hasn't been received or has been dropped. No matter how large k is.
You can take your sweet time fixing the missing packages in the middle, as long as the overall file transfer doesn't get delayed.
The closest I can think of is the old FSP protocol, which never really saw wide use. The client would request each individual chunk by offset, and if a chunk got lost, it could re-request it. But that's not quite the same thing.
"Effort" generally refers to some sort of persistence in the face of difficulty. Dropping a packet upon encountering a resource problem isn't effort, let alone best effort.
The way "best effort" is used in networking is quite at odds with the "best efforts" legal/business term, which denotes something short of a firm commitment, but not outright flaking off.
Separately from the delivery question, the checksums in UDP (and TCP!) also poorly assure integrity when datagrams are delivered. They only somewhat improve on the hardware.
The dropping part isn’t the effort; the forwarding part is.
> the checksums in UDP (and TCP!) also poorly assure integrity when datagrams are delivered.
That’s true, but it’s becoming less of a problem with ubiquitous encryption these days (at least on WAN connections).
Anyway, so the question is, if typical IP forwarding is "best effort" ... what is an example of poor effort, and what exhibits it?
That's what I thought "unreliable" meant? I can't really tell what misconception you are trying to avoid.
Ultimately, TCP vs UDP per se is rarely the right question to ask. But that is often the only configuration knob available, or at least the only way to get away from TCP-based protocols and their overhead is to switch over to raw UDP as though it were an application protocol unto itself.
Such naive use of UDP contributes to the stigma. If you send a piece of data only once, there's a nontrivial chance it won't get to its destination. If you never do any verification that the other side is available, misconfiguration or infrastructure changes can lead to all your packets going to ground and the sender being completely unaware. I've seen this happen many times and of course the only solution (considered or even available in a pinch) is to ditch UDP and use TCP because at least the latter "works". You can say "well it's UDP, what did you expect?" but unfortunately while that may have been meant to spur some deeper thought, it often just leads to the person who hears it writing off UDP entirely.
Robust protocol design takes time and effort regardless of transport protocol chosen, but a lot of developers give it short shrift. Lacking care or deeper understanding, they blame UDP and eschew its use even when somebody comes along who does know how to use it effectively.
I think the biggest mistake people make with UDP protocols is allowing traffic amplification attacks.
Oversimplified.
Best effort is a dumb term. There’s no effort.
reliable = it either succeeds or you get an error after some time, unreliable = it may or may not succeed
I concur with the people who think "best effort" is not a good term. But perhaps TCP streams are not reliable enough for TCP to be rightly called a reliable stream protocol. As it turned out, it's not really possible to use TCP without a control channel, message chunking, and similar mechanisms for transmitting arbitrary large files. If it really offered reliable streams that would its primary use case.
best effort implies 'no, we're not going to be climb that curve and get try to get to 100% reliability, because that would actually be counterproductive from an engineering perspective, but we're going to go to pretty substantial lengths to deliver your packet'
Reliability is arguably more of a statement about the availability metrics of your underlying network; it doesn’t seem like a great summary for what TCP does. You can’t make an unreliable lower layer reliable with any protocol magic on top; you can just bundle the unreliability differently (e.g. by trading off an unreliability in delivery for an unreliability in timing).
UDP is at-most-once; TCP is exactly-once-in-order.
In the end, best-effort is just saying "unreliable" in a fussier way.
>But it does not make UDP inherently unreliable.
Isn't that exactly what it does make it?
If that's not it, then what woud an actual "inherently unreliable" design for such a protocol be? Calling an RNG to randomly decide whether to send the next packet?
> UDP makes its best effort
TCP tries really hard, too, you know.
but, "best-effort" implies that it's doing some effort to ensure delivery, when it's really dropping any packet that looks funny or is unlucky enough to hit a full buffer
i like "lossy", but this is definitely one of the two hard problems
In a datagram-first world we would have no issue bonding any number of data links with very high efficiency or seamlessly roaming across network boundaries without dropping connections. Many types of applications can handle out-of-order frames with zero overhead and would work much faster if written for the UDP model.
Clearly websites, audio and video generally don't work with out-of-order frames -- most people don't want dropped audio and video.
Some video games are happy to ignore missed packets, but when they can they are already written in UDP.
On the contrary, for me it’s hard to imagine video game with missed packets as the state can get out of sync too easily, you’ll need eventual consistency via retransmission or some clever data structures (I know least about this domain though)
Even congestion control can be optional for some applications with the right error correcting code. (Though if you have a narrow bottleneck somewhere in your connection, I guess it doesn't make too much sense to produce lots and lots more packages that will just be discarded at the bottleneck.)
Oh, it looks like UDP packages are also buffered?
Update: it looks like I learned something new today. UDP can also suffer from bufferbloat. (I thought UDP was done without buffers for some reason..)
So your argument is that software isn’t written well because TCP is too convenient, but we’re supposed to believe that a substantially more complicated datagram-first world would have perfectly robust and efficient software?
In practice, moving to less reliable transports doesn’t make software automatically more reliable or more efficient. It actually introduces a huge number of failure modes and complexities that teams would have to deal with.
I think you can make a better argument:
Your operating system should offer you something like UDP, but you can handle all the extra features you need on top of that to eg simulate something like TCP, at the level of an unprivileged library you link into your application.
That way is exactly as convenient for 'normal' programmers as the current world: you get something like TCP by default. But it's also easy for people to innovate and to swap out a different implementation, without needing access to the privileged innards of your kernel.
(* in keeping with the HN's title guideline: "Please use the original title, unless it is misleading or linkbait" - https://news.ycombinator.com/newsguidelines.html)
I dont agree with the premise of this article, UDP isnt for unreliability, it provides a tradeoff which trades speed and efficiency and provides best-efforts instead of guarantees.
It makes sense depending on your application. For example, if I have a real-time multi-player video game, and things fall behind, the items which fell behind no longer matter because the state of the game changed. Same thing for a high-speed trading application -- I only care about the most recent market data in some circumstances, not what happened 100ms ago.
- Local discovery (DHCP, slaac, UPnP, mDNS, tinc, bittorrent)
- Broadcasts (Local network streaming)
- Package encapsulation (wireguard, IPSec, OpenVPN, vlan)
Package encapsulation is bad over TCP because if the encapsulated data is TCP itself, you have congestion control twice. On congested networks, this results in extra slowdowns that can make the connection unusable.
I do wish QUIC allowed carrying streams that were useful for realtime in conjunction with allowing reliable streams. Using MPEG-TS over SRT to have it just spam metadata to handle the unreliableness is janky. It would be far nicer to have a reliable stream for metadata, then an unreliable one for realtime streaming.
UDP and TCP have different behaviour and different tradeoffs, you have to understand them before choosing one for your use case. That's basically it. No need for "Never do X" gatekeeping.
This is literally the point of the article: if you want to create a protocol over raw datagrams, you have to implement a lot of things that are very hard to get right, so you should just use QUIC instead, which does them for you.
I don't do this stuff any more, but I worked on OSI transport and remote operations service mapped to UDP and other protocols back in the 80s
Is the author saying that with QUIC I can send a "score update" for my game (periodic update) on a short-lived stream, and prevent retransmissions? I'll send an updated "score update" in a few seconds, so if the first one got lost, then I don't want it to waste bandwidth retransmitting. Especially I don't want it retransmitted after I've sent a newer update.
For something like a scoreboard you don't want the message to be discarded once any packet loss is detected. As long as message N is the most recent one, message N should be retried until it succeeds. You just want to stop retrying N once N+1 becomes available - or ideally even successfully delivered.
In your scoreboard example, you probably just want an actually reliable stream for that - it's not like scores are rolling in so fast that it's beneficial to deal with it as unreliable.
Video games tend to use UDP for the same reason everyone else mentioned does: timeliness. You want the most recent position of the various game objects now, and you don't give a shit about where they were 100ms ago.
The proposed solution of segmenting data into QUIC streams and mucking with priorities should work just fine for a game.
> Is that even supported across all the various gaming platforms?
QUIC itself is implemented in terms of datagrams, so if you have datagrams, you can have QUIC.
This is only true for games that can replicate their entire state in each packet.
There are many situations where this is infeasible and so you may be replicating diffs of the state, partial state, or even replicating the player inputs instead of any state at all.
In those cases the "latest" packet is not necessarily enough, the "timliness" property does not quite cover the requirements, and like with most things, it's a "it depends".
QUIC optionally promises you that, you are free to opt out. For example, take a look at the QUIC_SEND_FLAG_CANCEL_ON_LOSS flag on Microsoft's QUIC implementation.
Let me rephrase the problem here for a simple FPS game:
The entire world state (or, the subset that the client is supposed to know about) is provided in each packet sent from the server. The last few actions with timestamps are in each packet sent from the client. You always want to have the latest of each, with lowest possible latency. Both sides send one packet per tick.
You do not want retransmission based on the knowledge that a packet was lost (the RTT is way too long for the information that a packet was lost to ever be useful, you just retransmit everything every tick), you do not want congestion control (total bandwidth is negligible, and if there is too much packet loss to maintain what is required, there is no possible solution to maintain sufficient performance and you shouldn't even try), and none of the other features talked about in the post add anything of value, either.
It reads like someone really likes QUIC, it fit well into their problems, and they are a bit too enthusiastic about evangelizing it.
You’re right that a custom reliable UDP solution is going to wind up QUIC-like. On the other hand it’s what games have been doing for over 20 years. It’s not particularly difficult to write a custom layer that does exactly what a given project needs.
I don’t enough about QUIC to know if it adds unnecessary complexity or not.
You also don't want to share the routing performance, congestion treatment, OS queue, and probably a lot of other stuff.
For instance, suppose that you wanted a special traffic class that's prioritized over everything else. You really don't want just anyone to be able to send you that kind of traffic, since a DoS attack can easily starve all the other queues.
Your best bet is usually to match traffic based on other characteristics, e.g. port × protocol.
Still, it seems like it would be a useful idea to have an IP bit for "minimize buffering and drop instead, packet is time-sensitive", and that seems safe since the user is requesting fewer resources, not more.
Think about it this way: a 100Gbps link can process a 1500 byte packet in about 0.12 microsecond. If you have an average of 1,000 packets in that buffer at steady-state, that buffer is contributing a fraction of a millisecond to your overall latency. Meanwhile, if your home router's buffer has the same 1,000 packets for a 1Gbps link, that's 12 milliseconds of latency.
If you had a way to tell a router to never buffer packets, you'd encounter packet loss much sooner, and for no good reason. Buffers are great for smoothing out (small) bursts of traffic (although if you have large bursts on very short timescales, you can easily overwhelm these buffers -- this is why you typically have pacing on the OS level).
If instead you want the router to add no more than X amount of latency, that's suddenly much harder to dictate. (The "correct" value of X depends on your application requirements as well as the number of hops through the network between you and the server you're talking to.)
(Also, while you might think that this should be strictly beneficial for network operators since the user wants fewer resources, encoding exceptions like this into router policy ends up using more of a specialized and expensive kind of memory called TCAM. Assuming that it's even possible to do such a thing.)
A lot of Enterprise messaging is based on UDP, I think on the presumption that corporate networks are just simply going to be a lot more reliable
https://en.m.wikipedia.org/wiki/IEEE_802.2
Back in the olden times, the Windows 3.11 and Windows 95 could even run services atop it using another protocol: https://en.m.wikipedia.org/wiki/NetBIOS_Frames
In fact, the link layer protocols that dealt with congestion and access to media more meticulously, eg Token Ring or FDDI, were generally much more efficient with the use of the media - using near 100% of the potential bandwidth, whereas Ethernet on a shared medium already gets quite inefficient at 30-40% utilization due to retries caused by collisions.
However, the trade off between additional complexity (and thus bigger cost and difficulty of troubleshooting) was such that the much simpler and dumber Ethernet has won.
However, a lot of similar principles are there in Fibre Channel family protocols, which are still used in some data center-specific protocols.
or Never use on a network where congestion is above a certain level
or Never us on a network where this parameter is above a certain level - like network latency
or only use on a LAN not a WAN ....?
Most end-user software ends up running on a variety of oversubscribed wifi and cell networks
One thing you can do is build in "bad network emulation" into your software, allowing the various pieces to drop packets, delay packets, reorder packets, etc. and make sure it still behaves reasonably with all this garbage turned up.
But people do appreciate software that doesn't have this restriction, and even go out of their way to talk about it (more on mobile networks than on wired connections).
https://en.wikipedia.org/wiki/Real-Time_Streaming_Protocol
Cheers =3
RTP is a core part of WebRTC, for example.
When you're doing a video call in a web browser, you're using WebRTC, including RTP. In fact, this RTP-via-WebRTC is the only way to send UDP packets from JavaScript!
RTSP is still used by older streaming systems and hardware ecosystems that are slow to change, such as network-connected security cameras. But in newer applications, WebRTC has mostly replaced it. Of course, the QUIC effort is in part an attempt to replace WebRTC, so the wheel continues to turn!
https://github.com/mpromonet/webrtc-streamer.git
I remain unconvinced UDP based streams will ultimately remain in the long-term, but webRTC certainly made it easier to peer a connection. ;)
UDP adds minimum over raw sockets so that you don't need root privileges. Other protocols are build on top of UDP.
It's better to use existing not-TCP protocols instead of UDP when the need arises instead of making your own. Especially for streaming.
As time has progressed increased from nothing to fec to dual-streaming and offset-streaming to RIST and SRT depending on the criticality.
On the other hand I've seen people try to use TCP (with rtmp) and fail miserably. Never* use TCP.
Or you know, use the right tool for the right job.
1. Routers do all kinds of terrible things with TCP, causing high latency, and poor performance. Routers do not do this to nearly the same extent with UDP
2. Operating systems have a tendency to buffer for high lengths of time, resulting in very poor performance due to high latency. TCP is often seriously unusable for deployment on a random clients default setup. Getting caught out by Nagle is a classic mistake, its one of the first things to look for in a project suffering from tcp issues
3. TCP is stream based, which I don't think has ever been what I want. You have to reimplement your own protocol on top of TCP anyway to introduce message frames
4. The model of network failures that TCP works well for is a bit naive, network failures tend to cluster together making the reliability guarantees not that useful a lot of the time. Failures don't tend to be statistically independent, and your connection will drop requiring you to start again anyway
5. TCP's backoff model on packet failures is both incredibly aggressive, and mismatched for a flaky physical layer. Even a tiny % of packet loss can make your performance unusable, to the point where the concept of using TCP is completely unworkable
Its also worth noting that people use "unreliable" to mean UDP for its promptness guarantees, because reliable = TCP, and unreliable = UDP
QUIC and TCP actively don't meet the needs of certain applications - its worth examining a use case that's kind of glossed over in the article: Videogames
I think this article misses the point strongly here by ignoring this kind of use case, because in many domains you have a performance and fault model that are simply not well matched by a protocol like TCP or QUIC. None of the features on the protocol list are things that you especially need or even can implement for videogames (you really want to encrypt player positions?). In a game, your update rate might be 1KB/s - absolutely tiny. If more than N packets get dropped - under TCP or UDP (or quic) - because games are a hard realtime system you're screwed, and there's nothing you can do about it no matter what protocol you're using. If you use QUIC, the server will attempt to send the packet again which.... is completely pointless, and now you're stuck waiting for potentially a whole queue of packets to send if your network hiccups for a second, with presumably whatever congestion control QUIC implements, so your game lags even more once your network recovers. Ick! Should we have a separate queue for every packet?
Videogame networking protocols are built to tolerate the loss of a certain number of packets within a certain timeframe (eg 1 every 200ms), and this system has to be extremely tightly integrated into the game architecture to maintain your hard realtime guarantees. Adding quic is just overhead, because the reliability that QUIC provides, and the reliability that games need, are not the same kind of reliability
Congestion in a videogame with low bandwidths is extremely unlikely. The issue is that network protocols have no way to know if a dropped packet is because of congestion, or because of a flaky underlying connection. Videogames assume a priori that you do not have congestion (otherwise your game is unplayable), so all recoverable networking failures are 1 off transient network failures of less than a handful of packets by definition. When you drop a packet in a videogame, the server may increase its update rate to catch you up via time dilation, rather than in a protocol like TCP/QUIC which will reduce its update rate. A well designed game built on UDP tolerates a slightly flakey connection. If you use TCP or QUIC, you'll run into problems. QUIC isn't terrible, but its not good for this kind of application, and we shouldn't pretend its fine
For more information about a good game networking system, see this video: https://www.youtube.com/watch?v=odSBJ49rzDo, and it goes over pretty in detail why you shouldn't use something like QUIC
What things are you thinking of here?
> 2. Operating systems have a tendency to buffer for high lengths of time, resulting in very poor performance due to high latency.
Do they? I can only think of Nagle's algorithm causing any potential buffering delay on the OS level, and you can deactivate that via TCP_NODELAY on most OSes.
> 5. TCP's backoff model on packet failures is both incredibly aggressive, and mismatched for a flaky physical layer. Even a tiny % of packet loss can make your performance unusable, to the point where the concept of using TCP is completely unworkable
That's a property of your specific TCP implementation's congestion control algorithm, nothing inherent to TCP. TCP BBR is quite resistant to non-congestion-induced packet loss, for example [1].
> If you use QUIC, the server will attempt to send the packet again which.... is completely pointless
Isn't one of QUIC's features that you can cancel pending streams, which explicitly solves that issue? In your implementation, if you send update type x for frame n+1, you could just cancel all pending updates of type x if every update contains your complete new state.
[1] https://atoonk.medium.com/tcp-bbr-exploring-tcp-congestion-c...
I'm always curious where frame proponents are going to buffer in-progress frames. OS? I have huge memory concerns. Also take note you want to drop in-order requirement and it means multiple in-progress frames. It's a very nice DoS vector.
And WebSockets do work fine while you're testing.
But WebSockets are TCP, so when you roll things out to real-world users, latency is higher than it should be a lot of the time. Connections randomly drop. You have to try to figure out why different OS configurations are behaving differently. You start building your own keepalive logic. You start trying to figure out how to build metrics and monitoring for audio-over-WebSockets.
The answer is to use UDP, and use a protocol built on top of UDP designed for real-time media (WebRTC).
It's been interesting to me that I've had this exact same conversation with several dozen engineers over the past few months. It's a good reminder that things that seem obvious when you've been doing something a long time aren't obvious to people new to a domain. Even if they are very experienced in some adjacent domain. (In this case, mobile app and web app development.) It's made me think about where my knowledge boundaries are and what "obvious" things I'm ignorant about. (Lots of them, I'm sure.)
If you're interested in WebSockets, WebRTC, and why UDP is the right low-level approach for real-time media, I wrote a a primer about that a few months ago here:
https://www.daily.co/blog/how-to-talk-to-an-llm-with-your-vo...
A reliable, unlimited-length, message-based protocol.
With TCP there's a million users that throw out the stream aspect and implement messages on top of it. And with UDP people implement reliability and the ability to transmit >1 MTU.
So much time wasted reinventing the wheel.
If you just keep everything in the same order, then you never have to worry about any re-ordering.
I can see why people picked streams as the one-size-fits-all-(but-badly) abstraction.
There’s a bit of reinvention needed, sure, but adding a framing layer is about the easiest task in networking. If you want a standard way of doing it, these days you can use a WebSocket.
Edit to add: oh, I see, you want reliable but unordered messages. That would definitely be useful sometimes, but other times you do want ordering. If you don’t need ordering, isn’t that pretty much what QUIC does?
EDIT: Internet being broken is also why there's a grease extension for QUIC
This generalization isn’t true at all. Streaming data over TCP is extremely common in many applications.
When you download a large file, you’re streaming it to disk, for example.
I guess you could get pedantic and argue that an entire stream of something is actually one large message, but that’s really straining the definitions.
Nowadays perhaps video streaming, which may be lossy, so TCP is not needed for it (i.e. what is now called streaming does not have the reliability requirements of the TCP streams), might transfer more data on the Internet than file downloads, but that is not a desirable technical evolution, the replacement of downloading with streaming is just a means to extract more money from those who are satisfied by such a service.
For most applications some re-orderings of messages don't matter, and others would need special handling. So as a one-size-fits-all-(but-badly) abstraction you can use a stream.
> Do we have any actual streaming use of TCP, with purely streaming protocol.
But to give you a proper answer: the stream of keyboard inputs from the user to the server in Telnet (or SSH).
Wouldn't each input be a single albeit too short message? But this level of granularity really makes little sense...
You can also stick your higher level messages into a structure that's more complicated than a stream, eg you can stick them into a tree. Anything you can serialise, you can send.
Of course, the downside to this is that when you don't need these strong guarantees, you are paying for stuff you don't need.
(1) zero-copy, zero-allocation request processing
(2) up to a 2x latency reduction by intermingling networking and actual work
(3) more cache friendliness
(4) better performance characteristics on composition (multiple stages which all have to batch their requests and responses will balloon perceived latency)
If you have a simple system (only a few layers of networking), low QPS (under a million), small requests (average under 1KB, max under 1MB), and reasonable latency requirements (no human user can tell a microsecond from a millisecond), just batch everything and be done with it. It's not worth the engineering costs to do anything fancy when a mid-tier laptop can run your service with the dumb implementation.
As soon as those features start to matter, streaming starts to make more sense. I normally see it being used for cost reasons in very popular services, when every latency improvement matters for a given application, when you don't have room to buffer the whole request, or to create very complicated networked systems.
If you want large messages just redefine the length from 16 to 32 bits.
It's been used for millions of messages per day for 25 years and hasn't been changed in a long time.
Admittedly it isn't as common as TCP and that's a shame. But it's out there, it's public, it's minimal and it makes sense.
This protocol guarantees too much. It's still a stream.
This now means that innovations based on Quic can occur
Also, the overhead of a UDP packet is 8 bytes total, in 4 16-bit fields:
- source port
- dest port
- length
- checksum
So, we can save a max of 8 bytes per packet - how many of these can we practically save? The initial connection requires source and dest ports, but also sets up other connection IDs so theoretically you could save them on subsequent packets (but that's pure speculation, with zero due diligence on my part). It would require operational changes and would make any NATs tricky (so practically only for IPv6) Length and checksum maybe you can save - QUIC has multiple frames per UDP datagram so there must be another length field there (anyway, IETF is busily inventing a UDP options format on the basis that the UDP length field is redundant to packet length, so can be used to point at an options trailer). QUIC has integrity protection but it doesn't seem to apply to all packet types - however I guess a checksum could be retained for those that don't only.
So in sum maybe you could save up to 8 bytes per packet, but it would still be a lot of work to do so (removing port numbers especially)
Is it possible for an out of order packet to get delayed by a millisecond? A second? A minute? An hour?
Have you written code to be robust to this?
If every message is getting dropped, the set of unacked messages fills up and the connection stops accepting messages from the application, similar to TCP in that situation.
I mostly work with TCP so haven't had to deal with unreliable channels generally, but I do use a similar approach at the application level for reliable delivery across reconnects.
And due to various political & cultural & industrial lobbying, more flexible but complex on their face approaches (OSI) were rejected.
TCP meant you could easily attach a serial terminal to a stream, a text protocol meant you could interact or debug a protocol by throwing together a sandwich, an undergrad, and a serial terminal (coffee too if they do a good job). You could also in theory use something like SMTP from a terminal attached over a TIP (which provided dial-in connection into ARPAnet where you told the TIP what host and what port you wanted).
If you look through some of the old protocol definitions you'll note that a bunch of them actually refer to simplest mode of TELNET as base.
Then in extension of arguably the same behaviour that blocked GOSIP mandate we have pretty much reduced internet to TCP, UDP, and bits of ICMP, because you can't depend on plain IP connectivity in a world of NAT and broken middleboxes.
I've seen a few implementations of a reliable UDP messaging stream (using multicast or broadcast). It uses a sequencer to ensure all clients receive the correct message and order.
It can be extremely fast and reliable.
a) It added yet another area of expertise to my portfolio.
b) I did it because it was needed to facilitate another product in the company where I worked. The alternative at the time would be paying $350,000 for middleware from Vendor-X. Not acceptable.
c) Vendor-X people noticed us, loved what we did and made us a partner to resell their middleware and we've made a ton of money on deployment, configuring and consulting services.Where to keep in-progress frame data? I case of TCP it's an app and buffer for only a single frame or app even could make "stream processing" for in-progress frame.
So why the F.. would I bother with anything else?
The reality, of course, is somewhat different and muddy. For example, if you have network outage following later by a software crash or a reboot, then all the TCP buffer worth of data (several kilobytes or upto some megabytes - depends on your tuning) is mercilessly dropped. And your application thinks that just because you used TCP the data must have been reliably delieved. To combat this you have to implement some kind of serializing and acking - but the article scoffs us that we are too dumb to implement anything besides basic stream /s
I'm not arguing that TCP is useless, just that UDP has its place and we - mere mortals - can use is too. Where appropriate.
When you are using UDP the correct way to handle out of order delivery is to just ignore the older packet. It's old and consequently out of date.
Figuring out how to solve any mess caused by mistransmission necessarily has to be done at the application level because that's where the most up to date data is.
I will point out that "author of the article" is one of the core contributors in the IETF Media-over-QUIC working group (which is an effort to standardize how one might build these real-time applications over QUIC) and has been working in the real-time media protocols space for 10+ years.
The author recognizes that the title is clickbait, but the key point of the article is that you probably don't want to use raw UDP in most cases. Not that UDP is inherently bad (otherwise he wouldn't be working on improving QUIC).
Possibly you just haven't seen this problem yet because your game isn't high-profile enough for them to bother with and it doesn't provide enough amusement for them to get a kick out of it.