Hooray for the Sockets Interface
blog.apnic.net
blog.apnic.net
TCP/IP won (IMHO) because the implementation was open source, and because no one was going to make a buck by locking others out (or charging them for the right to build an implementation). Plus there was a working reference implementation available for the taking
Everything else was commercial, while by being prevented to sell UNIX, AT&T was fine with giving its source code to universities for a symbolic price for sending the tapes, and the Lion's book.
It changed quite fast the moment AT&T was allowed to claim ownership, however just like with IBM and PC clones, AT&T could no longer take control over UNIX.
[1] https://en.wikipedia.org/wiki/FTP_Software [2] https://en.wikipedia.org/wiki/NetManage
1, bidirectional sockets probably should've been ≥2 fd's, not 1.
2, non-blocking semantics and poll/select suck.
Can't blame anyone in 1983 for not getting "async I/O" right since we're still struggling; at least now we have some decent answers (like io_uring on Linux.)
If we're just strictly talking about low level things, another good one would be synchronization primitives, which I guess we now have some of in Linux verbatim at this point, if only for the sake of emulation. (And of course futex2.)
Think of ctrl-alt-f7, f8 etc. This still works today, you can run X servers as different users on different VT numbers and switch between them. All that's missing is better UI.
I was talking about this:
https://news.ycombinator.com/item?id=49093002
And it winds up being mostly a thing you have to implement in the compositor, but none of them do so far to my knowledge.
And of course the same issue applies if you wanted multiple physical seats juggling sessions between them.
It's kind of the opposite, systemd-logind provides better seat management than without. I attempted to implement the same concept using libseat and seatd unsuccessfully though I do not remember exactly what I got caught on. (not in kwin but a toy compositor; I have been at this problem for a little while now.)
That said, most desktop systems have only implemented systemd-logind in a limited way so far that basically just does what the VT system + display manager was already doing, that's what would need to be worked on in order to make this a reality. I was able to accomplish a prototype with only patches to Kwin, kfreerdp and plasma-login-manager, no need to patch systemd-logind or anything like that.
> Who even uses systemd?
Well for one thing, almost all of the major distributions; Ubuntu, Debian, Fedora/RHEL, Arch Linux, NixOS, openSuSE? and of course SteamOS now.
The only major non-systemd Linux distributions I can think of are Alpine Linux and Gentoo. And Android if you wanted to count that, but I think it is special enough to be considered its own OS that just is Linux-based.
I certainly wouldn't DIY this; nothing would support my DIY version, which would defeat the purpose.
Linux strengths: file system, open source ecosystem, ssh, lockless algorithm support
The main features that were better in Windows NT than in the UNIX-derived operating systems were inherited from the DEC VAX/VMS operating system (1978) (e.g. WaitForMultipleObjects) and a part of them had been inherited from the even earlier operating system DEC RSX-11M (1974-11) (Dave Cutler also had a major role in those operating systems, so there is nothing surprising about this; Microsoft had to pay a big compensation to DEC, on the order of $ 100M, for the features that were obviously taken from the DEC operating systems).
UNIX was very simplified in comparison with the operating systems that were used on bigger computers, and for some things its derivatives never caught up with those older systems, except after 2000. Only "futex" (2002) and "io_uring" (2019) have advanced the state of the art clearly beyond what already existed for AIO on the IBM mainframes around 1964/1966.
It's an evergreen interesting fact that Dave Cutler took perhaps a bit too much inspiration from prior VMS work to NT, but I also think there's no need to put an asterisk on it anymore than there is UNIX or anything else; it does ultimately stand on its own. Like great artists, great software architects steal.
Opaque handles are often a blessing.
Maybe UUID would be overkill (it certainly would have been back in the day of 16-bit unix machines) but something fairly large would have helped a lot.
Even just incrementing and rolling over a 16-bit value would have helped, instead of situations such as closing stdout and stderr, opening a file, and now random error logs sprinkle into your PDF or whatever.
Yes, "you are holding it wrong", but it should be hard to "hold it wrong", not easy.
I'd argue that it's not that the semantics suck, it's that there are too many of them. We have O_NONBLOCK, multiple multiplexers (for sane reasons--I don't begrudge 1980s folks for not thinking about fd set size and copy overhead for select(2) either), and others. What's worse, they don't all work with all FDs--not only are regular files not nonblock-able in the same way that sockets are, but all sorts of other FD-exposed capabilities (signalfds, timerfds, memfds, pidfds) do or don't support nonblocking semantics and multiplexing in all sorts of weird ways.
If those old system designers had stuck with keeping the async IO syscall space small (e.g. "you only get select/poll" or "you only get read(fds) and read_noblock(fds, timeout)") and consistent (by drawing a hard line at "if you expose something as an FD, it must support all APIs that generically handle FDs"), we would have ended up in a better place. Sure, that would have slowed down some implementations (e.g. vfs drivers), but would also have massively sped up development against a lot of these APIs.
Ah, well, hindsight is 20/20 I guess.
Non-blocking I/O: "O_NONBLOCK" with "EWOULDBLOCK" (or "EAGAIN")
Signal-driven I/O: "O_ASYNC" with "SIGIO"
Synchronous I/O multiplexing: "select()"
All 3 had various defects, especially with a large number of concurrent I/O actions.
In UNIX-derived operating systems, after 1983 there have been many other attempts to implement something better than these 3 (starting with System V "poll" and with POSIX AIO, and then with various incompatible approaches in Solaris, FreeBSD and Linux), but none were good enough and most were seriously inferior to methods of doing asynchronous I/O that existed in some IBM and DEC operating systems decades earlier.
In my opinion, only io_uring has finally solved the problem of multiplex asynchronous I/O in Linux, and in a manner much better than in all older operating systems. I consider all the many older alternatives that exist for liburing as obsolete.
Of course, Microsoft then went on to introduce IORing in 2022, continuing their legacy of never being afraid to adopt good ideas. But still - it would be wrong to suggest Linux is currently behind on async I/O. It hasn't been for years.
For a concrete example, I should be able to select/poll/epoll any file descriptor. The multiplexers can short-circuit return for things where readiness is inherent (e.g. block device files/vfs, I’m not asking for the moon a la waiting for NFS shares to report ready or something).
I should be able to issue non blocking reads on special file descriptors (timerfds, signalfds, eventfds, and so on).
This is analogous to my other unachievable fantasy: everything that exposes a filesystem interface must work with inotify/kqueue/whatever. No exceptions for procfs/NFS/etc.
At some point this fantasizing just becomes me griping about what made it to market/the worse-is-better philosophy generally, though. And yet it moves, I guess!
And yet, sockets are a terrible interface. They don't provide a way to get the details of the underlying connection for features like migration, checkpointing, or introspection. E.g. there is no way to get the current sequence number for TCP (there is "connection repair" mode now, but it's Linux-specific).
Well, you can say that sockets abstract the low-level details, but then these details hit you in the face when you need to do protocol-specific name resolution.
I now believe that we could have switched to something like IPv6 two decades ago if the socket interface simply allowed binding multiple address families to one socket and handled the name resolution internally.
Sockets also cemented the "one connection - one address" model that is _still_ dragging back the IPv6 adoption. MPTCP or QUIC are still barely supported.
I think Happy Eyeballs can't be implemented at that layer. It's inherently involved up to layer 7.
After all, why would you expect every program under the sun to deal with all the possible network protocols? Who on Earth would expect every software to be rewritten to use IPv6 or SCTP?!? We don't have special open() calls for files on floppy disks or RAM disks, after all.
Sockets already have these "automatic" behaviors. For example, for the connection source address selection. Very few programs explicitly select it, although it's possible.
GAI was a late addition, and it's also awkward in itself.
If you mean that one connection should be able to use multiple address families and addresses, then yes. It would be great, but at a much higher complexity than just making several connections at once and then discarding all but the first succcessful one.
Really? With unlimited ability to tamper with the stuff below the sockets layer, you couldn't make a reliable network? I'm surprised.
We do have ARIN IPv6 space, but we also have thousands of customers who do not know what that is or care about it either. Twice in twenty years now we have had a customer ask about it, and we have offered to work with them to help pilot support through our network (since if they have a need then that need could act as a standard against which to measure support) but always within a day they simply find some other way to satisfy their need without IPv6.
"Get more money from their customers" aside, we do still have a duty to prioritize the needs of our customers, and I'm not pulling budget away from initiatives to improve availability or increase bandwidth or lower latency simply to chase a buggy moving target technology with no light at the end of the tunnel by way of material improvement for our customers.
No "killer app". <1% marketing draw if you advertise it as a feature. Nobody's connection gets faster, or higher uptime, instead you just get more complicated support tickets when some website hasn't configured their AAAA records correctly or when some third party's IPv6 support (be that a site online or a home router or an application on the user's laptop) ruins a connection which works fine once the client disables IPv6 support on their end.
I would be overjoyed if the tech actually worked as a drop-in replacement, because that is a crazy amount of addressing space. I would be overjoyed if it were actually backwards compatible instead of a mishmash of dual stack and/or CGNAT nonsense. I would be overjoyed if the originally planned path MTU discovery and IPSEC support got off the ground, but those both died in horrible ugly ways.
As it is, it has picked up enough steam that some time over the next 15 years we will see a tipping point where the incentives to switch finally benefit all parties and the non-mobile industry starts to fall in line. Where customers actually begin to benefit directly and it becomes worth prioritizing this tech above what they currently perceive as more vital. Where a great enough percentage of our middle-mile and last-mile tech finally support it that we can survive migrating wholesale away from the vendors that do not. And where many of the hellacious pain points of trying to support it today will have been further sanded down and be less damaging to power through.
But I said something similar ten years ago and while the needle has moved since then it hasn't moved as far as I would have liked.
The problem was that the early IPv6 migration strategies all focused on getting IPv6 connectivity to clients as fast as possible (like 6to4). And once you got that IPv6, a lot of stuff just stopped working. It became all-or-nothing, and the option to just disable IPv6 and fix all the issues has always been there.
Happy Eyeballs was standardized criminally late, in 2012. For some reason, the idea that a connection can be something tentative was not a part of the mindset at all.
I'm guilty of that as much as everyone else, I remember spending a lot of time working on connection roaming for 3G/WiFi switching.
> Many early modems were acoustic couplers attached to telephone handsets using Velcro — one part was a microphone, the other a speaker. You connected by dialling the phone...
>
> Tools like the Unix-to-Unix Copy Program (UUCP) used scripts that called cu or similar tools to establish connections, then transferred data...
>
> UUCP was how we sent mail, read network news, and received (small) files. For much of the research community, it effectively was the network...
>
> Sockets changed all this.
>
> Source: https://blog.apnic.net/2026/07/28/hooray-for-the-sockets-interface
So much has changed... yet so much has not... All are pure marvels and based on dear miracles...A DNS name can have a LOT of A records associated but a socket has to pick one. This is a severe limitation.
Rather, it is a feature.
Standards as well as all manual pages and other documentation say that getaddrinfo() should be called via a loop by clients. And results show that almost all clients indeed do that. Therefore, on the server-side, you can do round-robin DNS, randomization (as suggested by the post linked), fall-backs, CDN things, and many other things.
You can hedge your bets and open a connection and send data to _all_ of them, and pick whichever returns faster.
This of course requires you to know about the application-layer protocol.
For HTTP, you need to restrict this strategy to GET methods, for example.
Hence why it can’t be part of the socket interface.
You can always build higher level abstractions on top of sockets if you need them.
This is even used by browsers, this trick even has a slightly creepy name: "Happy Eyeballs".
TCP does not require that the act of connecting must be pure, so you cannot indiscriminately apply Happy Eyeballs at the socket API level.
More concerning is the risk of exhausting server socket resources when probe-based connects don't hang up quickly if they don't want to use a connections. Lots of servers/load balancers aren't well-tuned to force-close connections if the first byte doesn't arrive within a short time.
Instead we converged on a very simple primitive that makes no such assumptions and can be composed into a wider variety of high level abstractions in user space.
TCP connection cookies solved this problem in 90-s! They fell out of use because dedicating a couple MB of RAM to track a few hundred thousand connections is not a big deal anymore.
For non-public networks, sure. For things like TCP-to-RS232 connectors that can only service one connection at a time. Happy Eyeballs needs to be disabled for them, but that's just one socket option.
And if the Happy Eyeballs protocol was standardized earlier, these kinds of apps arguably would have adapted to it.
I say that not out of purity concerns, but because of how DNS and TCP work. DNS multirecord selection (or not) should be up to the user. DNS timeouts should be surfaced granularly and differently from e.g. TCP SYNACK timeouts. Once a TCP stream is open, if a novice user opened it conceptually to a domain rather than an address, there's no intuitively correct answer to what happens if the domain's resolution changes while the socket's connected.
I've used this in networks and it enables amazingly easy endpoint mobility.
Cause what you're asking is not just to do the regular DNS lookup inside `connect`, but to have the OS manage roaming and continuously updating DNS while the connection stays up?
I don't know about that. I don't know. That's a lot of complexity deep in the kernel and ossified.
I'm happy with the QUIC solution that allows you to migrate IPs and it's all userspace. It would be very hard to evolve a network protocol that had to be crammed into 3+ different kernels.
You can have sockets without DNS. You can pick whatever strategy you want when there are multiple A records. You can use SRV records instead. And most importantly imo, it mirrors the listening API.
The ones we have suck or are non-standard.
Another option is for sockets to accept multiple addresses and address families in "connect" calls, so that only one connection wins.
War story: many years ago I was developing a service that runs computational tasks inside containers. That was when K8s didn't work well with cloud providers.
The service was written in Go and worked on AWS. Everything worked fine for me and for our production deployment. Then people started using it with a Python client that used the REST API directly, and for some reason some people couldn't reach it when they were working from home.
Reason: I misconfigured IPv6 on AWS by not enabling the default route, so all IPv6 connections were failing. Go has HappyEyeballs enabled by default, so it worked just fine.
But Python did not have it back then. So people with IPv4-only connectivity (including our office) had no problems. But people with IPv6 were getting connection failures.