Each socket should have two file descriptors
cr.yp.to
cr.yp.to
For starter, a | b |c doesn't create two file descriptors. It creates four, two per pipe. Then there is a problem with the central argument. For regular files, and in fact all method of creating file descriptor, except pipes, opening in RW mode creates a single descriptor. The reason for pipe() special behaior is two pronged:
1. The raison d'être of pipes is for sequential communication between producer/consumer processes. There is no other reason for their design. As such, it made sense to break the convention of a single FD.
2. At the time of their design, the 70s, memory footprint was a crucial part of any design. Thus, sharing buffers between the producer and consumer FD was primordial, and the best way to make it happen is to create the FD at the same time.
The other problem is that the pipe analogy for TCP is wrong. Which client / server protocol over TCP ever had single direction data transfer? Close to 100% of all TCP usages are between a single client and a single server, communicating both ways.
The design for socket is practical for their intended purpose. Trying to shoe-horn an artifical problem an a design will, unsuprinsingly, no yield proper result.
He never said there was no reason that sockets behaved differently than other descriptors, only that the reasons didn't warrant the broken abstraction.
And finally, I think you missed the place where he claims the abstraction breaks down. He's not complaining that you use "open" to get a file's descriptor and "pipe" to get a pipe's two descriptors, any more than it's weird to get a socket descriptor for "socket". He's saying that once you have the descriptor, the single file descriptor for sockets forces the OS to implement a socket-specific call simply to get a FIN sent properly.
The reality is that sockets could be represented not with integer descriptors but with character arrays or structs and the Unix interface wouldn't be any less incoherent than it is now with accepting descriptors, a select call that works for sockets and not files, setsockopt, shutdown, recv, connected Unix domain sockets, ioctls, and so on and so forth.
Still, it's possible that open/read/write/seek is just not the right abstraction in all cases. http://yarchive.net/comp/linux/everything_is_file.html is interesting reading.
I'd summarize those posts in two main points:
1) Linux isn't a research project.
2) Bringing a Plan 9-like realization of everything-is-a-file into modern Unix really just creates an even uglier chimera than Unix already is.
Design by effectively nobody can be at least as bad as design by committee.
As others pointed out, Plan 9 is much closer to a proper realization of the everything is a file design philosophy. Unsurprising, since it was grown entirely within Bell Labs under the supervision of the original Unix team. It's unfortunate it hasn't reached a point of being ready for mass adoption.
I don't think it's unfortunate, having used it for a month straight I'm not convinced at all it's the right thing, and in fact, is exactly the opposite direction from where I think we should be going (as far away from the filesystem as possible).
I suspect that Plan 9 could be a good dev environment, but is not a great deployment platform—it just doesn't offer enough better than developed *NIXes to port code over.
But it turns out that what we have is good enough, and it never gained traction.
It's a research platform, and no real effort has been made to turn it into something more practical. If someone were to pick it up and turn it into something people wouldn't hate using, I think we'd see a community on par with NetBSD or OpenBSD.
Here's another similar example of how "everything is a file" falls down that I wrote about in 2005 when complaining about terminals (http://www.advogato.org/person/habes/diary/6.html)
Then there's the whole mess of pseudo-terminals.
If you are like me, you might wonder at first why
pseudo-terminals are necessary. Why not just fork
a shell and communicate with it over stdin/stdout?
Well, the bad news there is that UNIX
special-cases terminals. They're not just
full-duplex byte streams, they also have to
support some special system calls for things like
setting the baud rate. And you might think that a
change of window size would be delivered to the
client program as an escape sequence. But no, it's
delivered as a signal (and incidentally, sent from
the terminal emulator as an ioctl()).
Of course, you can't send ioctls over a network,
so sending a resize from a telnet client to a
telnet server is done in a totally different way
(http://www.faqs.org/rfcs/rfc1073.html)Also, "everything is a file" is still true even when some files have additional operations possible on them. It's misleading to say otherwise.
This is still quite a good paradigm to follow - the semantics for waiting on, duplicating, closing and sending file descriptors to other processes are generally well-defined and well understood. For example, Linux exposes resources like timers, signals and namespaces through file descriptors.
Before that, you had completely different syscalls depending on the device you were using.
First, the missing background:
djb likes unix.
the unix philosophy is to compose small programs together to solve problems.
djb's own programs illustrate this really well. They are all small, focused tools. This allows each program to focus on their particular task or domain.
The primary method of composition in unix is the pipe in a shell. Each pipe has two descriptors. One for read and one for write.
It is very easy to create a pipe and handle pipe IO.
The article:
At some point, djb wanted to have some programs live on the network. This expands the composition beyond a single machine. If you just try to treat a socket as a standard pipe, you encounter the problem he describes.
Any program utilizing a pipe requires two file descriptors. If someone built a trivial 'netpipe', they could just 'dup' or 'dup2' the socket file descriptor to make it look like a normal pipe. The problem with that is now the socket won't close until both fds are closed. This means the remote end won't detect EOF. This means the 'netpipe' program has to be very clever in order to detect EOF and do a proper close on both so the remote can see the last bytes and then EOF.
[1] http://web.archive.org/web/20030805143958/http://cr.yp.to/tc...
Edit: HTTP Last-Modified header looks like it might be right:
Last-Modified: Tue, 10 Jun 2003 23:44:11 GMTHeck, we could even create a centralized system for managing our new namespace. OMG, maybe we could charge people money for the names? Yes! We're rich!
And the result: Hundreds of millions of "parked" domains serving up cheap advertising. Simply brilliant!
1. See Hobbit's comments in netcat source code for a differing opinion.
I often wish that people like djb or Hobbit (=low tolerance for nonsense) had designed the systems that we are now stuck with.
Though they are only applications, netcat and ucpsi have aged well and remain a pleasure to use.
At the very least, I'd say there is strong coercion. If not in favor of using names, then certainly in favor of deprecating the use of IP and port numbers (e.g., for email). Lemme guess, now you'll say "No one is forcing you to use email."
s/ucpsi/djb'"'"'s & applications/The more fragmented an IP packet gets along the way the less likely it is it will reach destination, so you have to take into account path MTU size and split your datagrams accordingly. You also want to send as many datagrams as you have available in as few IP packets as possible and you want to do slow start for the same reasons TCP does it.
Result: your datagrams need to become a stream of bytes to be handled efficiently by any transport protocol sitting on top of IP.
EDIT: theoh says it better https://news.ycombinator.com/item?id=6080324
http://www.joelonsoftware.com/articles/LeakyAbstractions.htm...
this was written about by Joel in 2002.
This article probably dates before that, but is last modified circa 2003 (see other comment on that)
djb's argument here is that TCP sockets are more like pipes, with separate read and write buffers, and separate read-side and write-side close operations. This makes sense, but what about UDP sockets? What about operations that apply to the socket as a whole, like bind(2), listen(2) or ioctl(2)?