Rob Pike comments on Ryan Dahl's rant
plus.google.com
plus.google.com
Rob Pike - It's a different kind of mess, and for different reasons, but the Unix/POSIX/Linux systems of today are messier, clumsier, and more complex than the systems the original Unix was designed to replace.
It started to go wrong when the BSD signal stuff went in (I complained at the time), then symlinks, sockets, X11 windowing, and so on, none of which were added with proper appreciation of the Unix model and its simplifications.
So let the whiners whine: you're right, and they don't know what they're missing. Unfortunately, I do, and I miss it terribly.
* Use a dispatch queue and dispatch_async a block to handle computation. Since the queues work from a shared thread pool and are managed at the kernel level, a long-running Fib calculation won't hold up other blocks from processing on any concurrent queue (it would still hold up a sequential queue...obviously).
* Use dispatch_sources to do your IO. Now you don't have to worry about clumsy callbacks, and your IO operations are perfectly happy co-existing with your long-running computations.
* Use dispatch_read/dispatch_write to persist to disk and you don't have to worry about different latency and throughput characteristics of files vs sockets
...and to the point about signals:
* Use a signal dispatch_source to do things in response to signals that you would never have thought possible. This is possible because libdispatch sets up a trampoline to catch the actual signal and then turn around and queue your dispatch_source block. It's not a complete panacea, but it's close...
...now if only Linux would implement kqueue
Note also that signalfd() can do something similar to your dispatch_source.
Internally, libdispatch is built on a system call named "kevent" (part of FreeBSD, first introduced in 4.1, meaning it was introduced in 2000). This system call will block while waiting for an event.
When a system call blocks, it is subject to a signal interrupting it, in which case you need to check for the EINTR error condition to restart the system call manually, unless you have SA_RESTART set on your signal handler /and/ your system call is one of a small set of "standard I/O" calls (note: kevent is not on this list[1]).
[1] http://www.mail-archive.com/freebsd-net@freebsd.org/msg01788...
Unfortunately, looking at libdispatch's source code, one will note that none of the usages of kevent() are correctly handling EINTR. I noticed this, for the record, while debugging the "dispatch_assume_zero(k_err)" that I had determined was being hit while running some (buggy) code of mine inside of libdispatch (it checks for EBADF, and that's it).
tl;dr signals are insidious and libdispatch didn't care ;P
Signals are just a software interrupt mechanism, doubling as a way to kill runaway processes. All the things you are complaining about are problems with interrupts in general. Yet all computers since the 1960s still have interrupts, because they allow you to do things you can't do without them. And that's still true when you're talking about signals. Signals allow you, for example, to implement preemptive multithreading (with SIGALARM) or transparent transactional persistence (with SIGSEGV) or buffer overrun checking (ElectricFence, with SIGSEGV) in a userspace library.
However, they've clearly been expanded far beyond what they are necessary or even good for, and lamentably are still not usable for the primary use of hardware interrupts: I/O.
I must admit it is not trivial to replace this system with another equivalent without changing a lot of how Unix works. For instance signals are used in order to interrupt a program: if you don't make this trappable (like kill -9) then programs can't recover or prepare in any way before quitting. An alternative can be to just allow a signal handler to fire that will never return.
From this point of view you could even just have two signals, with the only goal of stopping processes, one in an hard way (SIGKILL) and one in a soft way (SIGINT), firing a signal handler (but the process will terminate when the handler returns). This way you don't need handling the interruption of syscalls nor writing "safe" signal handlers. As you said signals are interrupts. If the interrupt handler can never return a lot of problems disappeared.
All the other tasks currently performed with signals should use a more general and saner message bus, with select(2)able file-alike interface.
This is a more flexible system than Unix integer signals, but in practice the extra flexibility isn't often used, as a process can simply mount a ctl file into the namespace or post a pipe in /srv.
Reading your links now.
Edit: Yeah, that's pretty much what I was envisioning. I need to play with plan9 more:-\
Incidentally, this is node.js' model.
Same problem applies to one of Plan 9's key design features: private namespaces.
I'll quote jwz: "Convenient though it would be if it were true, Mozilla is not big because it's full of useless crap. Mozilla is big because your needs are big. Your needs are big because the Internet is big. There are lots of small, lean web browsers out there that, incidentally, do almost nothing useful. If that's what you need, you've got options... "
"Plan 9 from User Space (aka plan9port) is a port of many Plan 9 programs from their native Plan 9 environment to Unix-like operating systems."
and 9base:
http://tools.suckless.org/9base
"9base is a port of various original Plan 9 tools for Unix, based on plan9port."
Unix said that, but in fact meant "everything has a unique descriptor associated (an integer)".
Directory is not a file, it's a set of files and other directories. Unix doesn't even tries to camouflage this simple observation and returns EISDIR.
Adding sockets and windows outside the filesystem is one of the things Rob is specifically complaining about here, and in fact complained about at the time as well.
oh yes that is so elegant...
Same argument goes for find. Its usefulness compared to du justifies its existence.
While the design of these additional features violates the "UNIX way," it doesn't violate pragmatism. Too often in our field, perfect is the enemy of good enough. Is BSD's model and implementation of sockets perfect? Surely not. Is it good enough? From the purists perspective, maybe not. From the pragmatists perspective, absolutely. I probably wouldn't be typing this today (on my macbook pro) without the implementation of BSD sockets.
Today it's usual to structure your main() as some kind of select()-like loop; even if you don't need one for your main workflow, at the worst you can spawn an I/O thread and put it there. But back then, you didn't have threads. Lots of programs didn't read your input but still wanted to catch signals, e.g. to exit cleanly on kill -15. Many others read input, but without an event loop - they would just try to readchar() periodically when they had nothing else to do.
Was there an obviously better design back then?
This is now starting to look considerably inelegant, and we haven't even talked about implementing the equivalent of synchronous signals provoked by a program's own action (SIGILL, SIGFPE, SIGSEGV, SIGBUS...).
I can't help but juxtapose this current dialog (which now includes one of the Unix forefathers) with the idea of "Worse is Better" (http://www.jwz.org/doc/worse-is-better.html). Maybe it's the Jersey in me but at the end of the day working, shipped software is all that concerns me.
It's a Babylonian tower made of mud (some would say camel poo). But it's brightly painted and has a good view, and all your friends live in it.
Personally, I don't wanna move out either, but I still have to turn my head every time you see the paint come off somewhere…
But, taking me typing into this text-box on HN as an example: What would the Unix way be? (If every tool does one thing and does it well, with text as input and output, and pipes to join it all together.)
Would I really have an unholy long command-line of a bunch of tools piped together (but accessed by clicking an icon)?
It's basically the model for CGI scripts.
Let's cut to the chase: Why would you have a textbox and a specific site for what amounts to a simple discussion? There's really no big conceptual reason why we couldn't do this via email, nntp or some similar protocol.
Leaving that aside, the more difficult question is how you'd get the textbox on your screen, i.e. what's the "True Unix" way of "web browsing"? There aren't that many examples of a rather graphical, highly interactive programs that believers would classify as really Unix-like. Maybe something roff-like, where you have a pretty universal display language and different (server-side?) tools are used to create something that would be too hard to express in the language as is (cf. tbl, pic), but then how would you make that interactive without doing the same stuff as HTML/CSS/JS? I was quite fond of the concept of NeWs, where you'd distribute your application over the net as PostScript, and I think Pike's Newsqueak went in a similar direction.
And how would you handle things on the server side? I don't think it would be that much different from what we're doing now. HTTP is quite resource-oriented and thus maps closely to a file system (the path-like nature of URLs is no accident). A CGI/PHP model for simple "files" would suffice, and you could have an almost arbitrarily complex application that appears as a file system, just like you can have that now appearing as a bunch of HTTP resources. People wrote "big" C applications, even though they could've theoretically done it all with a bunch of shell scripts, awk and ed. The "one thing and one thing only" mantra never was that religiously adhered to.
So, in conclusion, I don't think we're that far off right now, especially on the server side. If you look at what Pike's complaining about, it's mostly how you do it. Quite often the wheel is needlessly reinvented or has too many spokes, it's not that driving somewhere is wrong.
I'd argue that our current system is closer to "Unix" than it would be too Lisp Machines, Smalltalk and other more homogenous systems.
By the way, Plan 9 doesn't have sockets, it has /net: http://man.cat-v.org/plan_9/3/ip with the dial API built on top: http://man.cat-v.org/plan_9/2/dial
This means that for example, Plan 9 applications that do networking are not tied to a specific network stack, you can mount multiple network stacks concurrently, you can have 'virtual' network stacks (for example running in user space and proxying to a remote host, or doing other neat tricks), and the apps don't need to care.
Also when Ipv6 was added to the Plan 9 stack, no application code had to be modified, because the API nicely abstracts network addresses.
"If you want to see a system that was more thoroughly _designed_, you should probably point not to Dennis and Ken, but to systems like L4 and Plan-9, and people like Jochen Liedtk and Rob Pike. And notice how they aren't all that popular or well known? "Design" is like a religion - too much of it makes you inflexibly and unpopular."
.. oh wait.
It is rather that good design make you popular and not so good design doesn't.
Property list files: a big f*n win compared to the ad hoc mess in a Linux/BSD /etc directory.
Hacks around extended attributes: any reason not to like those? Or just because in 1977 a file was just a file, and that's the way it should be, god damn it?
XML-based init system: a sane init system. And XML added in for standardization.
OS X is a mess in several ways, but those are not it. And the "monolithic BSD kernel bolted on top of a Mach 'micro'-kernel" sounds like a win-win situation. Monstrous why? Because it doesn't fit some idealistic model?
One of the great things about Plan 9 is that everything implements the same interface. If you can interact with one thing, you can interact with everything. Using new parts of the system becomes obvious because you already know the interface.
A concrete example of why this works: In Plan 9, process information is available via (guess what) the file system, so the Plan 9 debugger just reads the state of running processes from those files. Because file systems are automatically exported over the network (also part of the file system), you can debug a running process on another machine without the debugger knowing anything about the network. Nobody had to implement network debugging - it just worked right away because of good design.
You know you have a good design when these kinds of complex behaviors just "fall out" without any additional work.
Last year I even discovered a kernel bug in their file descriptor passing implementation that, AFAIK, still isn't fixed. Various signal handling properties are more buggy than on Linux.
IIRC, a lot of their code came from FreeBSD back in the day. I wonder why they didn't keep it in sync?
It is just that the end users are different people to the ones on Mac (and would indeed usually be power users or developers there).
What's the alternative? OS/2 had nice features; it got crushed. BeOS had some nice features; it got crushed. (Yes, I'm aware that they're probably still around in some OSS version.) Plan9 keeps being mentioned, but do regular users understand it? (Is there a "this is why we do things, and this is why that's good" written for people who haven't written their own compiler?)
Are people working on experimental new OSs to "fix problems" in existing systems, or have we gone way past the point of no return? Is there any work on new microprocessor architecture? How much stuff in my modern OS is there because of legacy 8086 stuff?
I think, but I'm not sure, that the length of time it's taken people to get (for one example) IPv6 rolled out shows that change is not likely.
Plan9 LiveCD and installer: (http://cm.bell-labs.com/plan9/download.html)
Haiku: (http://haiku-os.org/)
Yes, you can start with the main Plan 9 paper: http://doc.cat-v.org/plan_9/4th_edition/papers/9
And follow up with The Organization of Networks in Plan 9: http://doc.cat-v.org/plan_9/4th_edition/papers/net/
And The Use of Name Spaces in Plan 9: http://doc.cat-v.org/plan_9/4th_edition/papers/names
This will give you a good overview of the basic design decisions in the system and their rationale, for further details on how Plan 9 deal with issues from toolchain design to authentication and security see the rest of the papers: http://doc.cat-v.org/plan_9/4th_edition/papers/
They are a wonderful read even if you never touch Plan 9, they are full of insights, ideas and criticisms of existing approaches, and many even include discussion on how to apply them to existing nix systems (sadly most of this has gone almost completely ignored by the nix community).
>> Plan9 keeps being mentioned, but do regular users
understand it? (Is there a "this is why we do things,
and this is why that's good" written for people who
haven't written their own compiler?)
> Yes, you can start with the main Plan 9 paper:
http://doc.cat-v.org/plan_9/4th_edition/papers/9
So, he asks about regular, ipad-toting, angry-birds-playing, how-do-I-turn-on-my-printer users and you link to a technical paper containing buzzwords such as "compilers", "internet gateways", "distributed systems", "POSIX", and "remote procedure calls". Do you see the problem with this?Do your regular iPad users understand Unix, iOS, BeOS, or OS/2? No, they don't. They understand point and click interfaces on top of them. The operating system is irrelevant to people who want to Play Angry birds. Users who can't turn on their printer are irrelevant to a discussion on the merits of one operating system over another. The user interface on top the system and its programs is a separate topic.
But as an "iPad user", I don't know any of that. If the feature isn't included in the OS, "turtles all the way down" and enabled by default, then I will never even know that such a thing is possible (and even if I did, wouldn't know how to get it).
So it's really exactly for the non-expert users that OS development is so important. Everyone else (and by that I mean the minority), is informed enough to figure these things out no matter what OS you give them.
(a) Ryan Dahl is qualified to speak about these issues (i.e. he worked on the original Unix design or something)
(b) Rob Pike's comments are a big deal to Ryan, and
(c) whether Ryan is right or wrong?
I'm one of the "new guys" that doesn't know what he's missing, and I'd love to put this into context.
(b) Rob Pike's comments are a "big deal" because Rob Pike is one of the "original neckbeards" that worked on Unix, Plan 9, and has recently been developing the Go language at Google. i.e. He helped developed the OS that Ryan is complaining about.
(c) "right or wrong" depends heavily on what you are trying to accomplish. Much like anything in the real world, there is no right or wrong answer.
Do you know that 'neckbeard' is a slur? It refers to someone who tries to grow a beard but can only grow hair on below their jawline, which is typical of adolescents.
FWIW, Rob has never had a beard. Ken, dmr, and bwk all have/had full beards, not neckbeards.
I am kind of surprised to find myself writing this post, but here I am.
Ok, maybe Chuck Moore complaining about Oberon after that. Can we go any deeper?
(Also: Discussing this on the web is a wee bit ironic.)
It was fun back then.
I think I'm going to grab my time machine and warp back.
And he has earned the right to do so.
Every command that operates on files needs an extra flag to tell it whatever it should follow symlinks, or operate on the target, or on the symlink itself, etc.
Note that the history shows up in the ln command which originally only created hard links, and then got the -s parameter when symlinks were introduced.
Symlinks may be ugly and/or dangerous relative to the original file system concept, but they solve problems. They are not "all good", but life would be worse without them. E.g. - application by application implementation of location aliases a la IIS virual directories (or whatever they called it) for web applications.
ADDED: Note that I avoided hard links after that, because I was worried about what would happen if I modified one file and deleted it, thinking it was a link to another, when it was actually a copy. A person could lose a lot of work if that happened. At least with a symlink, I know that something is or is not a link.
>Symbolic links make the Unix file system non-hierarchical, resulting in multiple valid path names for a given file. This ambiguity is a source of confusion, especially since some shells work overtime to present a consistent view from programs such as pwd, while other programs and the kernel itself do nothing about the problem.
etc etc.
filesystems are a mess, and it'd be great if something could fix them.
What happens if you move a symlink? Will it still point to the same object? That depends on whether it's relative or absolute. What if you move the directory containing it? Do you have to recursively check for the correctness of links every time you operate on a directory? Sounds unreasonable to me... What if it links to a location which isn't mounted? Whats the definition of '..'? If you cd into a linked directory and do an 'ls ..', should that list the parent of the target or the link? How would you implement that?
The point is, symlinks are a messy kluge which probably wasn't thought out very well. You can look at what Rob Pike has to say about it himself on http://cm.bell-labs.com/sys/doc/lexnames.html
You have file at /some/path, then you put newer version of that file to /some/path, replacing the old file, and the symlink still points at the same address, and it doesn't matter what actually happened with the file itself.
It's useful enough to warrant its own abstraction.
What you mean to say here is that it was a DAG and now it's a directed graph.
(that is, not only "not having anything like them").
Pity.
I don't care.
Learn to use what you have or get something else and stop fucking complaining. The End.