But for the contemporary software coders the ingenious idea is of course rubbish. They would replace everything by node_modules.
The problem is the lack of innovation holds us back from solving complex use cases. The sockets API in particular is a gigantic mess, because the protocols are more complicated than people imagined forty years ago. For example it is simple to send a datagram. What if you want to get notified of the timestamp when the datagram hits the wire? Now you will 1) read a lot of man pages, 2) write 100s of lines of gross C code, 3) read the source code of the Linux kernel to figure out which parts of the manual were lies, then 4) fix your program with some different code and a long, angry comment. None of they would be necessary if the last quarter-century of network features hadn't been jammed into the ioctl and setsockopt hack interfaces.
In other words there's plenty of room for operating systems to reconsider interfaces in light of the systems and protocols we have now, instead of Space Age imaginary use cases.
Innovation is awesome, when it is in the pursuit of achieving something better. (Better, not just newer for the sake of rewriting everything Yet Again.)
What burns me out about this industry is that a huge amount of "innovation" is just a pursuit to create complexity for the sake of complexity and resume-driven development.
Thus we get to classics like the "Command-line Tools can be 235x Faster than your Hadoop Cluster" (a well-known one, but countless similar scenarios play out everywhere continuously).
To all of you setting up your 20 node cassandra cluster behind your enormous kubernetes setup, all just to serve your 100rps app: you're not innovating, you're doing it wrong.
(Not "you" you. The whole software engineering industry.)
Yes there are, but those 100k's of people doing that are working in a small handful of megacorporations.
For 99+% of the companies out there, they don't need the complexity. But it gets shoehorned in because resume.
Edit: I won't elevate the Unix model to a pedestal, but in his favour I will say that the relative simplicity of the model and it's inherent difficulties to extensibility helped limit scope creep and overengineering.
my daily job :(
Most recent horror I discovered: `mincore` is intentionally broken, always returns positive results when the calling task is running in a mount namespace. Because the hallowed unix file abstraction is fundamentally bad.
Because they're not just files. There are additional semantics and concerns that you cannot abstract away with read/write calls. You also need ioctls, you need send/recv, you need parsers on top to do something useful with the data, and so on.
The reason that the simple model dies out is because it's not as useful as it seems in practice.
It's a good framework for text-based computing, but nowadays we have so much more problems to cope with. This original idea never addressed networking, binary versioning and control, configuration management, async IO, multithreading, GUI, etc.
>Every increase in expressiveness brings an increased burden on all who care to understand the message.
There are plenty of new things that are largely compatible with unix philosophy. MQTT is one of my favorite examples, it's a message bus that follows a lot of unix philosophy and I find it a joy to work with compared to stuff like DBUS. Obviously it doesn't fill the same role as DBUS, but still.
http://doc.cat-v.org/bell_labs/utah2000/
At this point it's network effect, familiarity, and the free price tag, not a quasi-religious sacrament. Anyone can try to take that different approach, but that doesn't mean it will go anywhere. Some very smart people have worked on alternatives to Unix, or different directions for operating systems -- including the original Unix team with Plan 9, Wirth with Oberon, the Lisp machines. No Unix Inquisition shut those down, they just didn't offer a 10x improvement along enough axes.
I'm keen for any potential successor, but a lot things come back to effectively text, and anything we could build on top of it. Binary object formats seem like an alternative for faster parsing of structured data, so while that ability has always been it, it's a matter of people actually sticking to it. Maybe some coordination between the program and the shell?
Another approach is that the OS is basically a complete environment where everything is code. You can see the idea in the Smalltalk environments where you can theoretically interact with every object by sending messages. Lisp machines come to mind as well and one could even consider early personal computers that booted into BASIC as an idea of this (though in BASIC instead ov everything is a file everything is memory)
Anyone know how to add a new stdout-like interface to unix-like OSes?
https://man.freebsd.org/cgi/man.cgi?query=libxo&sektion=3&ap....
See ps(8) for example use: https://man.freebsd.org/cgi/man.cgi?query=ps&apropos=0&sekti...
Personally I like YAML, since you can add type-hints to data which you can use to turn simple text types into more complex types, but I can see why that standard wouldn't take off.
Most of the unix-philosophy people I know are interested in stuff like that, it just has to be implemented in a thoughtful way.
But that's sort of part of the problem, no one is going to agree on a common universal data-type.
You mean getting programs to ingest arbitrary data structures? OK, JSON is not arbitrary - it is limited to a certain overall format. But it can be used to serialize arbitrarily complex data structures.
JSON doesn't care so much about lines and doesn't necessarily represent an array of records, so it doesn't fit into this box.
Imagine a magical `ls` that can emit text/plain, application/json, what have you:
ls -f text
ls -f json
ls -f csv
ls -f msgpack
...
Now instead of specifying formats on both ends: ls | jq # jq accepts application/json => ls outputs as json
ls | fq # fq accepts a ton of stuff => "best fit" would be picked
ls | fq -d msgpack # fq accepts only msgpack here => msgpack
ls # stdout is tty, on the other end is the terminal who says what it accepts => human readable output
Essentially upon opening a pipe a program would be able to say what they can accept on the read end and what they can produce on the write end. If they can agree on a common in-memory binary format they can literally throw structs over the fence - even across languages, FFI style - no serialisation required, possibly zero-copy.We know how to do that:
- https://www.rfc-editor.org/rfc/rfc2616#page-71
- https://www.rfc-editor.org/rfc/rfc2616#page-100
And I mean, we really know: the last one we already do! Tons of programs check for stdin and/or stdout being tty via an ioctl and change their processing based on that.
It'd allow a bunch of interesting stuff, like:
- `cat` would write application/octet-stream and the terminal would be aware that raw binary is being cat'd to its tty and thus ignore escape codes, while a program aiming to control the tty would declare writing application/tty or something.
- progressive enhancement: negotiation-unaware programs (our status quo) would default to text/plain (which isn't any more broken that the current thing) or some application/unix-pipe or something.
- when both ends fail to negotiate it would SIGPIPE and yell at you. same for encoding: no more oops utf8 one end, latin1 on the other.
For me, this would mean an OS that supports both static and runtime reflection, and therefore code-generation and type generation. Strong typing should become a fundamental part of OS and system design. This probably means having an OS-wide thin runtime (something like .NET?) that all programs are written against.
UNIX's composability was a good idea, but is an outdated implementation, which is built on string-typing everything, including command-line options, environment variables, program input and output, and even memory and data itself ('a file is a bag of bytes').
The same thing has happened to C (and therefore C++) itself. The way to use a library in C (and C++) is to literally textually include function and type declarations, and then compile disparate translation units together.
If we want runtime library loading, we have to faff with dlopen/dlsym/LoadLibrary/GetProcAddress, and we have to know the names of symbols beforehand. A better implementation would use full-fledged reflection to import new types, methods, and free functions into the namespace of the currently-executing program.
James Mickens in The Night Watch[1] says 'you can’t just place a LISP book on top of an x86 chip and hope that the hardware learns about lambda calculus by osmosis'. Fair point. But the next-best thing would have been to define useful types around pointers as close to the hardware as possible, rather than just saying 'here, this is the address for the memory you want from HeapAlloc; go do what you want with it and I will only complain at runtime when you break it'. Pointers are a badly leaky abstraction that don't really map to the real thing anyway, given MMUs, memory mapping, etc.
There are so many ways we could make things a bit more disciplined in system and OS design, and drastically improve life for everyone using it.
Of the major vendors, I'd say only Windows has taken merely half-hearted steps towards an OS-wide 'type system'. I say half-hearted, because PowerShell is fantastic and handles are significantly more user-friendly than UNIX's file descriptors, thread pointers, and pipes. But this is still a small oasis in a desert of native programs that have to mess with string-typing.
[1]: https://www.usenix.org/system/files/1311_05-08_mickens.pdf
Long time ago I asked my karate sensei how come I see some of the black belts doing katas in a different way. His answer was that to improve on something one must first have deep mastery of how it is done.
(i.e. "Shut up and do the katas correctly. If you ever get to black belt then you can improvise.")
Always found that answer to apply to everything. In particular it applies to a lot of the poor quality libraries and code we see today as a direct consequence of not understanding the past and what has been tried and failed and why.
The fact that we're talking in 2024 about using the file abstractions of Unix created in the 70s is proof of how extremely powerful they are. Nothing is perfect, but that's quite an achievement.
To write great code, I believe, means you have to be a domain expert of the domain for which your code solves problems.
https://www.theregister.com/2023/12/25/the_war_of_the_workst...
Note that this is loosely based on part of a FOSDEM talk I gave 4 years ago.
If you impose any structure, you're telling hackers what to do at the system level.
If that clashes with their favorite language run-time or whatever, you're effectively declaring war.