Kernighan and Pike were right: Do one thing, and do it well
medium.com
medium.com
But ... firstly, let's stop with the Unix worship. They didn't deliver, even in the domain of the terminal. My copy of `ls` has at least 40 flags (where I stopped counting). Secondly, people have tried composition on the desktop. That's what CORBA, COM, and OLE were all about. They all kinda sucked for various reasons. Finally, the place at which the author arises has been around for as long as Unix: it's Emacs.
It's the same for many programs that started small but grew loads of knobs with age. You can't claim that this was the Unix vision, because what we call Unix today isn't even technically Unix.
I'm not saying Linux is bad or anything; I'm just pointing out that it's much more organic than Unix ever was, and that the resulting system is not necessarily carefully orchestrated.
"Usable in an ad hoc situation" gets us back to composability. If you can compose existing parts to meet a novel situation, that's better than having to carefully architect a solution. And the key to being able to do that is common interfaces. Unix did exactly that. And, for all the criticism, it works pretty well.
Can a well-designed interface do better? Sure. Can a well-designed interface deal with something outside the design parameters as well? Probably not.
And yet everyone still uses `ls`.
Why would anyone care about that difference? If I'm reading the listing, I don't care. My eyeballs don't care if the photons came from a GUI or text in a shell. But if I'm trying to pipe the results to another program, then I care. GUIs don't pipe well.
And people like me pipe the results of ls all the time.
Anything that you'd want to colour, the data should already be there.
Further I'm not sure the Unix way inherently requires completely unstructured text
How those bytes are interpreted is up to the endpoints of the streams. Sure, a lot of endpoints interpret byte streams as text. That's not Unix, that's a particular endpoint.
You will notice that Unix does not even offer 'text' files (like Windows does). Or ISAM files, or random access files. or Fortran carriage control. Just binary files. Because the Unix way is streams of bytes, and nothing more.
CSV, line delimited, is that unstructured?
Further if you wanted to go with JSON or something you can, there's nothing preventing it. Pipes are based on unstructured text so you have that flexibility. If you want to add JSON ls and filter that through JSON sed to JSON less you can.
Finally 'the Unix way' is a cultural thing. It isn't set in stone. It can, and probably has evolved. So while you could argue that it meant that to some people at some time, that doesn't make it true today or in the future. The down side is that we have these disagreements over what 'the Unix way' is. But any philosophy that's been around for 40 years, in tech no less, has to be flexible, that's why it's stay relevant.
I wish! Back in the day, tools like iostat and vmstat printed tabular data using fixed width fields with no spaces between. This worked fine until the computers got bigger and faster, then the larger numbers filled the fields so they ran together making the output unreadable. A surprising number of vendors had this problem.
By BSD 3 in 1980 it had 11 options. https://github.com/dspinellis/unix-history-repo/blob/BSD-3-S...
The thing is, we can see even from the 1970s 'ls' how the Unix model doesn't meet the goal "to chain these simple programs together to create complex behaviors".
There is no option to escape or NUL terminate a filename, making it possible to construct a filename containing a newline which makes the output look like two file entries.
The option for that was added later.
There's also the issue that embedded terminal codes will be interpreted by the terminal.
There are many flags that change the behavior of the file system calls that would be made...such as getting additional information.
To support all the formatting commands, the piped output of `ls` would need to be very verbose and structured too...like json or something.
They had to draw boundaries of functionality at some point.
So the unix idea of "one thing" was probably defined a bit more vaguely or more end-user use case related. Convenience seemed to be king.
It would have be nicer if they could have added a shell-like pipe syntax to C (and designed the code in a more modular and stream-like way).
In the `ls` [source code][1] you can see simple blocks of steps like:
if (!file_ignored(file)) {
...
sort_files ();
print_current_files ();
clear_files ();
}
For example the `file_ignored` function in `ls` would have made a nice reusable library (with standardized configuration params) and a cli tool (with standardized flags).Missed opportunity to unify the shell and C, but I guess SmallTalk was kicking around at that time which went a whole lot further.
It's funny that we are still very much lacking this unification...
We don't have a high-level interpreted language that can also perform well as a system's language.
New systems languages are popping up all the time, but they are more and more hardcore (looking at you Rust). They don't embrace any aspect of scriptability.
The best effort I can think of is https://www.modular.com/.
Surely though, a native TypeScript-based language is the way. It by far has the nicest syntax out of the popular scripting languages today.
We have come so close...like with Dart...but still everyone seems to be avoiding the inevitability.
[1]: https://github.com/wertarbyte/coreutils/blob/master/src/ls.c
Agree with UNIX worship point, I also had it, I was young and naive, lacking experience in the history of computer systems.
After WinDev managed to botch Longhorn efforts, its .NET ideas were redone in COM for Vista, and it has been like that ever since.
WinRT/UAP/UWP, is basically COM with a new base interface IInspectable, .NET metadata instead of TLB files, and app boxing.
And it is the foundation of WinAppSDK, whose goal is to port UWP subsystem into plain Win32, although their current execution leaves a lot to be desired.
As pjmlp mentioned, this part was very successful and all new Windows APIs are based on it.
The COM which failed and which you are thinking about was related to OLE and the idea of embedding Excel sheets inside Photoshop drawings, ...
that's still a thing isn't it?
On Windows it works both ways with OLE 2.
But composition needs to be understandable and controllable to be useful.
For it to be understandable, it needs to match the user’s anticipated mental model. Now, this isn’t an appeal to an intuitive notion uninformed by user education. But it does need to make logical sense to connect two things together temporarily to solve one task. And that logic needs to come from the user and not the developer. In other words, the developer needs to provide concepts that conform to a user’s algebra. For a photo editor, that’s hard to provide to the average user. For some CLI tools like wc and cat, it is easy.
For composition to be controllable, the primary method of configuration needs to be in the contextual relationships among the tools and not in the tools themselves. That is, using cat and wc to count the number of words in a text file is primarily enabled by the filesystem and shell features, none of which require user configuration during the operation. Configuring wc to count instead the number of lines doesn’t distract too much, but imagine counting the number of lines emitted from a streaming serial device: you want it unbuffered, but only kinda because of other issues that arise, like multiline error conditions. In that case, the value is disrupted by edge cases and you have to configure the nouns (the serial stream at least) by looking it up on Stack Overflow. It’s not just easier to consider building a monolithic tool, but arguably the right thing to do because you need to configure everything so much. Plus, this is too much to expect the user to handle through a simple composition.
In my opinion, desktop composition doesn’t work because the tasks are often beyond the limits of composition. And without a culture of composition among users, it’s harder and harder for remaining desktop applications to do it. Especially as younger users are unfamiliar with concepts like the filesystem, system keyboard shortcuts, “saving” files, offline operation (i.e. you have to reason through a problem and cannot lookup the answer), etc. Better to go with the monolithic app model, which can still use plugins under the hood or multiple processes for developer productivity. And I’m sad about that.
GNU is Not Unix.
I count 13 flags on both Unix V10 [1] and on Plan 9 [2], its spiritual successor.
-1 list one file per line
-C list entries by columns
-t sort by time, newest first; see --time
-r, --reverse reverse order while sorting
-S sort by file size, largest first
--sort=WORD sort by WORD instead of name: none (-U), size (-
S), time (-t), version (-v), extension (-X), widthA consequence of the proliferation of arguments in the GNU tools is imho that it made the GNU manpages too verbose and information dense to the detriment of their usefulness.
Back then, Companies and Colleges were the ones who got UNIX. Companies went with AT&T, and IIRC Colleges when with BSD. So AT&T won out due to better financial backing and maybe a possible threat of AT&T going after BSD.
There are 10 options listed for "ls":
l, f, a, s, d, r, u, c, i, f, g
My recollection is that this card was for Unix V7.
CLI flags are typically arbitrary. What if they weren’t? How would they look like?
Take big monolith. Refactor into 1001 pieces. Glue together again.
If glue is inflexible, then result is monolith+glue.
If glue is thick, then result is mostly glue + little bits of monolith.
If glue is superb (well designed & flexible), the whole works the same as monolith, but pieces can easily be rearranged as needed (and individually tested!).
Designing good glue is HARD! And then some.
Good point.
You have to hand it to them though, that shell commands are still the least-typing way to do something compared to every programming language out there.
Like if you gave the user an always-on JavaScript or Python repl, shell commands are still going to win every time.
There are so many JS functions I have that I wish would just get automatic cli interfaces...instead of having to go through the ceremony, or use a repl, or something like that.
The article doesn’t really quite get there, but is close. Almost everything else in there is wrong and should be ignored. Let’s see… *nix command line never did follow DOT, that’s the definition of shitty, not enshitification, the graphs are pulled out of thin air and represent the author’s feeling, not data, microservices don’t solve problems in large projects that various other approaches don’t solve (microservices tends to combine an certain approach to architecture, deployment, organizing dev teams and assigning responsibility to them within a larger organization, organizing code, change management, etc… the relevant things here aren’t specific to microservices).
BTW, the microservices graph shown near the end is a lot like spaghetti architecture, where everything depends on everything else. (Not that you have to do it like that, but you’ll need some higher-level organization to manage it.)
Do you saturate the resources of one machine and need to split things off? Do you have multiple teams each taking care of their own stuff? Do multiple services, it's fine in those cases. For almost any other reason you are just adding complexity, boilerplate and additional failure conditions.
If you think that microservices solve the "spaghetti code problem", well, good luck to you, you'll need it.
seem very confusing to grug"
Who's hiring grugs, the lazy programmers that try to do the most efficient thing in the least amount of work? I'm not smart enough to set up a whole georeplicated k8s cluster for your blog. I'll use Apache, maybe put Varnish if you get on the frontpage of HN, ok?
(Work smart, not hard. DevOps today is the exact opposite of that)
Fundamentally, don't make life hard for yourself, or others.
The best engineer (and not only, see that famous quote by Kurt von Hammerstein-Equord) often is the lazy one. It was a known notion that somehow disappeared around the turn of the millennium.
We're still around. We're just so far up the hierarchy now that the wisdom can't be pushed down easily.
The wisdom only applies to the places where it applies, and there is no wisdom there on telling where those places are.
Well, except for the problem of too much grug.
That is a very good point. In fact, good monolithic code layout is a precursor for microservices, as only once you have identified your boundaries and isolated your concerns can you begin splitting them out into their own services.
I will say this however - microservices might not solve the "spaghetti code problem", but it definitely helps isolate it. When we get consultants in to speedboat a new system, we give them their own separate service. Saves a lot of time and de-risks our beautiful monolith.
If hacker news is correct, microservices is the #1 worst idea of software engineering.
And whoever promotes them is either a kid with no experience or a complete moron.
If you have a problem with bad consultants then solve that. Trying to solve it by jumping into microservices you will just have bad consultants writing your microservices equally badly, or even worse.
But you're probably not the person I should telling this, developers aren't usually in charge of that.
And on pricing specifically, both shitty and good consultancies charge about the same. The ones that charge well below market are an outlier that will guarantee bad product - those are to be avoided.
lijok wrote:
When we get consultants in to speedboat a new system, we give them their own separate service. Saves a lot of time and de-risks our beautiful monolith.
One intern was told to deploy the app and I guess he made a mistake and updated k8 configs which deleted all the pods. Tell me how code reviews are supposed to catch it.
If you think "why not restrict config actions" etc, it's because the dev can go there and tweak those actions too. And each step added means less velocity of development. It's a game of cat and mouse. With isolation, you can limit the blast radius and loss of revenue.
So you don’t have a staging environment to check config changes before deploying to prod? Or you don't put k8s manifests (or helm charts or whatever)in source control and just update prod through CLI? And you let interns do it alone?
Your problem isn't interns, it's a terrible deployment processs with weak controls. Even senior engineers are going to break prod regularly in conditions like that.
[...]
> One intern was told to deploy the app and I guess he made a mistake and updated k8 configs which deleted all the pods. Tell me how code reviews are supposed to catch it.
These sound contradictory to me. Part of having automated CI/CD, is not having to manually deal with k8s to do a regular deploy.
Our team spend most of the time designing fucked up messes that run over poorly designed APIs, slowly, that impact customers. If it's not that it's upgrading 100 services worth of dependencies constantly, debugging contract violations and weirdness or performance issues.
There are companies that are very successful and have split their service architecture into domain entities, called microservices.
Are they successful because or despite this decision?
Is both true in some sense?
A more natural way to split up a server architecture is to use computational boundaries. These are found by thinking of how data is processed and flows through the system as opposed to separation of high level domain concerns.
But this requires a computational design and not a domain/feature centric one.
The biggest problem with the whole IT industry is seeing outliers promoted as the successful path and people taking on faith arguments for technical decisions instead of rational decision making processes.
If your organisation doesn't look like that, you can only fail at microservices. Either the company restructures around the code, or the code looks like the company. Trying to make pasta from mash potatoes is doomed to fail.
Plus it's easier to run a hundred monoliths than it is to run 100 services. This is on the basis that running 100 instances of something the same with no inherent complexity in each instance is much easier to automate, build and manage than 100 things that are different and have to talk to each other.
Of all the reasons microservices are bad this isn't it. Your first point can be handled mostly automatically with half decent SRE tooling. Lock files on any modern language (including Python via poetry) fix dependency weirdness. Contract violations are a problem with your developers not the style of programming. As for performance I've been on both sides and being able to independently scale microservices via kubernetes is a god send. With proper tooling microservices work great.
The real problem with microservices is what you mentioned before that IMO. Complexity for complexity's sake. Some things are better as monoliths and some things are better as microservices. It's the same problem you see with normalization vs denormalization. Microservices are unambiguously faster in almost all cases. "Do one thing" means that one thing can be optimized to death, in isolation, without fear of interrupting other services. It also means team structures are easier to manage. Microservices don't mean pull in the entire universe of RESTful libraries and use exactly one feature for 1 or 2 endpoints.
But this introduces another problem:
A second managerial problem is that microservices are often used in scrappy startups with < 200 engineers. On any complicated platform this means one engineer might need to know how 10 services work. That's not feasible. You need to split it so a team owns a couple services and "does one thing right" (those services) and nothing else. It's the overapplication of microservices to every problem that is the source of most woes.
Startups will religiously apply microservices. The thought, of course, being that they can hire/fire faster and an outsourced engineer can probably pick up the necessary work quicker. Of course, this never happens, because instead of 10 microservices you have 300 and no one even knows where half of them are. 99% of the time it's not the pattern that is the problem. It's cargo culting. Just like everything else in the industry (looking at you electron, leetcode, functional programming, agile, etc).
It has it's purpose. It's almost universally over-applied owing to the cargo cult and holier-than-thou attitude of a lot of functional programmers. I see it happen in Python where comprehensions are often replaced with maps. While it is nice, the comprehension is more idiomatic and far more widely understood.
No, of course not. We divide a single machine in an uncountable number of virtual ones; write some code to make sure they don't talk to each other; write some code to make them able to talk to each other; write some code make more or fewer divisions on the run, automatically; and write some code so we can set them up automatically too every time we get a new split.
That's how you use software scalability and create a simple, predictable ops environment.
(Well, ok, I'm having a hard time solving the first one with more code, but did hear this claim applied to it more than once, so I'm keeping it general.)
There is a huge amount of people that sees nothing wrong with my rationale up there. I don't get it either.
I have gone through the migration of monoliths to microservices (yes, at scale with multiple teams and requirements to scale individual components etc.) and it solved a lot of problems.
The benefits outweigh the costs in my opinion (and drastically so).
When people say "you can write well separated components in a monolith..." I can only say: of course you could, but you are not going to (and certainly not everybody at your giant ass company is going to).
In my head MS works well for - Teams with different approaches. - Different tech stacks
I cannot think of another good reason in which MSs would be a better solution than a good monolith (and communication).
Would I want to build a world wide streaming service, I would probably need some custom microservices and I will not need some kind of plugin architecture for them. Would I want to build a slick editor that non-programmers can use for knowledge organization I will not wire something together with awk and gnuplot. If I want to quickly search some log files and sort and count the errors in them I will not look for an obsidian extension nor build a microservice. The question of how to architect such software is still an open one.
I was happy about the clear thesis statement in the byline of the article:
> Extensible programs like Obsidian have achieved a Holy Grail of software architecture, after decades of failed attempts
However, by focusing on categories like "Applets", "Unix programs" or "microservices" the article in effect did barely touch the topic of software architecture and offered no evidence to support that byline.
I have recently had the thought that this is mostly what software architecture is...at some level.
It's about where we draw the boundaries around code to make it easy for humans to understand.
Imagine taking an existing repo, and extracting every block of code into a function, and then moving them into their own separate file in one big folder. You could extend it to all the libraries of your code, and of those services we communicate with over the network too.
This would look a bit like a debugger symbol table. It's closer to how the computer understands the program.
The app still runs. It's just hard for people to understand what is going on.
Software architecture is simply the grouping of these functions to some extent.
It's not a foolproof analogy, because there are some decisions to be made about implementation details of things, but at some level of abstraction you would find the same operations need to be run.
Just an interesting thought.
> The question of how to architect such software is still an open one.
I find a problem we face with architecture discussions is that its always about tradeoffs that are not immediately apparent.
And it takes a lot of mental effort to remember why something is a bad idea.
The way we discuss it is limited by plain text. It's hard to demonstrate a system evolving over time in a concise manner.
Architecture should be evaluated by throwing a spec at it, and then changing every part of the spec (including adding perf requirements) and seeing how long it takes to make the changes, and how many bugs it has.
I was there, in the 1990's, as a working professional not a hobbyist. The popular architecture was, for better or worse, structured as layers, not spaghetti.
Some software from the 80s might meet that definition, but I think it's uncharitable (almost to the point of insulting) to assert that until the 2000s most programmers didn't know how to write anything but spaghetti code.
C is rough to refactor sometimes, but it's nowhere near the kind of "push a couple return addresses and jump into the middle of a function" that used to exist.
Composability lends itself to focusing a tool on a task but doesn't necessarily require it. If you build tools focused on composability i.e. output that is regular and easily parsed and properly using input and output channels, you'll get a system that is very extensible and responsive to end user needs.
If you instead focus just on the "single purpose" tools you'll often end up getting ones that are too limited. They'll end up tightly coupled with other programs/service since their functionality alone doesn't do anything all that useful.
The conclusion is: "There are no easy answers". I wish I was kidding.
The classic example is shell pipelines, but they are actually quite shit because they only deal with unstructured data (there's finally some work to fix that in Powershell and Nushell but it took many many years).
A better modern example of doing one thing and doing it well is apps that support plugins.
----
I think the whole "do one thing and do it well" is terrible advice. A "thing" is not well defined. It's equivalent to "don't have too many features" which of course leads to "how many is too many" and you're on your own.
It's one of those bits of advice like "premature optimisation" that is more often used to excuse thoughtless design than to motivate good design.
This has NOTHING to do with monoliths. If you cannot architect a monolith, you’re going to have a hell of a time with micro services. This just comes down to being an awful software engineer.
Another reason things were cleaner was because we could write components in the language that was best suited for it. In our monolith version, everything had to be Python because we essentially needed its data science ecosystem, but that meant every other component in the system was fighting against Python’s package management, performance, and reliability problems for no material gain.
I think you can get similar rails in a monolith world via strict anti-shortcut culture, but technical controls are a lot easier than political controls IMHO. Similarly, you could probably build some Frankenstein FFI regime in a monolith to support multiple languages, but that seems strictly worse from a maintainability perspective.
All that said, there are definitely costs to building with microservices—I’m not saying they’re a panacea or even better than monoliths in the general case, I just don’t buy the “if you can’t do it in monolith you won’t be able to do it with microservices” line.
> Large codebases will eventually reach an “enshittification point” — the point at which bugs are introduced faster than they can reasonably be fixed.
I think of enshitification primarily as an organisational/business phenomenon rather than a technical one.
It doesn’t necessarily emerge “bottom up” as the result of technical debt but from top down as a result of misalignment of values between the brutally commercial goals of the business and the values/goals if its customers. Tech debt and development velocity do play into this but I don’t think it’s the primary driver of it.
The term is a useful one at that level of abstraction imho because we already have quite a lot of language to describe things like this at the lower, more technical level.
The other two important criteria implied by this new term are that it is deliberate and that it is taken for the benefit of the “author” (company) at the expense of the user.
Bugs are inadvertent and have no intent so don’t match either criterion.
Google’s so-called “integrity” system is DRM for the benefit of advertisers (and thus Google), not users. Intel’s on-chip security manager isn’t for the user’s benefit and adds security risks.
Encrusting photoshop or Word with a million features or back-compatibility with a mistake made decades ago may make the program worse, but are in service of giving customers what they want. It may make the application worse in some ways but the motivation is pro-customer, regardless of the ultimate result. (Note: I chose these two programs bc I don’t even like either)
> Bugs are inadvertent and have no intent so don’t match either criterion.
Yes, good distinctions. Very different from inadvertent bugs or routine technical debt.
I do think it's odd to link to a definition of a word as the author does, and then use it differently without explanation (potentially depriving the reader of some interesting insight into the phenomena under discussion, if the reasons are good).
Secondly, doesn't 'misuse'really just mean 'used differently from how most people use it'? In which case, I think the descriptor probably applies, given the word's recent coinage and limited usage
Let's not pretend this word has decades of deep cultural and etymological history behind it. It's not a technical term or a term of art, it's a meme. It's something Cory Doctorow came up with in a blog post last year that only a few tech people care about. It's a hipster nerd poop joke.
And because it fits cleanly into a description of anything hipster nerds consider to be "turning into shit," it inevitably will, until people get tired of the meme and move on.
Not necessarily because it's better. But because the incentive to enshittify it is lacking.
Closed / proprietary ecosystems tend to be nice for a while. And then disappear without warning. Or turn into crap, like a big sinking ship taking all its users down with it.
FOSS / open platforms tend to fork and/or evolve. And (if they survive) improve over time. That I can deal with.
A good example of the forking / gradual evolution I can live with.
I will second this. I can't count how many projects were started with a clean design and clean early releases (what worked well), to be later messed up with business changes because the original (business) idea didn't work. And those sudden changes are expected to be implemented usually "yesterday", which adds additional momentum to the project's downward spiral.
If the point of a project was to run a business, I would argue the project was messed up from the start if it could not adapt to the changes. Very little software survives contact with the real world.
Scrappy prototype code is written with developer velocity as the most important property. I’ll throw everything in one file. Stick to tools I know well. Copy from other projects. Use crap variable names and generally make a bit of a mess. Something made fast can be remade fast if it’s wrong.
Code you expect to last a long time should have tests, CI and documentation. If I’m serious, I’ll often rewrite the code several times before deciding on a final design.
A good engineer should be able to switch between styles based on the needs of the project. And a good company will make those needs explicit to the engineering team and they’ll be transparent and trustworthy when expectations change. (Eg if a prototype is going to stay in production long term, they need to tell the engineering team).
It’s fine for businesses to pivot. And it’s fine for engineers to make quick prototypes to test a business idea. But everyone needs to be on the same page the whole way along.
So, it’s not as general a term as you are implying and definitely the usage in this article doesn’t make sense.
> It’s not as general a term as you are implying
My point is that even if the term isn't "supposed" to be general it definitely sounds like that to a non-trivial number of people, which is why these sort of discussions crop up almost every time the term gets mentioned in comment threads on this site.
And you said "I'll ensure it gets done by Friday."
Would you expect to be fired for failing to take initiative?
As an aside, it's interesting that the phenomenon of expectations around the meaning of terms goes beyond obvious relationships to existing known words but even to the way a word sounds: https://en.wikipedia.org/wiki/Bouba/kiki_effect
Codebases can enshitified, this is a deliberate action. Having a zombie corpse codebase where bugs grow faster than features is not necessarily enshitification.
An enshitified codebase is created when a product person or a TL spends an undo amount of technical debt on new features or even better anti-features.
Is someone keeping a list of features that qualify for the label? For example browsers like Safari make it very difficult to export bookmarks to another browser like Firefox, even though a simple json file would suffice.
Emshitifaction?
It nicely captures a lot of the usually "fuzzier" things of technical debt, like how code even with no changes and no changing requirement somehow accrues technical debt (you get better, your idealized effort reduces).
That'd also play into the whole metaphor of technical debt.
edit: typo
Linux people seem to tolerate bash partly because they DO have a special purpose modular system. They algorithmically manipulate text in batch rather than realtime mode to produce an output on a regular basis(I'm not sure for what, besides compiling and building, but they all say they have use cases). Apparently it works great if you process lots of text files.
Perhaps a GitHub awesome list. I've been meaning to start one for very common standards in general(18650 batteries, tripod thread, etc).
We have LADSPA and LV2 for audio, shell commands for text, etc, What else do we have, and what should we have?
From my own experience teaching newbies about bash, the only hard part about all the flags was getting newbies comfortable with ignoring them at first and just trusting that `ls` lists files in the current directory, `sed` let's you manipulate text in a smart way, `grep` finds things, and so on. Once they got that in their head and understood that the flags just allow very specific operations that you previously needed to write a lot more code for, most people got it pretty fast and could punch out quick shell scripts to help with their work. But the base commands without a slew of flags was still quite useful for these newbies.
Again though, that's just my experience and it was fairly narrow in scope for accomplishing specific troubleshooting tasks, so that may factor into why I saw success here; I don't know if the people I taught further evolved their scripting skills past basic troubleshooting, though I am fairly confident these persons understand the "kitchen tools" nature of the GNU core utils as many ended up writing scripts on their own for situations we never discussed, and the scripts were pretty okay.
Edit: Typo, changed Exceptions to Extensions in the first paragraph
Why not? What is your worry? That starting a process is too slow? That data transfer is too slow? Something else?
I agree that process startup may be meaningful overhead in the rare cases when the actual computation is very fast. But sharing uncompressed data between programs is essentially zero-cost.
Separate processes look particularly elegant and modular to me. They are much easier to debug and profile than bolted-in dependencies, since the running time and input/output of each process are trivial to isolate. Of course, having both a library and a cli interface for the same computational brick is always better, but the unix-style tool is an essential thing to have.
Process management gets very annoying, very quickly. When you run a helper program soon you have to consider process management. What if it crashes? What if it gets stuck? What if you crash and it remains running? What if the user uses the same program elsewhere? None of this gets work done.
A text stream is an awful API. Many times the called program doesn't intend to provide an API, or doesn't commit to stability. Your code breaks because version 5.0.2 fixed a typo, and you relied on the typo for your parsing.
A text stream is an awful API. Instead of a nice protocol you get to parse lots of text, deal with escaping and quoting. You better hope the called program does it competently. It may well not, then you have a problem.
A text stream is an awful API. No normal program will dump or read a JPEG over stdin/stdout, so if you need to communicate some sort of binary data now you need another communication channel. That may involve commandline arguments, killing, restarting the program, and reestablishing the state. More process management fun.
A text stream is an awful API. You'll find yourself doing things like assembling multiple lines of output into a single coherent concept, and trying to detect where something ends when the program doesn't necessarily provide a clear indication.
A text stream is an awful API. Sometimes programs print stuff before it took effect, and may not ever give you a clear indication of "now it's been applied". You may need to somehow test for it, retry operations, insert wait states.
99% of the effort invested in this doesn't get work done. You're spending it on management that wouldn't exist if you were using a sane API like DBus or similar, where you don't deal with process management, where things are broken down into nice fields, and where the API is intended as an API.
I also believe we could solve many of the integration issues between multiple tools, by having them all optionally spit out a universal, easy / easier to parse format -- even JSON would be good here. So you could do
ls --foo --bar --json | jq ...
and be able to process the output to your heart's content, without having to do any kind of white space parsing; as a free bonus, you get automatic support for file names with embedded white space...The best we have for now is jc (https://kellyjonbrazil.github.io/jc/), which knows how to do this for a closed (but growing?) set of commands:
ls -l *.db | jc --ls | jq .Debugging async rocesses is definitely not an easy task. As is managing them in systems that have no management and supervision capabilities (aka nearly all of them).
What happens when your process errors out? Gets kille dby OS? Gets stuck? Fork bombs?
They are also quite expensive to start in OS terms.
Eclipse isn't universally beloved. I haven't been following Eclipse development very closely of late, but there were versions released in the early 2010s that were universally derided.
This is one example of cherry-picking I noticed in the article.
Also linear relationships in Unix - ever heard of heard of tee?
Also "enshittification" is turning into a meme at this point, I believe mistakenly applied here.
You export your phtoshop design as a .png file and you import it into word. There you go.
Windowed desktop programs like Photoshop and Word do offer interoperability, in the form of files. (Which by the way is another way programs communicate in UNIX, pipes are just a way for that communication to happen on memory).
I think you are comparing apples and oranges here. Maybe if Photoshop and MS Word used the same dedicated underlying binary/application to read and process the png, then maybe it would follow the UNIX philosophy a bit closer?
The unix philosophy is to do one thing and do it well, it doesn't need to conform with expectations about each binary achieving low-level tasks.
What even would be "reading and processing pngs"? Sounds like a dumb technical way to divide responsibilities. I rather have responsibilities of apps be user-defined like "editing images" and "writing documents".
The unix philosophy precedes the development of the UNIX ecosystem, it's anachronistic to understand an OS based on how applications ended up interpreting it.
It's like conflating a constitution of a country for its laws.
To count the number of functions in a rust file you could run:
cat main.rs | grep "^\s*fn\s" | wc -l
This both promotes the "one thing well" model, while revealing its inherent limitations.How many functions are in this rust file?
/* confuse things
fn fo fp
fn fm fl
*/
fn
main() {
println!("Hello, world!");
}
The above grep reports two, when there is only one.It can be made to work if everyone agrees to follow conventions, like always formatting the Rust code, and indenting comments which look like Rust code.
But that means pushing complexity onto people, rather than into the tools. Which you could do at the beginning, when complexity is low, but it gets increasingly more difficult over time.
- Count non-comment tokens within each scope, to look for complex scopes: `foo --tokens --exclude-comments --group-by=scope`
- Count expressions (as opposed to just SLOC) per function: `foo --expressions --group-by=function` piped to `uniq`
- Get the functions and their argument count: `foo --arguments --count --group-by=function`
- Find very short or long names: `foo --names --sort-by=length` piped to `head`/`tail`
See Tree-sitter at https://tree-sitter.github.io/tree-sitter/ , and as an example project which uses it to parse different languages in a line-oriented manipulable way, see https://pypi.org/project/code-ast/ .
At the very least, the hard part - parsing the different languages - is done.
I normally use Python for things like this, but I have gotten burned more times than I care to admit with simple scripts having dependency issues.
I realized that shell languages are not only smaller languages, but are more limited in scope. This means there are less dependency issues to handle.
On top of that PowerShell scripts are easy to stitch together with pipes, so doing things like parallelizing and sending a job to the background as a task, then checking on it later is a lot easier.
Furthermore the stitching together of shell feels very functional, and because the cmdlets only do one thing getting it to be parallelized is infinitely easier. To this day PowerShell is the only code I’ve ever parallelized on purpose.
I then realized this was all possible due to the self contained nature of the cmdlets, which goes back to the Unix philosophy invented in 1969, about doing one thing and one thing well
But near the end, there is a gem, which is the introduction of the Hub and Spoke design.
Now Plugin architectures aren't new by any means. But they've not been discussed as much nowadays, and I think the idea of "Hub and Spoke" describes it very well.
VsCode is in many ways a philosophical successor to Emacs and sits on the opposite end. Yes, it has plugins but it is designed from the top down to work in a particular way, and everything integrates with it. VsCode Extensions don't function on their own, they're not composable tools, they don't expose any agnostic interfaces (except for the LSP). They're designed to work well within the existing ecosystem of VsCode.
It's essentially a lesson in systems thinking. Paraphrasing Russ Ackoff, the complexity of a system is the consequence of the interaction of its parts, not the parts themselves. Microservices and unix tools always ignored this, which is why they fell out of favor. I can split one complex thing into a dozen simple things, but that doesn't make my life easier, because then all the complexity is in glueing them back together. And even worse improving one part doesn't mean you improve the system you care about.
Ackoff always used to give the example of a car. You take the best part of all cars in the world an put them together, you don't have a great car but a pile of junk, the parts don't fit. This is also the issue of unix tooling and microservice architectures, people optimize for the wrong thing.
But I think perf issues did it in from this design choice.
I'm keen to explore Emacs more.
Plugin systems are one of the hardest things in software (but also look the easiest from the outside).
In the world of Lisp and Smalltalk, you have a holistic view of what computation means. In Unix, complexity is pushed outward.
Web browsers, language runtimes, virtual machines and containers are all examples of people expressing their needs by simulating operating systems.
In particular, what is true about Obsidian which wasn't true about Eclipse some 20 years ago?
The example of Obsidian (in my eyes) confirms this. Markdown is a very simple and elegant way to produce simple documentation, but as soon as one needs more it just crumbles.
I agree that the big tools (Word, excel, photoshop, etc) are bloated. These applications have crossed the bridge into 'they do everything about X' long time ago and they work well in their very wide problem space. Just taking excel (or Calc) as an example, I cannot see how using simpler tools I could create pivot tables and aggregate values across multiple sheets.
In other words, you could implememt unix pipe like or plugin like microservices by applying certain constraints.
[1] And can be used to build all other data structures. A tree is a directed acyclic graph, a list is a tree with no more than one child per node, etc
Its just a thought sparked by that interesting diagram at the end of the post and motivated by the fact that the REST architecture is also specified as a set of constraints...
That is why we have restricted domains.
This is all theoretical ofcourse, but my sense is that the debate about microservices is colored first of all by the extra complexity of networked computation (which raises the bar for a succesful architecture) and maybe indeed the absence of well defined constraints that would guide people on different possible best practices for segmenting monoliths.
The WP contact list acted as a shell for all person-to-person communication programs, showing a unified history feed of all communications you can read from a given person. Like RSS but for humans.
It seems like the obvious way to do "lots of small programs" on the phone -- one single place for managing your communication with a person, whether it's Twitter, email, StarCraft, phonecall, chess, WhatsApp, teledildonics, etc.
That's why as a small team it is wrong to adopt 'The Spotify model'. And yet, many companies will do just that. At every level of the corporate ladder there are different appropriate sets of tools and ways of working that will allow you to achieve particular results. Once you reach the Microsoft of the 1990's era level size (or Oracle, or SAP) you can ignore some of the truisms from the leaner days and do different things that would have killed a younger version of your company.
Just like an adult can run and jump further and faster than a toddler.
I think there are two sides of this.
There’s of course a much stronger need to integrate with established systems, no matter the unnecessary complexity, for those companies. They need to provide stronger compatibility guarantees as well as mitigate risk.
The other side is that once you have that many resources, it becomes less necessary to strive for simplicity and uniformity. The value for those companies shifted a long time ago from engineering to business processes.
Companies make people move quickly.
> "If they took enough time for this, they’d be late to market and miss a giant amount of revenue."
They arent worried about markets or revenue, the shareholders are.
> "At some point, the size of the code base and the number of developers eclipses any attempt and understanding and control. "
This is the result of corporate software development. Its software is shit because no one is allowed to sit down and think about problems. Just close the jira ticket and move on to the next one.
To seek refuge, find a software community that cares and participate there.
But as a user I just want the software that solves my problem from start to finish even if it does a worse job and costs more. Even if there are N pieces of software that solves my problem in N steps, that composition isn’t something I generally can or want to do. I just want someone to make a big ball of mud app that does exactly what I need. It’ll try to do whatever everyone else needs too, which is why it is a buggy ball of mud. But it’s still a lot more attractive than having N separate steps and doing any sort of composition myself. Because that invariably requires knowing about files, data formats compatibility between components .
The Unix principle works well in Unix because the people using it are computer people, perhaps even software developers. In the age of everyone using computers (so 1990 and onwards) this breaks down horribly. The big ball of buggy mud do-it-all Software wins hands down every time.
Well, if you're willing to pay, there's a lot more options.
You correctly identify the unix principal only working because the people doing it are computer people. It's designed for a situation where you are the one that has to solve the problem, not one where you can just pay your way out of it.
And while that might break down post 1990 once you start getting non-computer people using computers, I think your analogy breaks down post 2009 or so in the age of ads and everyone expecting everything to be free.
Obviously there is some in-between cases… occasional use is ok.
Just hire someone to make a big ball of mud app.
> it does a worse job and costs more
Yep. Just pay for it.
At my last company, we had developed years ago an API layer that allowed outside developers to build plugins to our web app. We also created a market place in which these developers could host these plugins.
It was an absolute constant headache trying to tease out the bugs that these plugins would introduce. They often slowed the app down to a halt, and because they’re so opaque, users would blame us for the slow experience. I suspect it was because of the fact that their performance impact was so difficult to hold accountable, plug-in developers were never incentivized to make their plugins performant and just hacked them together.
And this is universal. Every app I’ve used with a sufficient amount of capability for extensibility will invariably suffer this problem. Slack and Figma also suffer from this issue after you’ve installed a bunch of plugins. They get really slow and bug out.
It is inevitable: the more code you add to something, the worse it gets.
Unix tools have developed in tandem in the unix ecosystem over decades of rewrites. They rubbed the rough edges off of each other.
The major failing of most "systems design", architecture and companies is a reluctance to replace what is there with something better. No one likes a rewrite.
And this affects microservices too. Once a microservice exists it's boundaries are defined. sure you can rewrite it to be more peformant but you cannot chnage its interfaces or it's scope - something is depending on that.
The ability to reform those interfaces is what makes it possible to grind down and find the one thing to do well.
And that takes lucking into the right fundamentals and being willing and able to make large scale changes.
"Move fast and break things" might actually be excellent advice
There is no contradiction between monolith and DOTADIW. The code unit is the function, so each function does one thing and does it well.
I don't get why microservices is the opposite of large code base : having the code splitted in different files or in different repo, changes nothing about the sum of all lines of code, no ?
doctrine: Effective architecture shifts complexity to the core competency of the development and operational resources.
observation: software engineering suffers from the disconnect between the competencies in development and operation resources.
I mean it likely is quite good based on their velocity and quality. But if we're to learn anything I'd really like to see its source code.
(And if we need understand how to application operates under the hood, it’s entirely possible to use tools like IDA and Ghidra.)
Well it does but... its actual purpose is to concatenate multiple files.
All Unix commands can read files so you don't need any special tools for that.
E.g.
grep "^\s*fn\s" < main.rs | wc -l
or in grep's specific case just grep "^\s*fn\s" main.rs | wc -lclassic bad tool work.
EDIT: imagine a line starting 'pub fn'. correct solution? use parsers for parsing, not hacks.
Now that electron apps spawn their own browser process, then act like the central GUI hub and "do that right" - it seems to be quite the opposite from "do one thing and do it well" from my point of view.