A Generation Lost in the Bazaar (2012)
queue.acm.org
queue.acm.org
The horror shows I have witnessed and had to work with inside some of the companies that have employeed me over the last 15 years have been much worse than anything I have come accross in reasonably large and active FOSS projects.
The Peter Principle lives on strong in both camps.
edits: spelling
I think this failure is more what the article is getting at. What’s under the hood may be shit in many commercial products, in the end they ship something people can use. In the end most users don’t care that the code looks ugly or that the design is messy and full of suboptimal choices. They care that it works and more importantly is usable for them.
Commercial products that fail to find product-market fit die and disappear from the universe.
Open source products that fail just limp along forever in github.
OTOH, i do really believe that good UX is very hard to do without a strong authority in charge, which FOSS often (but not necessarily) lacks.
The "real people" I encounter rarely venture outside a Web browser and, sometimes, an email app. Most "software" they use are Web apps.
And the times I've introduced Free Software alternatives in an environment where they're supported and promoted ... Free Software is in fact often highly useable.
The issue to me seems far less one of utility than of business models:
- Desktop operating systems were dominated by a convicted anti-trust-violating monpoly, which tied (and still largely controls through strong vendor relations) PC sales directly with OS licenses, the infamous "Microsoft Tax".
- "Office Suite" sales were ... tied to that same desktop monopoly.
- With Web 2.0, AJAX-based Web apps and Software as a Service came to dominate major markets, most especially Web Search, but also email, maps, messaging, social media, online video, and the like. Software isn't sold so much as it's advertising supported, which is to say that there's a business model tied to proprietary development models. In much of this space there are also F/OSS alternatives, but network effects come to be a dominant factor, as well as interlocking monopolies (e.g., Web Search, online advertising, browser, and email, in the case of Google). Again, it's business models which drive the capacity to amass and direct immense capital flows, rather than F/OSS development itself, and several of the tools mentioned are in fact Free Software, most notably the Chromium branch of the Google Chrome Browser. Mozilla's Firefox is functionally comparable or superior to Chrome in many aspects, and is itself Free Software.
- In terms of overall usage, mobile devices dominate, and the principle issue here is that App Stores again tend to incentivise and favour proprietary apps. The ten most-used Android apps are Google Play Store (largely integral to Android), Google Maps, YouTube, Facebook, WhatsApp, Instagram (FB), Amazon, Facebook Messenger, Gmail, and Netflix. Four of those belong to Google, three to Facebook. All are SaaS offerings, though Gmail could be substituted by a third-party email app. The picture for iOS is generally comparable.
It is tremendously difficult to disentangle development methodology from business model and attendant access to capital, regards UI/UX quality; and given that several of the cases I'm listing here are F/OSS applications or software, I'd say your handwave is fairly strongly refuted.
Sure you could have a single person responsible for quality within a project, and that may even be a good idea (BDFL is a pattern for a reason). However you can't do that across organizational boundries. Ports will have lots of duplication because the entire point is to bring the wider world together including projects competing against each other.
Like the solution just doesn't seem applicable here. It feels like the equivalent of someone suggesting to fix the problem of firefox & chrome doing the same thing by appointing a single person responsible for web quality. The solution doesn't seem wrong so much as non-sensical.
I don't particularly see the connection to cathedral & bazaar either. Even before that book competing software projects had duplication.
As far as I understand, that—compared to the Linux distro which derives a similar range of capability from a primordial soup of small nuggets of software—was Raymond’s “cathedral” as well, and the original model of the BSDs, before the range of expected functionality was broad enough that distro-ish “ports” became unavoidable.
One example that’s just plain interesting to read to see what things people were exploring (because, let’s admit it, historical Unix source gets plain boring after a while) is AT&T Research’s https://github.com/att/ast.
We all know what trying to make Solaris free beer made to Sun.
Orange juice made from bitter oranges tastes really well when given for free.
Commercial UNIX was certified to run Big Corporate applications; Linux ran FLOSS LAMP stacks.
When x86 hardware became cheap AND reliable, and web apps took off, Linux became the better choice. The commercial UNIX vendors didn't adjust fast enough and were left behind.
And now we have a network effect, where very few developers/ engineers know HP-UX or AIX, and software is being written specifically for Linux that would have to be ported to HP-UX or AIX.
Some 'law' about always something, somewhere depending on whatever got built previously. Documented or not.
Of course there's the rewrite. Read: starting over with a better design than existing one. Yet another skill that few mortals possess when it comes to intricate machinery like an OS.
And then a whole bunch of functionality must be re-implemented. Which takes countless human-hours (always in short supply). If not years.
That leaves the option to replace pieces one by one with better ones, leaving overall structure (mostly) in place / working. Oh wait, that's what is usually done. Rip something out, plug the hole, move on.
Hyrum's Law in comic form: https://xkcd.com/1172/
This immediately reminded me of Helm charts (Kubernetes) and their implementation with Go templates, that work in the syntatic level instead of semantic, which makes it unnecessarily hard to operate it.
It disturbs me that Helm templates became a defacto standard for k8s. Something went very wrong there.
If you're not doing a unicorn style deployment (100% of users) then you're SOL.
K8s is probably the best/closest thing we have to a workable service "package manager" for server environments. It's a fairly reasonable way of packaging and deploying random stuff to random servers in a sustainable and secure fashion when all you're trying to manage is a small dev pipeline or a prototype/homelab type setup. Instead it feels like the community is trying to skip right over some of the most beneficial use cases for the tech and go straight to nonsense unicorn land.
I think the issue is that for almost everyone tens of servers is the sweet spot. Most teams and even most companies don't need to manage more than tens of servers (how many teams even in large companies actually use or even touch more than 100 servers in prod?)
But kubernetes can scale to thousands.
The leap past those 2 orders of magnitude chnages the game and makes the seeet spot out of reach for the majority. It's like we are taking F1 racing cars on our daily commute and stopping off for groceries
No for almost everyone one server with some packages installed is the sweet spot.
> What if it goes down?
You users will survive a little downtime, if it happens at the right time they wont even know
> But containers
For this use case? A solution looking for a problem
> Why?
Don't be a company with more servers than users
For most it really isn't.
For single node deployments of containers, Docker Compose has usually been enough.
For multiple node deployments, Docker Swarm (possibly with something like Portainer) has worked well for years: it's easy to create clusters, they don't require much in the way of resources and running workloads is pretty simple, definitely in part thanks to the Compose spec.
Of course, there have been issues over the years (CentOS 7 I think needed masquerading for networking to work properly back then) and people might be conflicted about some things - personally I love my "Ingress" being just a web server container with whatever configuration I give it, others might want Traefik or something else, whereas more advanced use cases like task scheduling and replicated storage across the cluster don't have anywhere near as many solutions as Kubernetes does.
Then again, none of those are a deal breaker for me. Of course, something like Podman can also be considered, or maybe Hashicorp Nomad if you don't mind HCL.
I don't know much about k8s, but this reaction reminds me very much of how I felt when I first learned CMake. "Wait we're replacing all this crap with.. this crap?"
Few forces favor convergence and standardization in open source. After a few decades of this, we have far too much duplicative stuff.
It's really hard to clean that up. At one time I tried to get the 6 or 8 Python packages that parse ISO8601 date strings standardized. All of them had bugs. It took 6 years of bikeshedding discussions.[1] Issue filed in 2012, patch applied in 2018. For something that's a few screens of code.
This isn't just true of standards, it's also true of implementations of a standard.
Thank you for that!
But, the reason for the described situation is no mystery, and may not have that much to do with the bazaar. Programmers like to reinvent stuff, as well as use what they’re familiar with. Those two facts alone explain a lot of the criticisms.
I suppose, in a properly managed cathedral, these tendencies might be somewhat mitigated. But how many cathedrals are properly managed, or have the luxury to “do things right”?
The whole “worse is better” debate is relevant here. Worse may not be better because it’s actually better, but because it’s more easily achievable, particularly by people without infinite time, expertise, and money.
It would be interesting to see a company provide professional development so that programmers feel less pressure to do so in their day to day jobs.
But it's probably worth pointing out that while cathedrals are built, with one architect's vision, not all people are capable of envisioning a cathedral. Many see only ugly 4-walled-flat-roof sprawl. They might build it really well, but its still sprawl.
Equally, things improve with practice and time. My early work may be "single vision", but I'm more skilled now, my vision is clearer and my execution of that vision is improved.
Wonderful software, like wonderful buildings, are wonderful partly because it is not common.
Good vision is partly driven by talent, partly by experience - good software is the result of both great architecture and a capable team willing to work in service of that vision.
Of course all good builders want to be architects, and few want to be remembered as the successor to someone else's vision.
The problem of duplicated dependencies in linux comes from doing too little splitting of packages, not too much. You get a lot less of this on a platform like nodejs where people properly split up functions into much smaller libraries, because the dependency management is much more effective. (Of course people on HN moan more when you do that, for no discernible reason, as if having 1000 dependencies is somehow morally inferior to having 10).
Ultimately a certain amount of duplication is necessary and healthy. There is a nugget of real criticism to be found here - the "Bazaar" nature of open-source OSes makes it very difficult to deprecate and remove anything, where a more centrally-managed system is able to make decisions like "we'll stop using or supporting Perl from 2020" and enforce them. But "herp derp 122 packages herp derp 22 tests" is not what we should be focusing on.
It's not morally wrong, it's wrong for auditing and general change management. Too many suppliers of small dependencies you have to check.
This is important for larger companies or for mission critical software in sensitive domains.
His beef is that people are lazy. If there is a perl tool to munge something you use it, even if the rest of the process uses Python. This is the reductionism of "if you start in awk stay in awk" which is good, honest and .. hard.
Ports is better than APT. apt, has the disease of 'I made a bundle for you but the name is not indicative of whats in it' -where Ports has metaports which are explicit what they pull in to make the "thing" they meta over.
APT has the "-dev" norm now. thats ok, but sort of stupid too. How can it possibly make sense to have to pull in the -dev thing to use its shared libs standalone in another package? (I have seen this)
Ports is the bazaar but with sub-zones. Like the Grand Bazaar in Istanbul, things tend to clump together. Sometimes you wonder why its in sysadmin/ and not text/ but in the end its birds-of-a-feather.
Homebrew ditched M4 for Ruby. Gak.
CPAN is the enemy.
I remember great pieces of software from before the Internet era. The grammatical tense that applies to them is "past". They are dead and gone and only exist as "hey, remember Code Warrior?" "remember the LISP Machine?" "remember Lotus Notes?" Millions and millions of dollars, hours, lines of code all just gone. They live on as inspirations but not as artifacts.
The "cathedral" model doesn't produce cathedrals, it produces photos of cathedrals that were abandoned and eventually collapsed and were forgotten.
The bazaar has produced code that outlives its authors, its sponsors, its original communities. You can Tanenbaum Linux all day long but it will outlive the cockroaches. That's the real monument.
In the article, PHK seem to assume that it’s a good idea that each user compile all their applications locally. I know this is the norm in some places (e.g. Gentoo?), but is it really a reasonable expection of users in general? I consider myself a power user, but if I want to install a tool, in most cases I just fetch it using Homebrew and I will get a binary, so no need to compile it at all. Compilation and packaging might be complex for some of those tools, but it’s “out-sorced”. IMHO a system such as macOS + Homebrew works well from the end-user perspective.
Do I say that just because I’m lost in the bazaar? Have I even seen one of these cathedrals that PHK refers to? Not sure. What would count as a cathedral? HP-UX? Symbolics’ Genera? Would OpenBSD count?
Strongly disagree. We need to get back to the utopian ideals of the early web where normal people were encouraged to cobble together weird and wonderful websites.
Seriously though, the problem is that the Cathedral is a Utopian fantasy that has never once materialized in our timeline and the people who attempt to instantiate it invariably create something much, much worse than the Bazaar.
It may look like it's made of stone, but those arches are made of painted poo.
Sure, the Bazaar is ugly and even offensive to the refined eye of the great priesthood, but these things are almost always far more useful because they come from necessity, closer to the person with real needs.
So true, I've yet to see anything of quality produced by groupthink.
Autoconf doesn't exist because "the bazaar" is full of lazy incompetents and only cathedrals have halfway decent design. It comes from the bad old days when "open" meant "spec documents are available under RAND terms" and had little to do with what we now call open source. Unix had splintered into lots of different idiosyncratic systems, each with its own configuration, and the GNU stuff had to run on all of them. Autoconf provided a portable way to paper over all those differences.
The same with libtool. The Unix splintering had already happened when shared libs became relevant for Unix. There wasn't even agreement among the various vendors on how to implement shared libs! For example, how do you resolve the fact that a dynamic/shared library could be loaded anywhere into a process's address space, making absolute pointers inside the library invalid? There were two main schools of thought:
* when linking the library, replace all exported pointers with placeholder values. Then when the dynamic linker loads the library, it swizzles the placeholder values back into pointers. (Windows)
* use position-independent code, taking advantage of relative jumps, calls, and data accesses to create an object file that can be loaded anywhere. (ELF)
I've seen some half-measures as well:
* give each library a fixed address where it's loaded every time (Linux a.out)
* punt on symbol resolution; hand the programmer a base address and have them call into the library at fixed offsets from the base (AmigaOS through 3.x)
Each of the implementation strategies for shared libs required a different constellation of compiler and linker flags. Again, libtool papered over these and made it tractable to write shared libs for a variety of operating systems. It is true that these days, compiling with 'gcc -fPIC' and linking with 'gcc -shared' will probably do what you want. But that assumes a modern enough OS and toolchain, something that GNU couldn't and still cannot in some cases necessarily assume.
The purpose of both these programs is to portably solve problems that were created by multiple, competing cathedrals. I don't like Autoconf or Libtool. They're unholy messes. But they're largely that way out of necessity considering the problems they solve, problems which would otherwise be nigh-intractable -- not because their developers were lazy or incompetent.
And no, Meson does not solve the same problem, it punts on working on any OS that's insufficiently modern.
The neat thing about the open source bazaar is that it can produce reference implementations for things everybody agrees should be in the finished product, reducing the friction of competing implementations. An example is the 86open consortium, an effort by various proprietary x86 Unix vendors (mainly BSDi, Sun, and SCO) to agree on a single binary standard for x86 Unixes. Well, the vendors got together fully expecting to get into the weeds with meetings and committees and stuff... but what they found was all their products had Linux/ELF kernel personalities already built in. So they declared Linux/ELF to be the binary standard and disbanded the 86open initiative!
IMHO, it still makes sense in 2023. Suppose you want to ship & build your program on the latest Ubuntu, OpenWRT, 5 year old NetBSD or whatever SunOS user might be running in his basement. In that case, Autotools is your only bet without you having to run all those OSes to test if your programs build correctly. If you scan through GitHub, most projects support only the "latest" whatever because modern build systems aren't capable of anything else without significant workarounds. And you get that for free with a few autoconf & automake expressions.
They are not perfect, but the problem they are solving is nasty.
Behind that scary door is the war of architects and their egos, splintering half-completed standards, and nightmares from the past like CORBA etc -- maybe-decent on paper, but good luck ever finding an open and usable and actual quality, working implementation.
What we need is to develop consensus on how to do things, and that, unfortunately, is messy, takes time, and involves people compromising all over the place. And I'm not convinced that committees are the right instrument for that either.
The essay's slagging of the .com era is funny to me as well; because I lived through that era and was absolutely one of the messy green hackers/devs that he is complaining about. Funny thing is though that things shipped, still, and it led to a lot of people being able to make web pages, write code, and get themselves out on the Internet who hadn't had that privilege before.
And it wasn't the quality of the software that caused the .com implosion after, it was the quality of the business plans.
And I also saw what came after, the kind of reaction to the .com chaos; it was all J2EE and EJBs and big enterprise application servers, and I never worked in a single shop where that actually led to anything quality and saw more than one where a team that was doing Perl+CGI+Apache ate itself by rewriting everything "the right way." A few years later, the trend swung back the other way; Ruby on Rails, web2.0 etc. etc.
Software development is messy. That's just how it is.