That Old NetBSD Server, Running Since 2010
it-notes.dragas.net
it-notes.dragas.net
I now understand why I'm not – and will never be – rich. My first (and last) employer complained about my preference for stable and reliable solutions, equating it with lesser profits. According to him, unstable solutions requiring frequent maintenance meant more revenue. To me, a job is well-done when it works consistently, not when it demands constant fixes.
This is a great point. Haven't you noticed the teams that have downtimes constantly are hailed like heros for saving the day countless times. Yet the guys that have services running 24/7 for years without a hiccup will be the first ones on the chopping block when finances get tight.I am also in that boat. I too like systems that never fail, defensive programming, proactive attention. I'll never be rich and most of my company has no idea what I do. It's still the right thing to do.
Boring that never fails = Quality.
These days they largely opt for the GSuite/Squarespace stack or similar to fulfil those needs around here.
Actually I have been considering investing & going BSD only to escape the insanity that has now also contaminated the "Linux world"
Now from my initial observations the opportunities in the saner world are quite scarce and hard to find. But I am pretty sure those companies often face a skill shortage too.
Probably complaining about the rat race / grind mindset which leads to lot of people without strong fundamentals gaming the system and getting high end jobs.
The whole Docker & container craze vs how things are done in FreeBSD worry me much more.
But this is not so much about the technology but more about the fact that as Linux became mainstream & the default choice it doesn't filter out anymore the companies who have no idea what they are doing.
New and thus inherently unstable things may allow order of magnitude jumps in efficiency, and, again, getting onto a growth trajectory.
Dropping VMS and Solaris in 1997 and switching to Linux was switching from stable software and hardware to unstable, but such that had a growth potential. Switching from bare metal and colocation to AWS in 2009, the same. Switching from highly advanced ICE to electric motors and huge lithium batteries in 2007, the same.
If all that you value is stability [ed: of your setup], the amount of wealth you own will also stay the same, at best. Stable things need to be stepping stones into the unstable, uncertain future that may bring something bigger.
The whole move fast, break things and grow quickly is a cancer that compromises everything not a commercial win. When you put yourself before your customers you fucked up.
I'm saying that if you have the same infrastructure over 13 years, and you did not have to upgrade it, or replace its parts, or even migrate off it, you're likely not growing.
All the above things can be done with minimal disruption for the customers, ideally in a completely transparent way.
Of course growth and getting rich is not the only worthy target. Say, OSS is not about getting rich, but it's wonderful and worthy without doubt, in my eyes.
This is of course an extreme, probably unusual example.
Run around panicking with your arse on fire constantly firefighting issue after issue (almost all of which you should have anticipated and already dealt with) making your constantly "red" project go "green"? You're a fucking hero that wins any internal little award/prize for rescuing the project.
Still annoys the crap out of me after 20 years.
Not really my experience, as long as the stable services costs of operation are good. I have witnessed more chopping on the R&D side, where things are never "stable and reliable", but innovative and expensive.
No need to change or outsource something that works technically and financially. Until it becomes irrelevant.
You built everything in AWS via ClickOps? Sounds like an SRE problem.
Your DB is just a giant JSOB object in Postgres? Sounds like a DBRE problem.
It’s extremely frustrating, as someone who has been / is both of those other groups, to see projects being hastily thrown together, knowing that one day it’ll be your problem to solve.
Of course, some managers will be dumb and equate instability with revenue.
Not sure how to be better at making money.
I chose a 1.42Ghz Power PC Mac Mini. I installed Linux, and was very happy with how well it worked, how tiny it was, and how it took a fraction of the the rack space that my web server and other servers took. I thought I might even just use those in the future.
Fast forward a couple years, and the load started increasing. I had used XFS for the mail partition, it ran Qmail and used used Maildirs which tended to accumulate thousands of files per mail directory, and the server was starting to choke. I also avoided rebooting it for years. If I remember correctly, by the end, the server had an 6 year uptime because I was so scared that rebooting it might brick it. But I had a major problem: this Qmail+Vpopmail+SpamAssassin+[dozens of custom tweaks] install had accumulated so many custom hacks, tweaks and patches that I never had confidence that I could do a real downtime-free cut over to a new system without a barrage of complaints.
So I put it off. And I put it off. Fast forward to about 2013 and I decided enough was enough, so instead of doing a fraught cut-over, I just ended email service. Problem solved. Best choice I ever made.
Needless to say, I avoid overly complex, patched configs now.
I was also sixteen when I did that, so I mean, of course I wasn't going to do anything serious with it.
Google and others will happily block your mail or send it to spam folder even if you never sent one SPAM email ever, and have all those technologies you mentioned set up.
Gmail et al have been spam filtering messages from correctly configured mail servers for a decade+ now. All the dkim, dmarc, and spf in the world won't help you if you aren't known to them.
Remember, even if you have everything 100% nailed down with SPF, DKIM, etc. you can still end up on a random blacklist, some of which are basically extortion shakedowns. Now, you can say, "well ignore those losers, who cares?" but sometimes you have a customer who directly or indirectly relies on those blacklists. I certainly do, that's how I found out!
That’s awful.
The person who had the company Gmail administration login had tried to open a ticket, had played with settings, had deleted and re-added the account on the copier many times, but nothing worked.
Never being someone who is content to just wait for someone else to, you know, do their job, and not wanting to deal with the proprietary bullshit that is Gmail, I decided to do something far simpler: set up an SMTP server.
A $12 PogoPlug, an SD card, and a couple of hours later (I had to compile Sendmail from pkgsrc), the printer could scan to email and deliver by smarthosting through a public SMTP server.
People will often tell you that something is a waste of time because they don't know how to do it or because "everyone does something else". Some people tell me that running email servers is a waste of time, but standing up a whole server is infinitely easier than trying to deal with Google.
Someone from the company called me many years later and asked me if I wanted my PogoPlug back. It had been in use for at least half a decade before they replaced that copier.
Turn off atime and devmtime, turn of daily and weekly cron jobs, and log to tmpfs, and SD cards will last for many, many years. I have one that's been running continuously since 2015.
At some point, I pitched the silly idea to IT that I'd figure out what frequency the mic ran at, tune it with gqrx and an SDR, then feed it back to my own microphone using a loopback device in PulseAudio. That day, we had our best-ever voice quality on the remote call, I ended up becoming critical path for all-hands meetings.
Fast forward a few months, the only thing that saved us from building a Raspberry Pi + SDR + Zoom box was another round of funding, with which we bought a proper conferencing system.
However when I look at the personal computers I run today, there is no "desktop" because I prefer textmode and I am running many tiny servers. These serve only me not the internet at-large.
The dream I had at the time of NetBSD 5.1 was a framebuffer where I could launch GUI applications from the command line in VESA textmode, without any context switch (Ctrl-Alt-F1, etc.) back to X11.
If I have the choice between "desktop" and "server" in 2023, I choose server. The "server" is a more useful metaphor for for me than the "desktop" ever was. Eye candy GUIs can be seductive, but IME servers operated with text commands and text configuration files are more powerful and ultimately more useful. As it happens, the server metaphor is quite common. For example, it's used to run the web.
NetBSD 5.1 was great. Arguably one of the best releases in the project's history.
Sometimes I play around with nghttp2 and it just feels totally inflexible. Bad ergonomics for how I use the web.
Using a browser from an advertising company shapes the way people think of the web. It limits thinking and especially thinking outside the box. Look for the replies that seek to discredit those who aberrate from use of the web as mechanism for delivery and receipt of advertising.
If I preferred what was "common", then I would not be using NetBSD. And if it was as popular I suspect it would be less good. Quite confident other NetBSD users have had that thought before. Some software developers stuggle to understand why others might not care about popularity because they live their lives desperately hoping/trying to get others to use the software they write. But I'm not a software developer and I like stuff that is not popular, especially non-popular software.
1. To be truthful, with aria2 it's technically possible but not designed for use in UNIX pipes.
2. Answer: Advertisers, but certainly not me.
That puts you in the vanishingly-small minority. To borrow a phrase from Tom Ptacek: to a decent first approximation, zero people including you are daily-driving text mode for their personal machines.
People describe "desktop-oriented" distros as such because they're filled with niceties such as support for multiple physical displays, sub-pixel LCD hinting, sane defaults for mouse/touchpad input, graphical wifi managers, and usually desktop environments. They ship with these things out of the box. Imagine recommending with a straight face that someone go to starbucks and manually edit wpa_supplicant.conf to get onto the wifi.
People describe distros as "server-oriented" because while you could achieve the aforementioned niceties with them, it's usually a hassle. In my experience BSD (or Gentoo) aficionados tend to eschew the term "hassle" in favor of sayings like "minimal" or "exactly what I want and no more" or the ever-popular but always meaningless "cohesive" when describing those operating systems.
In any case, the "desktop" metaphor might not be useful to you but consider that your use case is far from common.
Linux in a nutshell.
Editing config files is not specific to TUI, you can do it even with GUI editor, and you have to if you want to do anything more complicated with your computer, like route some traffic over VPN and some not, or have a dynamic VPN setup based on what wifi network you're connected to.
Also in the TUI land, there's iwd these days which has nicer end-user text mode UI than wpa_cli, with one command to start a scan, and another to connect. There's also wifi-menu, or nmtui. All pretty much equivalent to the GUI wifi selection menus people use on windows.
What would make it one of the best in your opinion?
I used to like 1.6 and 2.0 a lot, personally.
<?xml version="1.0" encoding="utf-8"?>
<!DOCTYPE html PUBLIC "-//W3C//DTD XHTML 1.0 Strict//EN"
"http://www.w3.org/TR/xhtml1/DTD/xhtml1-strict.dtd">
<html>
<head>
<title>503 Backend unavailable, connection timeout</title>
</head>
<body>
<h1>Error 503 Backend unavailable, connection timeout</h1>
<p>Backend unavailable, connection timeout</p>
<h3>Error 54113</h3>
<p>Details: cache-[redacted]</p>
<hr>
<p>Varnish cache server</p>
</body>
</html>
closedI ran NetBSD systems for years, but I don't think any of the BSD I have operated managed 13 years. Even 3-4 years felt like a big win!
I think how manufacturers of embedded OS looked at the field and went VxWorks, Linux or BSD is interesting too. By no means is it automatic you go to a linux kernel for a small device.
It's ironic that Java was supposed to be that idealized machine, but native virtualization caught up. The JVM has it's benefits, but x86 itself ended up eating java bytecode's lunch, which isn't something anyone predicted. Indeed, even javascript of all things is becoming a strong contender for this role, something no-one predicted either.
Instead, x64 virtualization easily allows to run not just your existing software, but your existing OS, at practically native speed.
Records are indeed a serious change. (Fibers, too.)
JVM runs a bunch of languages that are pretty dissimilar to Java, like Clojure or JRuby. It takes certain creativity, but it works pretty well.
The `struct` bit can refer to a few things: value typing, stack allocation, and explicit memory layout. I think records give you value typing, but not stack allocation or explicit memory layout. They are just syntactic sugar for immutable POJOs.
The failure was the dialup ISP had cancelled their dialup service. I found another one, signed them up for it temporarily and ordered an ADSL line and router for them and swapped that in a couple of weeks later. The compaq was retired.
Some trite syslog analysis suggested it managed 7 years of uptime in one stretch killed only by what looked like a power outage. There was no UPS on it.
I like follow-ups on predictions, so please allow me to request one. When asked in that thread, you said you would "probably controversially" choose Windows Server 2012 on a mid-range HP DL or ML server for a system meant to last over a decade. Almost a decade has passed. Do you think today that this would have been the right choice at the time? I am not questioning your choice as an anti-Windows thing or anything like that. I am genuinely curious.
Well over a decade has passed now and I have nothing to do with Windows whatsoever any more and am running fully Apple on the desktop and Linux on the server side of things. I wouldn't have anything to do with it any more. That was a completely wrong prediction. I inherited a lot of Hyper-V infrastructure with SCVMM which really finished it off.
What did work was CentOS on AWS EC2 though although I'm not sure that's good any more what with the whole RHEL controversy recently.
So for another future prediction which will be equally wrong: I don't have a clue any more and am trying desperately to find a way out of the industry.
They've never been upgraded, they've only been rebooted a few times by linode during maintenance. I have no way of building them again, I have no backups, and yet... They are the production servers for a web app with 250k monthly users (not a commercial venture).
Seemingly I'm fine with this... They've just been that reliable. A python so old I've no migration path, a django so old I've no migration path. But yet they keep working
The only thing updated in a decade is an API built in go, which still compiles on latest go despite being written for pre 1.0 go. And I did replace the postgres server as that had outgrown it's instance. But everything else, a decade old and still fine.
I logged in the other day to discover that SSH didn't initially work as ssh+RSA has been superseded and needed new ssh config just to keep connecting.
There are so many things on these servers that no longer make sense, graphite monitoring that goes nowhere, new relic integration I disabled years ago, linode Longview long deprecated. And yet still the servers work as load balancers, and app servers for ancient python programs.
Very bold… How do you keep it secure with such ancient software? What happens when it goes down and you don’t have backups? I can’t imagine a service that’s both unimportant and has so many users. So many questions.
Only 3 ports are open: 80 443 <another for ssh that isn't 20>
The nginx is just a proxy to the frontend and a file cache, I'm confident I could rebuild it in under an hour even without access to what is there today. Also confident that if it's compromised it can't do much real harm, I will nuke it.
The django is just a thin front end that calls an API, I'm confident that when it gets compromised it can't do further damage as it has no direct access to the DB. And TBH, if/when it dies it will be a kick up the butt to reimplementing this in Go (I've about 60% of that already done but I only look at it once a year for a few evenings)
> I can’t imagine a service that’s both unimportant and has so many users. So many questions.
It's a platform with 300 forums on it.
And I guess if it goes down I discover either the users didn't care (no big loss) or the users really care (and perhaps now the donations would be sufficient to cover the eng time for me to get someone to complete the frontend rewrite in Go).
The whole thing is provided for free, and I figure people get what they pay for. I used to feel very emotionally attached, but now I'm quite YOLO about it. I believe these things should be ephemeral, if it turns out it has run it's time then so be it.
Maybe you don't care about the data running on that server. It's still irresponsible and a disservice to the rest of the Internet to run an unpatched Internet-facing server.
It's not entirely unmonitored, Linode send bandwidth usage warnings and iops warnings, and my users have been the best downtime signal.
I'm very fine with these servers. They belong to hobby web, a low threshold for making something for yourself and others should be encouraged, but also the burden should be carried lightly, and this burden is carried very lightly and if it falls I'm fine with it.
There is nothing what stops you from, at least, dd'ing to some other server, even if at Linode, too.
> I logged in the other day to discover that SSH didn't initially work as ssh+RSA has been superseded and needed new ssh config just to keep connecting.
Yeah, I have a 2T PKI, originally from 2012R2. I needed to re-issue one certificate and I needed an older OpenSSL binary to split it to pem+key pair.
- https://www.theregister.com/2001/04/12/missing_novell_server...
- https://skeptics.stackexchange.com/questions/32502/did-a-com...
But the hardware of that day was maybe not that great, e.g. Compaq or HP servers, so from that standpoint I'd be a bit more sceptical of such a long uptime.
What I've seen ares workers opening a wall with a hammer in an hospital disregarding the fact there was a small shelf with 19" switches on the other side, that they left hanging by the ethernet cables, without even asking someone to call IT. They went on to take their lunchbreak like nothing happened while computers on that floor were without network.
I once saw my father furiously smacking the screen of the new laptop I brought him. It had stopped working. I quietly explained that the power light is off and he's not hooked up to the wall... yup, came back on as soon as he plugged into power.
I guess I see a commonality about people who don't know any better..
https://arstechnica.com/information-technology/2013/03/epic-...
And the forum thread as it looked when it was posted, with proofs:
https://web.archive.org/web/20130401030556/http://arstechnic...
Reboot at least once a month folks, even for those single node critical systems. Better it doesn't boot on a Friday night than mid morning on a Tuesday.
Also, monitor and send alerts :))
It's amazing how many servers that were seemingly running "fine" for months don't boot back up. Memory failures, disks that disappear, random power issues, motherboard/controller failures. As high as 1%.
> The external services were active but inaccessible, wisely kept hidden from potential threats
Probably a good idea.
Server status at 2023-08-27 21:00:32
System status: Database up for 2020.83 days.On the personal project side, I'm currently running NetBSD/cobalt 9.3 on a Cobalt Qube2 microserver, which is an old MIPS server appliance. It lives here:
Mostly it hosts my persistent IRC sessions, and provides a few network resources to some of the old computers in the shop. It requires very little maintenance, pretty much just does what it's supposed to.
I also ran their RaQ series as custom Linux router/firewall boxes for a long time. Still have two CacheRaQ 1s, which were designed to be caching web proxies and have dual Ethernet ports. I think they came out of production 5 or 6 years ago, ran fine with Debian mipsel until the customer upgraded their Internet connection and exceeded the little 150 MHz CPU's capacity!
When I read the article, it said it went down once due to an earthquake 13 years ago. Depending on what it is used for, tweaks may not have been needed, also it may not have been connected directly to the internet, but behind a firewall in another router.
So I say this is true, NetBSD is very stable and it does not need all the sub-systems Linux needs just to be useful. Some of those things are in pkgssrc (like dbus), but if not needed, it is not used.
Whatever it did, and what I know about NetBSD, I would not be surprised other NetBSD systems are still active in some hidden place forgotten about, doing its job without any "thanks" :)
I still use a BSD variant or Solaris clone when I want something super reliable, but now with Linux I am doing an experiment for a system I want around for 10+ years (I do have experience with Linux almost since its release, just think its something I have to handhold more).
I also like in the article the engineer talking with the customer about the reliability of the hardware - I guess the customer was proven right or just lucky for once!
NetBSD is fantastic, I also have a Manjaro box that has been handling some print stuff for me, been running for 3 years (with no intervention) and I often forget it's there until I'm reminded while moving boxes, just quitely working and waiting for the next request.
We had multiple linux instances running for that long (live patched via ksplice, back when Oracle still didn't cut support to other OSes).
Leaving some box unpatched in datacenter for a decade isn't impressive.
Also there is actual risk server not restarted that long just keels over after a power fail for hardware reasons, more frequent (say, every 2 years, if you have kernel patching) reboots at least reduce number of machines that can die at once, as rare as it would be.
> since 2010
I guess I am getting old, too. I have a server running Arch Linux (same installation but updated and rebooted, though) since early 2011 and I would have not called it old by intuition. I just realized it's over 10 years now and not something like 4 like I was "feeling" still. :(
Don’t we all search for stability and reliability