We’re recently running two machines (master and standby) at M5 Hosting. All of HN runs on a single box, nothing exotic:
CPU: Intel(R) Xeon(R) CPU E5-2637 v4 @ 3.50GHz (3500.07-MHz K8-class CPU)
FreeBSD/SMP: 2 package(s) x 4 core(s) x 2 hardware threads
Mirrored SSDs for data, mirrored magnetic for logs (UFS)You can get the software and language HN is based on here: http://arclanguage.org
Check out around this area of the code to see how simple it is. All just files and directories: https://github.com/wting/hackernews/blob/master/news.arc#L16... .. the beauty of this simple approach is a lack of moving parts, and it's easy to slap Redis on top if you need caching or something.
There is a modern maintained variant at https://github.com/arclanguage/anarki/tree/master/apps/news as well if you want to spin up your own HN-a-like and have the patience.
File syncing between machines is pretty much an easily solved problem. I don't know how they do it, but it could be something like https://syncthing.net/ or even some scripting with `rsync`. Heck, a cronned `tar | gzip | scp` might even be enough for an app whose data isn't exactly mission critical.
Does anyone know of other open source applications with similar architectures like this?
There's a good reason everyone else just uses a relational database, and it isn't because everyone else is addicted to unnecessary complexity.
With filesystem as the storage you don't even need Redis, OS would cache the most recent files anyway.
Other than that, I think it's delightfully ugly and lightweight.
Also, HN apps tend to make it harder to send interesting things to Roam or the laptop or Safari's reading list, the website makes that really convenient.
Having to separately configure individual sites or web apps for dark mode is a nonstarter anyway; if you could do that, would you really want to?
Ideally, you should be able to set your device to dark mode, and everything would follow: every app, every site in the browser.
Some combination of setting your OS to dark mode and using a dark mode extension in the browser sort of approximates that, imperfectly.
We still need the workaround via extensions or Userstyles for the ones that don't implement that, sadly.
[0] https://developer.mozilla.org/en-US/docs/Web/CSS/@media/pref...
The idea that a website needs to be “rich” to be usable is one of the dumbest things the industry has convinced itself of in the last 20 years (following only ‘yaml is a smart way to encode infrastructure’).
It takes substantial wisdom to arrive at an 80% solution and cease fannying about.
What do you prefer?
* doesn't encode Norway to false;
* most formatters for JSON are deterministic.
* doesn't deserialize into arbitrary objects;
YAML, in in constrast...
* YAML is insecure by default and will deserialize into arbitrary objects;
* YAML knows that there's no such thing as wall clock time, there's only number of seconds since midnight;
* YAML has 22 ways of writing true or false, and the parser will silently replace your "strings" with false.
* There are 63 ways of writing multi-line strings;
* A truncated YAML file is still a "valid" YAML file.
JSON is one strict subset, but one that makes smart trade-offs for strictness and machines like error detection and syntax-typed types.
We decided on a different subset of YAML for our users that were modifying config by hand (even more strict than StrictYAML). Some of the biggest features of YAML are that there is no syntax typing, and collection syntax is simple (e.g. also true for JSON, false for TOML).
For example, a string and a number look the same. This seems bad to us developers at first, but the user doesn't have to waste 20 min chasing down an unmatched quote when modifying config in a <textarea>. Beyond that, it's the same amount of work as making sure the JSON is `"age": 20` instead of `"age": "20"`, one just has noisier syntax.
I think the StrictYAML docs have a great breakdown of the advantages: https://hitchdev.com/strictyaml/why-not/
We decided against TOML because nesting is too confusing. https://github.com/toml-lang/toml/issues/846
For example, if it was more user-friendly, it could have links to jump between root comments, because right now very popular top comments tend to accumulate most interactions, and scrolling down several pages to find the next root thread requires effort.
TY!
Yes, I've heard that SO runs on relatively simple and modest infra. And agree that would be a good example.
>HN is not user friendly
How so? I find the HN UX a refreshingly simple and effective experience. It might not have all the bells and whistles of newer discussions fora, but it doesn't obviously need them. I'd say it's a good example of form/function well suited to need. Not perfect perhaps, but very effective.
YMMV of course.
So, I think there is a lot more in common than you think between HN and SO.
And they're tasked with building a product that can handle Google-levels of demand, though they currently only have two customers, neither of them paying.
It indeed is imperative, but not for technical reasons.
Anything HN has had to implement, Reddit has to implement at a generalized, user-facing level, like mod tools.
Frankly, we underestimate how hard forums are, even simple ones. I learned this the hard way rebuilding a popular vBulletin forum into a bespoke forum system.
Every feature people expect from a forum turns into a fractal of smaller moving parts, consideration, and infinite polish. Letting users create and manage forums is an explosion of even more things that used to be simple private /admin tools.
I agree, most software is deceivingly simple from the outside. Once you start building it, you become more humble about the effort required to build anything moderately complex.
Its not that the mod tools are constantly being used, its that there's now potentially far more code complexity for those tools to even exist.
Since I've been here they've added vouching for banned users (and actually warning people beforehand) thread folding, Show HN, making the second chance pool public, thread hiding, the past page, various navigation links and the API. They've also been trying to get a mobile stylesheet to work. They've also mentioned making various changes for spam detection and performance. And the url now automatically loads a canonical version if it finds one, and the title is now automatically edited for brevity. And I've probably missed a few things.
And HN isn't a simple application by any means. Go look at the Arc Forum code - it isn't optimized for readability, or scalability or reliability, but joy - for the vibe of experimental academic hacking on a Lisp. It's made of brain farts. Hacker News is probably significantly more complex than that for being attached to a SV startup company and running 'business code' and whatnot.
Engineers.
[^1]: https://www.freebsd.org/cgi/man.cgi?query=carp&sektion=4
I'm gonna write an alternative which will be WebScale.
j/k of course.
In the microservice or serverless arrangements I've seen, data is scattered across the cloud.
It's common for the dominant factor in performance to be data locality, most times this talk about data locality is about avoiding trips to RAM, or worse disk. But in our "modern" distributed cloud things, finding a bit of data frequently involves a trip over the network. In the monolith world what was once invoking a method on an account object, has become making a HTTP POST to the accounts microservice.
What might have been a microseconds operation in the single server world, might become hundreds of milliseconds in the distributed cloud world. While you can't horizontally scale a single server, your 1000x head start in performance might delay scaling issues for a very long time.
A most excellent paper related to this topic that I think should be mandatory reading before allowing anyone an AWS account is http://www.frankmcsherry.org/assets/COST.pdf :)
If you know some details of the services you're going to host on that hardware, the things you can do while saving a lot of resources is considered as black magic by many people who only deploys microservices to K8S systems.
...and you don't need VMs, containers, K8S and anything.
After understanding these parameters, you can limit the resources of your application by running it under a cgroup. Doing this won't allow a service to surpass the limits you've put onto it, and cgroup will pressure your service when it nears its resource limits.
Also, sharing resources is good. Instead of having 10 web server containers, you can host all of them under a single one with virtual hosts, most of the time. This allows good resource sharing and doing more with less processes.
On the extreme case, I'm running a home server (DNS server, torrent client, synching client, a small HTTP server and an FTP server with some other services) under a 512MB OrangePi Zero. The guy works well, and never locks up. It has plenty of free RAM, and none of the services are choking.
Distributed computing only makes sense when you're starting to deal with millions of daily users.
A new server could have 4 × 64.
Also, to distribute the load you can use an ‘A’ DNS record per server:
The other risk I guess would be that all the NS are AWS/route53 severs, so if that went down they'd be up for about a minutes (looks like the DNS TTL is 2 minutes).
You could host your own NS servers in two different locations on two different providers, you could have a third hot spare server ready to go too, that would allow the service to survive an earthquake flattening San Diego (off site server ready to go), and cope with the loss of AWS DNS. Whether the cost/benefit ratio is there is another matter. I think the serving side of route53 is fairly reliable though (even when the interface to update records fails on a frequent basis), and the cost of being down isn't particularly terrible.
Exactly, Amazon US EAST makes the same point several times a year! :)
I only use VMs and route53, and form what I've seen when us-east-1 is down I can still resolve my domains
What is the upside for HN to add this engineering complexity to avoid a couple of minutes of downtime?
For Amazon too it probably doesn't matter -- if I occasionally can't buy something I'm almost certain to simply try again 30 minutes later
For some services a couple of minutes of downtime isn't good enough. Imagine if the superbowl went dark for 5 seconds just as the winning play happened.
Well, two, presumably. Otherwise you can't reboot without taking the website offline.
Yes you can, since 2008. https://en.wikipedia.org/wiki/Ksplice
Not hi loaded of course, but still hundreds active users every single day interacting with the app.
Recently had to upgrade to the next tier because of growth.
Modern servers are super fast and reliable as long as you know what you’re doing and don’t waste it on unnecessary overheads like k8s etc.
It's kind of incredible that "a few hetzner dedis running k8s" still has better reliability than the Cloud™ does.
K8s does this, kind of, on a higher level, where the downtime is often more noticeable because there's more to restart. But if your application is already highly fault-tolerant, this is just another point of failure.
Of course you can take over some of these concepts, but I don't know any other language that is so well designed around the actor model.
Because people adore stable income with no risks, and because most programmers out there are hacks who will only ever learn one or two ancient languages that lost their competitive advantage 10 years ago.
Take a look at how insanely far you can go with Elixir + Phoenix on a server with 2 vCPUs and 2GB RAM. It's crazy. At least 80% of all web apps ever written will never get to the point where an upgrade would be needed.
So yeah, people love inertia and to parrot the same broken premise quotes like "But StackOverflow shows languages X and Y dominate the industry!".
Nah, it doesn't show that at all. It only shows that programmers are people like all others: going out of the comfortable zone is scary so our real industry is manufacturing bullshit post-hoc rationalizations to conceal the fact that we're scared of change. A mega-productive and efficient change at that.
Social reasons > objective technical reasons. Always has been true.
/rant
For me, the most important thing is I don't have to take care in my deployment pipelines about how many servers I have and how they are named/reachable, to stop Docker containers, to re-create Docker containers, how to deal with logs, letsencrypt cert renewals... all I have to do is point kubectl at the cluster's master, submit the current specification on how the system should look like and I don't have to deal with anything else. If a server or a workload crashes, I don't have to set up monitoring to detect and fail over, Kubernetes will automatically take care of that for me. When I add capacity to the cluster for whatever reason, I enroll the new server with kubeadm and it's online - no entering of new backend IPs in reverse proxies, no bullshit. Maintenance (e.g. OS upgrades, hardware maintenance) becomes a breeze as well: drain the node, do the maintenance, start the node, everything works automatically again.
It's not incredible, it's normal - it's been always like this.
Sometimes you will share the CPU.
Is there evidence cloud compute is more than a few percent slower than on-prem?
Not sure this is true for AWS anymore.
Even a t3.nano gives you 2 vCPUs, which I've always interpreted as meaning they give you 1 core with two threads to ensure that you never share a core with another tenant.
Of course, t2 instances still exist which give 1 vCPU, but there's no reason to create one of those when a t3 gives 2 vCPUs at a lower price. The only reason to be running a t2 is if you created it before t3 existed and just haven't bothered migrating.
Edit: Google's Borg is a very different beast.
Edit: no need to patronize me. I worked on massive scale deployments otherwise I would not be commenting.
The compute overhead to orchestrate these clusters is well worth the simplicity/standardisation/auto-scaling that comes with Kubernetes. Many people have never had to operate VMs in the hundreds or thousands and do not understand the challenges that come with maintaining varied workloads, alongside capacity planning at that scale.
The FAANGs operate the first kind, k8s is mostly aimed at the second kind scale, so its designed "for scale", for some definitions of scale.
We use it to serve ruby with 50 million requests per minute just fine. And the best part is the Horizontal Pod Autoscaler which saves our ass during seasonal spikes.
While serverless/lambda are great I do think K8s is the most flexible way to serve rapidly changing containerized workload at scale.
So I'm not disagreeing with your assertion, but I'd perhaps scope it to saying it's useful overhead at significant organizational scale, but you can certainly operate at a significant technical scale without such things.. and that can be quite fun if you have the team or app for it :)
But for 99% of projects i see, it’s a waste of time and resources (mostly people resources, not just cpu). HN is a perfect example of a project that doesn’t need it, no matter the traffic.
If you need some additional flexibility and scalability over “bare metal” setup, you can go far with just docker compose or swarm until you have no choice but use k8s
Again, if you know what you are doing.
I migrated our services from a very “pet” oriented architecture to a 4-node Proxmox cluster with Docker Swarm deploying containers across four VMs and it worked great. Services we brought up on this infra still to this date have 100% uptime, through updates and server reboots and other events that formerly would have taken sites offline temporarily.
I looked at k8s briefly and it seemed like total overkill. I probably would have spent weeks or months learning k8s for no appreciable advantage over swarm.
I once wrote a software in Rust, a simple API, one binary in one DigitalOcean instance started by systemd, and nothing else. The things has been working nonstop for years making it the most stable piece of long running software I’ve ever written, and I think it all comes from it being simple without any extra/unnecessary complexity added.
I’m not bragging btw, I actually had to contact the user years after I wrote that because I couldn’t believe that the thing was still working but I hadn’t heard from them in years!
Those function as proxies and lower the perception of downtime as well.
I've never experienced the website being inaccessible even for short periods.
https://freakonomics.com/2012/08/this-website-only-open-duri...
I hope another elder will drop others here.
Other than that I have seen no downtime.
And once in a blue moon, HN gets hammered by unusual traffic and becomes somewhat unusable while logged in.
Not all "downs" are reflected there. Last time I remember having really bad perf or non-workable, don't remember - but opening in incognito, you would get cached results fast.
We're all used to effectively beta software that's constantly being updated every day and never final
This might be true on the infrastructure layer, but HN definitly uses "fancy" technology as Paul developed his own lisp that powers HN, Arc :) http://arclanguage.org/
> Arc is designed for exploratory programming: the kind where you decide what to write by writing it. A good medium for exploratory programming is one that makes programs brief and malleable, so that's what we've aimed for. This is a medium for sketching software.
Hell, this is one environment where I bet you could get away with a full async mount.
Does look familiar with interface?
It depends on what you mean by "online" and the service level :
- "online" meaning a HN server responds with something. In this more literal sense, HN always seems to be up.
- "online" meaning normal page load response times. In this sense, HN sometimes times out with "sorry we can't serve your request right now". That seems to happen once a week or once a month. Another example is a super popular thread (e.g. "Trump wins election") that hammers the server and threads take a minute or more to load. This prompts dang to write a comment in the thread asking people to "log out" if they're not posting. This reduces load on the server as rendering pages of signed-out users don't need to query individual user stats, voting counts, hidden comments, etc. This would be a form of adhoc community behavior cooperating to reduce workload on the server rather than spin up extra instances on AWS.
The occasional disruptions to the 2nd meaning of "online" is ok since this is a discussion forum and nobody's revenue is dependent on it. Therefore, it doesn't "need" more uptime than it already has.
PG operating the switchboard.