HN ran on a single box in 2018 has anything changed?
news.ycombinator.com
news.ycombinator.com
https://news.ycombinator.com/item?id=28478379
Also, it's still using the Arc programming language
I do get an occasional no-response from firebase, but my guess is that this is really on my ISP. Again, that is also quite rare.
Also there is a public fork of Arc Lisp at https://arclanguage.org/anarki, which is not guaranteed to be stable, and ships with a forum which is definitely not to feature parity with HN.
(I know, restarting the server != rebooting the machine.)
I had to go to one of the more "exotic" server rooms, placed in a basement and saw a very old desktop stuffed on the bottom of a rack. Covered with dust, open side panels, plastic turned yellow, you know how old and forgotten hardware looks like. This had a new-ish post-it with the message: "this runs all of the fax infrastructure. you turn it off, you explain it to the CEO".
They started bouncing paychecks before I got to see the drama come to a head, but I'm obsessive about proper levels of redundancy anymore. Don't build fragility into a system, and at least leave yourself a means to fail gracefully.
Not sure what platforms you've worked on but "6M requests/day" is peanuts in my neighborhood. (No offense to dang and team which do a fantastic job, but that's not high throughput compared to almost any platform I've worked on over the past decade.)
I have a site with ~ 2.5M req/day running on an m3.medium instance that rarely spikes 10% cpu. NGINX is frontline server probably handling a good % of those request directly before passing others via proxy.
It’s not trivial to take advantage of that but that is also much easier to handle than QPS that are distributed across your dataset (IE Facebook or Twitter where everybody is seeing different things).
I disagree that startups with millions of customers handle less traffic. There are probably some exceptions that prove the rule, but once you're into millions of customers you are most likely over 50K rps at peak times.
Product managers are constantly coming with cool ideas that are hard to shard.
At any rate they should be able to serve that easily off one box, so the architecture holds up.
Let's imagine that all the traffic occurs in a single hour (6M per hour) and is Poisson distributed. Even then to reach 10,000 rps would require a second with 6x the average traffic volume for that hour. Admittedly this could theoretically occur when some correlated event causes everyone to browse HN, but historically the site goes down during those.
For comparison our service (also largely US-centric) served 30M views yesterday with a peak volume of ~1000rps.
Yes, 10K is likely an overestimate by some factor, however it's in the realm of possible. We can argue how much of an overestimate, but I don't think this takes away from the main point that I was trying to make above, which is that they are very unlikely to exceed the serving capacity of a single host (I don't know how fast Arc is, but a well-tuned serving system should be able to handle 40K rps for simple to cache content, which HN surely qualifies as).
While 10k rps may be possible it would either require vastly exceeding 6M requests per day (and per hour) or some very non-standard traffic event (this could happen in practice, for instance email asset delivery is increasingly peaky due to modern email clients and can exhibit very large traffic spikes). As HN is primarily visited by human users spanning many organizations such a strongly correlated event seems unlikely.
Traffic is a bit of text, extremely basic CSS, no images, no tracking, no bullshit means that load times are probably instantaneous even on a 56k modem.
Content is a bit of plain text => easy rendering, the entire database probably would fit into memory, so not much in terms of i/o going on.
What I have utterly no idea is how the karma system works. Updating karma counts and comment sorting only on new submissions would help with performance, but on the other side karma "decays" over time, so there needs to be some sort of external cron?
In other words: they are serving exactly what the people who come here are looking for, and little else. Don't get me wrong, there are cases where images and moderately more complex CSS improve a website. On the other hand, many websites go to an extreme that degrades the user experience.
HN can do an incredible amount of easy caching. Besides a small box in the top right with user profile info, the entire page is completely identical for everyone.
People who've never stood up a "bare-metal" server in their life. Who weren't around to see the kind of traffic one shitty late-90s commodity-parts server with slow memory and maybe two processors if you're very lucky and spinning-rust drives, could handle when tuned properly and without horrendously bloated software or being subjected to a pile of bad DB queries written by people who haven't the first clue what they're doing.
For example, what ibm.com ran on in 1998 (<https://en.wikipedia.org/wiki/File:IBM_RS6000_AIX_Servers_IB...>)
Even better with a heat-dissipating case with an integrated M.2 SATA port in the bottom, like the Argon ONE M.2.
I use mine to run multiple .NET 6 microservices, and it's replaced an entire HP Microserver. The entire Pi setup will have paid for itself in under a year in energy savings alone.
I recently had an important infrastructure service die a dumpster fire of a death because it uses mongo, which quietly disables its journal on 32 bit operating systems. (Restored from backup, fwiw.)
Bandwidth maxed at 2400bps (yes, really) and HN was the only news site I trusted not to overload my dripping faucet of an internet connection. The modems built-in browser stripped HN down to the bare minimum to be legible.
I really appreciate the thought that goes into making a lightweight, information heavy site.
They might serve you well if you find some sites that interest you.
[0] https://play.google.com/store/apps/details?id=com.kiwibrowse...
https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...
https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...
Used to enjoy using userscripts + greasemonkey until the extension was removed from the chrome browser store. Tampermonkey for firefox fell out of favor over the Dark Reader extension as default but didn't know uBlock Origin can be used this way.
||youtube.com^$3p,frame,redirect=click2load.html
||youtube-nocookie.com^$3p,frame,redirect=click2load.html
Got this from gorhill, Ublock Origin's developer. He sometimes puts out tips and tricks on his twitter
* https://hckrnews.com/ -> my favorite thread browser
* https://hackerweb.app/ -> a mobile focused app w/ dark mode
I'd love this also, but it's outside of my domain to attempt to pitch in.
Bonus points: available in Firefox mobile, too.
I feel most websites can run on a single instance with decent caching. But owners want reliability and that adds complexity, cost, and technical debt.
Which reduces reliability
See my previous comment in this thread or search "Arc Lisp" and you can find the current distribution of the code and forum software HN is based on. But bear in mind Hacker News itself is somewhat proprietary for business reasons.
I mean, to Hacker News! I can promise most people less work. I just can't promise that to Dan.
I assume not because every comment made would not be visible until cache expired. Unless they have a way to expire the cache (which is not dependable on most CDNs?)
It is now synthesized on global platforms to seize the viral e-markets and aggregate brand schemas on the bleeding edge despite all the challenges in the supply chain.
edit/sorry was joking. I dont know anything about anything.