How to Serve over 100K Web Pages a Day on a Slower Home Internet Connection
cheapskatesguide.org
cheapskatesguide.org
The ideal world is one where environmental sustainability and server efficiency are hand in hand like this.
I also host a Discourse forum. It's pretty snappy for being in my closet (https://forum.webfpga.com) Beats paying DigitalOcean $15/month ($180/year!) for a machine with 2vCPU + 2 GB RAM (the minimum for Discourse).
I think more people should consider self-hosting, especially if it's not critical. It makes the Internet more diverse for sure.
I'd like to imagine that my content is not DDOS-worthy though :)
Also, with IPv6 you have plenty of IPs...
The DNS timeout is set relatively low, 5 minutes. I used Google Cloud DNS because it’s scalable, cheap (the setup costs under $2/month for low traffic) and has a great API and command line tool. Just took a half day of tinkering to get it all set up and most days I forget it’s there. (I should probably set up monitoring, but it’s fine most days...)
If using CloudFlare, it’s possible you could push your dynamic IP to CF as a DNS host. Haven’t looked into it yet.
And of course there’s the DynDNS approach but most of those approaches seemed more complicated or costly these days. I also use gcloud to update the DNS for LetsEncrypt roughly once a month, so it does double duty. The language for the gcloud DNS CLI tool is a bit technical but try it, test it out, and it’ll start to make more sense by example: https://cloud.google.com/sdk/gcloud/reference/dns
For a bit more security, the apps I’m running are in their own virtual machines and I’ve tried to enable default auto updates everywhere. I keep telling myself I’ll set up something more complicated to create a continuous deployment pipeline with test cases and automatic rollback but I haven’t yet gotten around to it. (Though it would be fun to set up my own virtual PPP server at some point for running test scenarios, etc.)
since i'm using mac, you may need to change ipconfig command to hostname -i/-I)
DO_INT={internal_record_id}
DO_EXT={external_record_id}
DO_BASE=https://api.digitalocean.com/v2/domains/{yourdomain}
DO_TOKEN={your_token}
DO_TYPE=content-type:application/json
*/15 */1 * * * (set -o pipefail; /usr/sbin/ipconfig getifaddr en0 | curl -sX PUT -H "${DO_TYPE}" -H "Authorization: Bearer ${DO_TOKEN}" "${DO_BASE}/records/${DO_INT}" -d "{ \"data\": \"$(cat)\" }" | tee /tmp/ip.int;) &>/dev/null
*/15 */1 * * * (set -o pipefail; curl -sf wtfismyip.com/text | curl -sX PUT -H "${DO_TYPE}" -H "Authorization: Bearer ${DO_TOKEN}" "${DO_BASE}/records/${DO_EXT}" -d "{ \"data\": \"$(cat)\" }" | tee /tmp/ip.ext;) &>/dev/nullA couple of thoughts though for Dynamic IP:
- I have my DNS on cloudflare, and I make a script that runs ever minute to checkif my ip changed, and if it did, than update the DNS (this was on a Cloud server and I didn't buy a static IP). - I also use pfSense, and it updates dynamic DNS as well.
My previous ISP, Comcast, required a business plan, statics were $5/month, and they even changed the static IP on me, once. (They issued the same IP to another customer. I complained, and they told me I had to rotate.) Did have IPv6, though.
I once had an ISP with a 3600 hour lease on a dynamic IP. For small things like game servers or short-term file hosting it may as well have been static, but the bi-annual IP change was usually pretty jarring.
Cool project by the way! I'm thinking doing the same with a Rpy..
(Of course, as many people will tell me, this already exists. But sometimes, it's okay to do things for the sake of learning... come on, HN)
Desktop browsers ignore the viewport tag and mobile browsers respect it. The page is rendered how you want everywhere for little effort.
ISPs there are usually monopolies or duopolies so you run the risk of having to buy slower/pricier service, or being totally offline, if one shuts you off.
The guy has less than 1Mbps upload bandwidth.
I don't particularly agree with it, but thats what it is.
If an ISP is concerned about traffic volume, take out the botnets hammering away at the doors — the occasional guy's private website getting slashdotted is small, temporary potatoes.
It does, however, have a policy against 'excessive usage', which is left conpletelt undefined and they have a policy of 'Acceptable Content' which is equally vague and ill-defined.
So I guess they do have recourse if i started hosting a pornsite and it became popular, but that's about it.
Thanks to restrictive firewalls, much incoming and outgoing traffic to services is tunneled over HTTP/HTTPS and uses the standard 80/443 ports. There are also legitimate reasons to have 80/443 open such as access to security systems, your home router, etc. They can't block incoming 80 or 443 without breaking a lot of things, so they won't. It's not the 90's anymore.
You really have to worry more about hitting your 1TB Comcast bandwidth cap.
In my benchmarks I have found that a laptop with 4th gen i5 beats a $80/month Azure VM. You are paying thoufh the nose for redundancy and stuff
Scaleway and OVH also have cheap bare metal offers if you don't like to share.
Also, despite living in the US, I like that my server is in in Paris. It’s EU and sounds super posh. Ha
ipfs resolve -r /ipns/notryan.com
> Error: context deadline exceeded
I tried the ipns multihash too: ipfs resolve -r /ipns/QmZaQZyPXjyXWB2FZUUXV9Ft5qciCDMX6CPP2FmiuPSyAq
> Error: context deadline exceeded
You may not be forwarding your swarm address port, 4001 by default, to your laptop.
I'm pretty new to ipfs so take this with a grain of salt.edit: It's worth noting that I can access other ipfs content using my gateway :)
edit2: I also noticed you're hardcoding the local gateway port to 8080 for your header background image. This port is configurable, mine is set to 5080.
The blue banner was just a hack where the referenced jpeg loads if available. It's a call to url("http://127.0.0.1:8080/ipfs/QmX63Pan72CB2vMhNiFMKxKgEhmffbkgE..."); which is a Monet painting I like.
Basically, if you've got IPFS you'll see a bar. If not, you'll see nothing. Did that at least work? If you had `ipfs daemon` running in the background on port 8080, you should've seen something.
Also how many people do you think are not running on port 8080? It's annoying for sure as a dev, when every program decides it wants its local web server to be 8080.
Good times. I did a writeup 5 years ago how it looked like: https://blog.haschek.at/2015-my-company-just-turned-10.html (down at "Our Hardware")
That way you can own your data and serve more visitors.
Maybe things like Mediagoblin or IPFS could do that job.
1. Page Loading Issues are irrelevant, unless it's because of large items served over limited bandwidth. 2. Static vs Dynamic webpages are irrelevant, if the pages themselves are small. Dynamic of course incurs some CPU on the server-side, but that is a machine not bandwidth issue. 3. Limiting the amount of data is obviously important. 4. Number of requests to the server is only relevant for the request-size, CPU and tcp-overhead (which can be alleviated via multiplexing). 5. Yes, do compress the pages. 6. Agree, website development kits often makes the pages much larger. 7. Certainly, but this shouldn't be necessary if you have good cache-headers.
One thing that was not mentioned, was ensuring that static items are cachable by the browser. This has a huge impact.
This does not mean I disagree with your other statements, I too found it a mix bag of somewhat obvious or slow-server specific (not just low-bandwidth)
Huh? Sure, it's not a bandwidth issue, but the article is about more than serving on low-bandwidth. It's about using limited hardware more generally.
From the first two sentences:
> The cheapest way to run a website is from the home Internet connection that you already have. It's essentially free if you host your website on a very low-power computer like a Raspberry Pi.
I guess the headline refers to bandwidth, and there's some bandwidth-specific framing, but the article is discussing more than just low-bandwidth situations.
I run some surprisingly-spiky-and-high-traffic blogs for independent journalists and authors, that kind of thing. Lots of media files.
There are two ways this typically goes: either you use some platform, and try and get a custom domain name to badge it with, or else you imagine some complicated content-management system with app servers, databases, elastic search clusters etc?
At the time wordpress etc weren't attractive. I have no idea what that landscape is like now, or even what was so unattractive about wp then, but anyway...
So we're doing it with an old python tornado webserver with one core, 256MB RAM VM at a small hosting provider. (I think we started with 128MB RAM, but that offering got discontinued years ago. It might now be 512MB, I'd have to check. Whateever it is, its the smallest VM we can buy.) The webserver is started-if-crashed using cron minute and flock -n.
The key part of the equation is that the hosting provider I use did away with monthly quotas. Instead, they just throttle bandwidth. So when the HN or some other crowd descends, the pages just take longer to load. There is never the risk of a nastygram asking for more money or threatening to turn off stuff or error messages saying some backend db is unavailable etc.
Total cost? Under $20/month. I think domains cost more than hosting.
The last time I even checked up on this little vm? More than a year ago, I think. Perhaps two? Hmm, maybe I should search for the ssh details...
My personal blog is static and is on gh-pages. A fine enough choice for techies.
And I think that more than half of the monthly cost is actually the amortized domain renewal fees etc, not the vms themselves.
So we could probably save a few dollars if we shopped around and moved? But I've just spent more time writing on HN today than I normally spend in a year on thinking about these old servers....
(I used RootBSD back then and they only gave you 256MB for $20 ;-)
Further, your users will get lower latency and faster downloads when accessing one of the globally-distributed edge caches. And lastly, you don't have to expose your own IP address, which puts at risk of DDoS.
Does anyone have a good argument for self-hosting an entirely static website?
Are you sure?
> The desired hostname is not encrypted, so an eavesdropper can see which site is being requested. [0]
ESNI is meant to fix this but isn't, yet, widely deployed (AFAIK).
---
Let's face facts. Most personal sites won't see a giant boost in traffic unexpectedly. Which means hosting and paying a cloud provider is completely unnecessary.
If you happen to become popular enough (the 1%) THEN you start to consider some scaling patterns.
There's also a distinct vendor portability advantage. If you want to move your self-hosted static site from your home connection somewhere else, it's pretty straightforward. But if I want to move my static sites from Amazon to some other cloud vendor, I basically have to start my Terraform from scratch.
Caring about the privacy of your readers.
On top of that, the Snowden revelations.
Having said that, I think NearlyFreeSpeech provides a great web hosting service for pennies a day for those who don't want to host at home. I've used them for years. https://www.nearlyfreespeech.net/services/pricing
(This assumes you already have an internet connection at home, of course)
You also might miss out on a public IPv4 address, but that's whole 'nother issue... FRP (fast reverse proxy) is a decent workaround for this.
I bought a Dell precision 4600m (3rd gen i7, 4c8t, 16gb ram, 480gb ssd + 750gb hard disk) for 50€+vat.
It was taking the dust on a shelf.
It's perfect as a low end home server.
I was doing this for some time and all the nuances with it (assuming electricity and internet are cheap, which isn't the case) makes it still unprofitable on the long run. ISPs usually have really bad upload, port blacklist (e.g. irc), connection unexpectedly drops during the night or day, DNS resolvers or forwarders usually are miss-configured and so on. Support for requests outside "my internet is not working" is often zero (e.g. "your DNS isn't configured properly").
I didn't even start with DDoS (ssh, http) attacks, port scans and even crawler hits (some crawlers by even reputable search engines can be nasty) - all these things will eat your monthly quota without even serving a byte of your site to human visitor. And don't forget visitor impatience these days - if site isn't loaded under a second or two, no more visits :)
Get a cheap openvz box from lowendbox which will cost you between 3-15$ a year.
KVM is a better choice for public-facing sites, even if it's pricier.
S3 hosting + Cloudflare SSL free tiers
(One of the reasons I moved it off my LAN was security - not providing an ingress point, etc.)
The thing with personal home websites is that there's really no actual problem if the site gets overloaded, or if it goes down for a day or a week or is only up intermittently at all.
These requirements of constant availability and massive scaling aren't universal requirements. It's okay if a wave of massive attention is more than your upstream can support. If people are interested they'll come back. If they don't it's fine too.
Your site will still be viewable even after you turn your local node off until all traffic goes to zero and the caches eventually expire.
If I’m having trouble viewing a page on someone’s hammered server I either look in Google’s cache or use links/elinks to grab the text (usually what I’m interested in).
Yes, it does make deploys easier. You don't "have to think" anymore. Well, that's the argument.
I self-host by creating a git bare repo on a server, ala `git init --bare mysite.com.git` and cloning that repo again for the web server (`git clone mysite.com.git mysite.com`). Now add this git hook in the bare repo:
#!/bin/bash
cd /home/ryan/mysite.com
unset GIT_DIR
git restore .
git pull
exec ./build.sh >/dev/null
Now my deploys take a second and become visible immediately. If you are intending on serving 1000s of requests per second, sure you might need CDNs and caching. But... if your hosting a portfolio/blog website, you might not need it. On top of that, demonstrating knowledge of backends is never a bad thing for a personal website.2. Using a good CDN ensures your site is fairly responsive on (hopefully) all continents. Your server probably isn’t. Your closet server — even less.
As for responsiveness, it’s a good thing to proactively prepare for an HN hug of death ;) It also starts to matter when the traffic crosses into 100k a day, as is the case in TFA.
https://www.grahn.io/posts/2020-02-08-s3-vs-b2-static-web-ho...
The platform is open-source: https://github.com/tinspin/sprout
Still going....
BTW, your drag and drop triggers on right click, so weird things happen when using the right click menu.
WEBP can do better (uploaded as PNG): https://i.imgur.com/w8XnHda.png
I only add compression on a need basis manually. CPU is the most scarce resource.
Also, gzipping a GIF is actually likely to make it bigger!
Where's the license?
Contrapoint: content must be static
Another option is to reverse proxy to a cloud provider that can provide a static IP.
also prevents leaking your home IP address.
Additionally, their uploader CLI is spyware, and their corporate stance on this is that you agreed to their uploader tool spying on you when you created your Netlify account.
They also can’t pull from private git hosting, only the big public ones (which are themselves ethically questionable), which makes using Netlify for builds a bit of an issue if you, for example, self host Gitea or GitLab.
I use a PaaS (sort of a self-hosted Heroku) called CapRover that pulls/builds from a self-hosted Gitea on branch change. It’s a relatively small change to drop a two-stage Dockerfile into an jekyll/hugo repo that does the build in step one then copies only the resulting static files into a totally off the shelf/vanilla nginx container in stage 2 for runtime hosting.
For my main website, I still have cloudflare in front of that.
What you’re doing is exactly the sort of thing I don’t want to worry about
It even sent a “telemetry disabled” telemetry event when/after you explicitly indicated you didn’t want telemetry sent, until I repeatedly complained (they brushed it off initially): https://github.com/netlify/cli/issues/739
I just don’t believe that the company has any meaningful training or priorities around privacy.
As for the types of content you can’t post, I encourage you to peruse their TOS:
An excerpt of some of their prohibitions:
> Content with the sole purpose of causing harm or inciting hate, or content that could be reasonably considered as slanderous or libelous.
I personally would like to be able to post political cartoons or other political content expressing and inciting hate toward, for example, violent or inhuman ideologies, and those would be posted for the express purpose of causing harm and damage to the political campaigns they target.
Netlify should not be policing legal, political speech on their platform.
You also aren’t allowed to host files over 10MB(!) so that rules out hosting your own music, podcasts, high res photography, most types of software downloads, or videos on your website, all pretty normal/standard things to host on a website in 2020.
Is this true? If so, then it's a good reason for me not to enable SSL/TLS on my sites that don't need it (e.g. read-only documents or blog posts).
That said, SSL is useful on read-only documents: it provides integrity. Without SSL anyone between your server and the end user can just change the page content, whether that's for misinformation, false-flag-attacks, inserting malware, scanning for censored content (which is known to happen) or for inserting ads (which also happens).
It basically boils down to the old KISS principle "keep-it-stupid-simple".
Actually, that's wrong in an important way.
KISS is "keep it simple, stupid", with the emphasis on the last word, referring to you.