100k+ page views a month for $5 with a self-hosted static site
runninginproduction.com
runninginproduction.com
My site sits on a $5 DO droplet and I have never even come close to hitting any kind of limit despite the occasional traffic spike.
Dynamic content management systems are a convenience for the developer/maintainer of the site but push additional requirements onto the hosting CPU (and the client for javascript heavy sites). The cost of these requirements is often not clear but usually much larger than you might expect.
Would be kind of interesting just to have a point of reference (even if it ends up being super beefy and overkill since you mentioned CPU load is low).
The bulk of the load will be negotiating the TLS sessions, and modern ciphers are both fast and hardware accelerated, and that’s assuming something upstream isn’t already taking care of the TLS.
At $CORP, the web application went down when some heavy Wordpress traffic slowed down the shared DB :/
It's not optized at all.
At the time I was tasked with evaluating it as a "cheaper" replacement for an expensive, managed CMS. My conclusion was it would end up being more expensive to accommodate all of this inefficiency with caching galore not to mention support and dealing with a loss of dynamic web sites.
I looked again last year, assuming it had greatly improved, had replaced all that with GRPC or graphql and nope, still horrendously designed.
I wonder if it's intentional or just a case of support for backwards compatibly.
Now, don't get me wrong, I use and like Wordpress. There is not more versatile yet approachable blog/CMS/whatever out there, the ecosystem is huge and it's my go to tool when I want a quick and dirty solution.
But it doesn't change the fact it has been badly designed from the start.
Popularity in software doesn't reward good design. It rewards solving problems.
It isn't harder to call a script that regenerates the static files.
It is just that dynamic content management systems are more common.
Anybody who was around during the first blogwave in the early 2000s knows the score: after a while, with static generators you end up waiting a lot after hitting “publish”. And god help you when you change the template over the whole site.
Of course, dynamic systems have their own tradeoffs, but they tend to scale better by volume content. You can then slap a simple cache in front and be done with it until views get in the hundred millions.
If there was a need for it, static site generators could be scaled to tens or hundreds of thousands of pages and still finish in a few seconds. Caching partial results is always an option.
There are generators that are decently performant, like Hugo (Go) [1] or Zola (Rust) [2].
For example, I generate my blog from markdown with a custom generator written in not-very-optomised Python. There are about 400 pages currently and it only takes about 2 seconds on a 6 year old MacBook to completely build the entire site.
Even if I write another 1000 blog posts each year, generating the site will still take less than a minute for the foreseeable future. Paying the cost of building the pages up front like this is still much better than having my users pay it by waiting 10s of seconds from the page to slowly be dredged out of a database.
There are sites that do have tens of thousands of pages, and perhaps a content management system is appropriate for them. Horses for courses and all that.
Adding a memory cache on top of the filesystem cache is probably a waste of time for no gain (unless your workload has really specific pattern maybe…).
In our case, I started and successfully exited a crypto trading service (tradedash.io which was sold to Bittrex.com) that made billions of dollars in transactions using nothing more than 2x $5 VPS. Our entire tech stack costed us less than $75 dollars per month. The backend API handled massive amounts of trading data as you can imagine using those two servers and they never even got close to being at full utility - even when tens of thousands of orders were placed and cancelled at the same exact second. We had thousands upon thousands of users hammering our servers with requests every second of every day and the dropplets did just fine. At the end of the day, it all comes down to building out your infrastructure properly.
Clean your code base, use only what you need when you need it and make sure things are as fast and secure as possible. Dont over-complicate things for yourself and dont add abstraction layers just for the sake of it. That to me is the essence of good software and the main benefit of that is less overall headache as I dont have to go in and engineer a solution every time something breaks in some random library or abstraction layer. Scale will also be a fortunate side-effect of that as well.
Other than that, there were actually very few tech pieces involved. I am an extreme minimalist by nature so I kept things as simple as possible in order to minimize the surface and complexity of the code base.
Overall, it took us 9 Months in total to build the application from a prototype phase to a full blown trading system. We did it so fast because we kept things simple and low in complexity.
In that scenario, the majority of hits will happen in the <24-hour period when the link is on Page 1.
The rest of the month could see daily traffic in the hundreds, and it might still amount to 100k+ if added to the traffic obtained during the viral spike .
E.g. my personal finance traffic would be mostly on weekday evenings. These were usually 9-5 working people, so not a ton of traffic overnight, holidays. Weekdays had more traffic than weekends.
I switched to static some years ago because the VPS running Drupal kept crashing with 100k visitors per month.
A few questions:
1. Do you have any protection against DDoS? I currently implement my own custom solution, people seem to be transitioning towards Cloudflare but I see such a move as dangerous. What's your mitigation and/or opinion?
2. You mention the ease of spinning up a new server - I personally just run a bash script for this. How do you automate the transference of your URL over to a new IP?
3. Have you ever experienced slow down and if so, how did this affect your site?
4. What are the most resource intensive aspects of your website? (I.e. What is using your bandwidth/CPU/RAM?)
After giving up, tried to use cloudflare. A few clicks later, they just ate the entire attack for me. All on their free tier.
That said, I strongly suspect cloudflare probably works hand-in-hand with the NSA or some shit, and it's a dangerous trend. But I just don't think there exists a viable alternative, unless you're a multi-million dollar internet company
> machines efficient blocking rules, I was saturating their
> connections.
I have a slightly more complex "load-balancing" method, as long as the network bandwidth itself can hold out (hosted in the cloud) I've been able to deal with all attacks so far. My current method is to have a "whitelist" and "blacklist" look up table to quickly begin filtering incoming connections. If you end up on the blacklist (reasons include too many connections over a time period, bad requests, too many requests without authentication, etc, etc) then you instantly get killed - not a single byte gets sent. Being on the whitelist (reasons include being an already auth'd user, low traffic density, unique requests, etc, etc) then you shortcut some additional checks. Connections not in either list (potentially either attackers or new users) go through an additional per-generated (the check and answer is statically generated, so it's not added to connection overhead) check to "test for humanness". Both lists are decayed over time and connections on both lists can be re-assigned if they start behaving badly. The additional checks are mostly only activated under high load.
After that there is a resource management layer that is essentially "service temporarily unavailable" on anything too heavy (large file GET/POST, large database read/write, non-essential database writes are dropped, etc, etc).
This is all then run through a high-level network simulation to test where the bottlenecks will be. Each "release" goes through the test before making it into production.
I can't give too many details, but there is also a slightly malicious protection layer - if we detect somebody using a known/obvious bot (which is harmful) we have a few methods to crash them (based on known vulnerabilities). Some of these also act against aggressive search bots.
I have a beautiful graph sitting somewhere showing an attack ramping up and then mostly disappearing when the defense was triggered. It was quite worrying at the time because we just assumed our system went down or we had accidentally started blocking real users.
> That said, I strongly suspect cloudflare probably works
> hand-in-hand with the NSA or some shit, and it's a dangerous
> trend.
I also highly suspect this.
> But I just don't think there exists a viable alternative, unless
> you're a multi-million dollar internet company
I agree and that is a massive problem.
With a REAL ddos attack you will get so many incoming traffic that your bandwidth is saturated before any software will come in effect. A real ddos is a hardware problem, and it cant therefore be solved by software.
> described as protection against a simple DDOS attack. With
> a REAL ddos attack [..]
I'm not entirely sure what is really classified as a "simple" vs "real" DDoS attack.
> With a REAL ddos attack you will get so many incoming
> traffic that your bandwidth is saturated before any
> software will come in effect.
I specifically said: "as long as the network bandwidth itself can hold out". If the network is saturated, it's saturated - it's game over, for anybody. But you'll find that most services will die way before this limit is reached. On a WordPress site for example you'll usually find the database will give out long before the bandwidth is saturated.
Even with legitimate users coming through some great filter like CloudFlare, it doesn't prevent a hug-of-death if you don't have some "smart" application level handling of high numbers of users. Like Linux OOM, when you hit your limits, the only choice left is to try and quickly and intelligently start killing without affecting the well behaved.
> ping, UDP, and TCP traffic.
Please bare in mind that the description is heavily simplified. If you're talking about bandwidth, of course this is the limit.
> It's really only effective against http traffic.
How so?
> Which means his upstream or provider is probably going to
> cut him loose for the duration of the attack.
There's nothing we can do about that, but being in the cloud offers at least some protection. DDoS'ing some random Pi at home is a little different from DDoS'ing an AWS server.
> A single person cannot effectively block a concerted DDOS.
Of course not, but that doesn't mean you have to make it easy.
These numbers are all rough, but they are all for the absolute core/unflexible C-code that barely complies with the RFCs.
> expensive.
Agreed and a very good point. IPs that end up on the "blacklist" are rejected before any handshake occurs. The description I gave was really quite simplified.
Moving to cloudflare gives cloudflare veto power over any attempt to access your site.
> To make up evidence or contrive events so as to
> incriminate (a person) falsely.
I disagree that calling CloudFlare Nazis is framing them, as the purpose (assumedly) is not to incriminate. Criticizing them by likening their behaviour to that of Nazis or suggesting that the company is filled with fascists may be in reality false, but a valid criticism regardless.
Even if cloudflare doesn't misuse their power, something worse happens. Cloudflare becomes an easy OFF Switch for laws, warrants and copyright trolls to track people or kick both content and people off the internet.
I’m not disregarding the possibility or risk, but there are far greater problems that deserve our communities attention more than CF.
The numbers aren't meant to be brag or impress anyone. It's just an example showing real numbers based on my current site's traffic.
I wrote it because over the past year at least 40 people have asked me how I host my site and what the process looks like to build it. Now when I get future emails, tweets or Youtube comments I have a place to link folks instead of having to wing a half-assed individual response.
Also in a world where everyone thinks you need a multi-galaxy hyper inverted Kubernetes cluster with DeepMind-level AI to auto-scale to infinity I think it's a breath of fresh air to know you don't need to do that to host a lot of different types of content.
In a world where you are promoting 100k views for $5 a month perhaps the people asking you for advice will be given a chance to consider the faster alternatives for exactly $0
Basic off the shelf webhosting with a half dozen middlemen taking their cut can handle that perfectly well :/
Sorry to be blunt, but it's entirely true.
A modern NUC has 4 cores / 8 threads, an NVMe SSD and offers performance equivalent to a $300/month Azure/AWS VM. There are backup/availability/connection concerns, but in terms of performance bare metal on one's basement can't be beaten.
Internal scheduled backup to hard drives that will be physically disconnected most of the time.
Standby redundant synchronized server deployed somewhere else for peanuts. It is not as performant as said NUC but will do while the emergency is fixed on main one.
[1] https://discourse.gohugo.io/t/transition-2m-posts-from-wordp...
Their free tier covers 50GB of transfer and 2mm requests. If you need an extra 50GB as indicated elsewhere in the thread it would cost about $4.25 a month, less if it's not the full 50.
I'm looking to move off of a cheap shared hosting provider and on to AWS for my ten-accidental-hits-a-month personal (static) sites, it doesn't make sense to even consider a webserver for a solved problem like this.
If you need dynamic content served from a CMS, put your CMS's public API behind cloudfront and bust it when the owner makes changes.
For websites that don't get edited a lot, you can use lambda to store the website content as a JSON file which sits on S3 (somewhere like /content.json).
If you have multiple editors, use Dynamo or Firebase as the content store (again, behind cloudfront).
With every one of these options, the most expensive part would be the domain name.
Some interesting reading on it
For my use case it was the simplest (and cheapest) thing that could possibly work.
Of course if your baseline is a wordpress site where every page weighs 3 MB and takes 2 seconds to generate when there's no server load, then this might seem like black magic to you
To buy another 100 GB I would have to pay $20 per month (4x as expensive).
Also I don't feel comfortable basing my entire business on Netlify. My blog and course landing pages are how I earn a living. It's the same reason why I don't use GitHub pages. I just don't like the idea of having to adhere to their TOS (even if I'm not doing anything wrong) and also be limited to their free tier's traffic constraints.
I'm sure they are a good company and I wish nothing but the best for them but I just don't see them as a good fit for me given the above.
I am not sure. I never used their service. I got my usage numbers from DigitalOcean's dashboard.
But I recently looked at Netlify to host another static site and they mention that 100 GB limit on their pricing page.
At 100k page view per month, you'll need to serve a 100MB page to reach that limit...
I guess it's because I have a lot of photo gallery images of travel trips and they get caught up in image search results. I'm not sure why it's so high to be honest.
It is going to go up quite a lot too because now that I have a podcast, each episode is around ~50MB for the mp3. I might end up putting the mp3s on a CDN in the end. Probably DO Spaces.
Nothing surprising here.
But a great blog overall!
Even a raspberry Pi can serve 100k page views in a matter of minutes (seconds?). But in a world of serverless-cloud-bigdata developers seem to have forgotten just how powerful modern computers are, and how much we have already been doing for decades with less powerful hardware.
Just to give you an idea: in 1997 Rob Maldo grew Slashdot to 100k pageviews a day on a single server [0] [1].
[0] https://news.slashdot.org/story/17/10/03/2330258/slashdots-2... [1] https://en.wikipedia.org/wiki/Slashdot
i've used a $20 DigitalOcean in the past that used to serve 5 millions pageviews per month. but that number is not very impressive at all.
at peak, there were at most 2,500 people online (number via Google Analytics) and that's not a lot.
I know I could have done better with other solutions but most of the editors are familiar with WordPress so I had to keep using it and optimize the heck out of it.
I've been toying with loadtesting a WP site. Out of the box it can't handle much even with a beefy server.
Currently looking at Wagtail as a hopeful happy medium between WP and static