Being on HackerNews, One week on
sharelatex.com
sharelatex.com
Let's say that the peak is 100 times more frequent than the average, so that would be about 10 visitors/second.
Being a systems developer, and not a web programmer, I work on software that handles tens of thousands of requests/sec. I don't understand why around 10 visitors/second would be a difficulty with semi-modern hardware. How many requests does that translate to?
This is a genuine question, can anyone explain what is it that takes so much work in handling a web request?
If you're willing to throw either time or money at the problem then it goes away easily. But both resources are typically limited.
Languages like PHP follow the same model, and as a result, every single page request is processed independently. All the raw data is retrieved from the database, is processed appropriately for output (e.g. turning content into HTML), is run through a templating engine, and assembled with the right CSS and JS so it can be served. This is attractive from a rapid development point of view, because you can deploy changes instantly and can scale it out horizontally just by adding more servers, without any additional work.
However from an efficiency standpoint this sucks, and this is why the most common fix is to place a static HTML cache in front of it (e.g. Varnish) as well as opcode caches, object caches, etc. This only works if all your visitors see the exact same thing (e.g. a HackerNews discussion thread). If you use 'write through caching', then you can control the rate of updates independently of the amount of traffic you receive, and you can handle traffic pretty well.
If your pages are dynamic, you need a different approach. You'll want to cache all the static chunks of each page, and assemble them together with the dynamic parts on-demand. The extreme example is Facebook: everyone sees something different. The only way to scale this out is to parallelize everything, with your first tier of web machines making many simultaneous requests to a farm of servers behind them, delivering all the pieces within a relatively constant time.
The problem is that such a parallel architecture is both unnecessary for a small web app, as well as involves a leap in complexity and know-how that is undesirable for small teams. Hence, there is an increasing technological gap between what hobbyists/start-ups do, and what the giants are doing.
Edit: it's also important to realize that the web loves 'inefficient' dynamic languages not because they're dumb, but rather because development is very rapid, very experimental, involves designers, UX experts and marketers, and you don't want to be forced to make long-lasting decisions early in your development process.
I would say a big lesson people should learn from you and others is, use load balancers. At least if traffic spikes, you can add another server (provided you have an image sitting by waiting and your code doesnt mind being load balanced).