I think scalability and reliability are mostly orthogonal issues and
should be discussed separately.
Scalability is a farce. In my experience applications are implemented
in such hilariously inefficient ways, because it's hard to see where
the time and especially the IO is spent. But adding more machines is
not a good solution, because it becomes even more difficult to
understand the performance behavior of the system. Also, it transforms
the system into a distributed system and therefore adds a whole new
level of complexity and failure causes. Most websites don't have the
some problems as Google and Amazon.
Optimizing the application could prevent the need for scaling out to
multiple machines. Optimization is not rocket science: Find out the
bottleneck and fix it. Often it is random IO, so either load the data
into RAM or change the algorithm to use sequential IO (e.g. through
batch computation).
Reliability means that the system must work correctly all the time and
there is no way to fix it (e.g. a satellite on a mission). Thus, for
most applications availability, which is the fraction of time a system
works, is a more appropriate metric.
There are two basic strategies for achieving high availability in a
system: perfection and fault-tolerance. Perfection is the default
programming model, which assumes that every hardware and software
component of the system works as expected. Fault-tolerant systems, on
the other hand, are hard and require that the system is designed
around this idea. Also, fault-tolerance makes it harder to change the
system. An often overlooked fact is that for many systems it is
actually much easier to achieve a particular availability goal through
perfection rather than trying to build a more complex fault-tolerant
system.