Numbers everyone should know
highscalability.com
highscalability.com
Using a register is like interacting with information in your own brain, using an L1 cache is like walking to a bookshelf in the room you're in, using an L2 cache is like walking across campus, using RAM is like going to Sacramento (we were in Berkeley), hitting the disk is like going to Pluto, and using tape drives is like going to Andromeda.
I never fully registered the consequences of hitting disk until then.
Core dump and flush need no explanation :)
This is a large reason I dislike cloud hosting, most of the hardware is crap and you could get much better performance from it with a little more expense. Also it's mostly pointless to bother making your writes sequential as the performance degrades to random IO because of all the other instances on the machine.
http://ee380.stanford.edu/cgi-bin/videologger.php?target=101...
The slides for the talk: http://www.stanford.edu/class/ee380/Abstracts/101110-slides....
SSDs about twice the price of 15k rpm drives, hardly out of reach for most. Also for VPSs, server axis 20mb unmetered SSDs are cheaper than slicehost's rotating disk based systems. http://serveraxis.com/vps-ssd.php
[0]: https://github.com/facebook/flashcache [1]: http://blogs.sun.com/brendan/entry/test
I think people really overestimate their data-store needs, and underestimate how much a big iron DB server can deliver.
Its probably worth thinking about scaling out on your data end, but until you start hitting some sort of limit on a 8 core 96 gig machine with server level SSD's its probably worth investing your time in other issues.
But, if you've built a site on the precondition that one database is enough, with a lot of joins so you can't split tables, and your hardware can't keep up.. you're in a really bad place. I say this out of experience ;)
Of course over a certain scale this is normal, so im not interested in a site with 50 millions users, but cases where a single DB is not the best solution.
Since these queries were built over ~10 years with the expectation of being able to join any of the thousands of tables, it would now be a tremendous undertaking to perform any sort of sharding or other ways of distributing load (other than master-slave replication, which was used widely).
It's been a struggle to keep hardware up to pace with the desired growth of the company. Even with a dedicated data center and top of the line hardware, doubling their user base would probably require a significant architectural undertaking.
Had the site been designed to support an arbitrary amount of distributed database servers from day one, it would now be trivial to grow horizontally.
There's a legitimate rant to be had about using elephant architecture to serve mouse traffic, and I think that's what annoys me most about the NoSQL fad. But that doesn't mean the problems mentioned in this article don't exist.
I'm not sure I understand. If I make an application that has to increment a counter, shouldn't I always be worried about concurrency?
What I mean is, sure, if I'm serving only a few requests, then of course the probability of running into concurrency issues is lower than a site with more requests/sec. But it's still a matter of chance, there is some probability that two requests will come in at exactly the same time and cause a problem.
The parent poster is worrying about performance of concurrency. This doesn't need to be worried about unless you are one of about 5-10 tech companies whose name is recognisable to people on the street.
When it comes to parallel operations over multiple machines ( most startups at scale will run into this problem ), these numbers really can come in handy.
If you're complaining about the Hacker News title, then:
a) It's the linked article title, so it would be inappropriate to change it, and
b) HN shows the domain name of the link in parentheses after the title. The article was pretty much in line with my expectations for a High Scalability article with that title.
Knowledge, properly applied, can be power.
Indeed, every programmer should know those numbers.
key = timestamp + UUID4
TL;DR: This would actually work really well I think!