Scalable SQL: How do large-scale sites and applications remain SQL-based?
queue.acm.org
queue.acm.org
Perhaps other DBs have similar systems, I'm not sure. But this is definitely going to change that way we scale our SQL at our company.
http://blogs.msdn.com/b/cbiyikoglu/archive/2010/10/30/buildi...
Also, $30K is the retail price for the license, no one actually pays that, and if you can write a little bit of code you can easily use a SQL Std license for $2-5K.
If you plan to go further than that, you should have an exit NoSQL strategy ready.
Downtime is not tolerated - I've been surgically torn apart by boards of major listed corporations over small amounts of downtime (it was something of a pleasure to be so professionally abused, strangely).
For all this we have found the best solution to use a monolithic DB with SAN-based replication, C libraries to connect to the SQL db, and a process-based custom-built application server with an embedded (unfashionable) scripting language for ease.
None of these things are fashionable, but no-one's beaten our performance.
The trouble (from my perspective) with all these social sites is that downtime is tolerated (Facebook is always screwing up; I consider them a partial joke technically) and the data is easily distributed (think Google) because the demands on consistency are low.
I'm sure our time will be up one day, but you go trendy at your peril. I'm an advocate of functional programming languages, but it's not realistic to get hundreds of these guys to switch seamlessly and reliably to a new paradigm without serious pain. Ditto NoSQL.
I'm going to guess some home-rolled column store similar to KDB.
[EDIT] re: Oracle stuff I've not heard of - there is tons of Oracle scale/performance stuff I've never even had to think about. I've worked somewhere with KDB+ in place to handle streaming market price data and interviewed at another investment bank where it was used as well. So my guess is based more on what I've seen and read rather than knowing you couldn't build such a system on top of Oracle.
Personally I would love to see if Oracle can match these other systems out there, purely from an interest point of view!
No-one's even mentioned the DB, and it's a major DB system owned by a three-letter company...
Quit teasing Out with it. Sounds like fun with lots of real hackability and responsibility built in. Would love to hear more of what you can share.
You mention a SQL db - have you guys put something together of your own or are you using a pre-rolled db?
It is fun, can be a lot of pressure, and the sheer quantity of issues and the speed with which they must be dispatched means I always have to learn and I often screw up analysis. Humility and the ability to prod people who know stuff and grok what they say quickly is essential.
We use a major DB (not Oracle). Again, think unfashionable, think speedy.
Sorry, don't mean to tease - I just don't want people to be able to track me down.
The reason I ask is because I am a pretty highly skilled Oracle guy, and some of the numbers I have seen around financial systems just seem so fast and big volume transactions per second. Maybe I could get Oracle to do them, but then the zero downtime scares me somewhat as until very recently, keeping Oracle up all the time, even through application upgrades is fairly tricky!
Oracle, for example, to my understanding, is the standard database in use at Amazon.com for their services; there are alternatives for key services, but Oracle is pervasive (at least as of 2009, when I last had insight into this).
Another interesting rumour I recently read (ymmv, of course): Apple runs the iTunes Music Store on Oracle's Exadata. Source is from a fairly trusted source in the Oracle DBA community (Daniel Morgan): http://forums.oracle.com/forums/thread.jspa?messageID=425520...
Clearly there are limitations to a shared-disk RDBMS or even a shared-nothing RDBMS, but the tech sphere somehow seems to be stuck looking only at the open source implementations of these things as the only "reality". For a startup, that's likely true, but for a larger company, there's a very different risk profile when evaluating a SQL vs. a NoSQL solution.
Some tables are sharded, others aren't. Some tables are imported into hadoop for hive queries, others aren't. Backups are run using a fragile program called zrm and an army of outsourced DBAs to manage it.
The army of outsourced DBAs also deal with replication failures, figuring out why inexplicable deadlocks happen, tuning every inane innodb configurable, and "rebuilding" servers when mysql decides to slow itself down and yearns for a complete dump/reload of all tables.
Friends don't let friends use mysql.
You'd likely use SSDs first for indexes, then data. Transaction logs are usually a sequential workload and 7200 rpm drives often suffice.
[1] http://archives.postgresql.org/pgsql-performance/2011-03/msg...
SQL servers at this scale are so reliant on Memcache and other related solutions that it seems like a large omission.
On an unrelated note, does ACM really still use Cold Fusion? And people like to knock PHP.
Until someone builds a cloud hosting platform with RDBMS that can do the necessary magic bits, that is.
[1] Yeah, yeah ... MySQL and their commercial licensing BS. It's been awhile but if memory serves and things haven't changed, you can't host a commercial web site backed my MySQL without A) releasing all your source or B) buying a license. IANAL.
Doesn't NoSQL require the same, only it shifts from being the job of a db engineer/sysadmin to that of the developers?
The problem is that the people who actually handle this correctly charge so much that I've never even seen their products in live use. Meanwhile free SQL databases are still really bad at it. But choosing an API that deliberately restricts you to only the worst possible query plans is not going to be part of the solution.
There's no weekend project that's going to be "fully scalable". Scaling isn't just grabbing a NoSQL backend and declaring victory. A lot of scaling is figuring out how you're going to use your data and creating data structures that allow for fast, efficient, and scalable access. Many NoSQL data stores force one to think about how you're going to access the data more than SQL databases since the NoSQL stores usually don't support things like joins. One still has to do a lot of work to design to implement it right.