How StackOverflow Scales with SQL Server
brentozar.com
brentozar.com
- Everybody's the DBA
- Do what it takes to get what you want
- Tune later, cache & separate now
- NewEgg your way out of problems
- Share for great good
For each point, he argues why the old rule no longer applies and what the new solution is.I felt a lot of the presentation was about tuning SQL Server without tuning SQL Server: caching, leave full-text searching to Apache Lucene (because it's not querying), and using SSDs to speed up performance without having to touch any code.
I havent had a chance to watch the video, but i hope to later on. What is interesting, even with just the link is that for a website like StackOverflow that MS SQL is a viable solution.
We have been using SQL Server in our own company for our projects and i was really starting to get annoyed with it. I found it to be heavy on resources, slow to respond and lets not forget cost. I just completed a project that i have been working on for the last 4 months, and the majority of the work was within SQL server. One thing i learned was that its actually quite a powerful beast.
When used correctly and in the right way SQL is a very capable SQL solution. I'm glad that we decided to stick to SQL Server. There is a lot i learned about SQL Server in the last 4 months that i had no idea it was capable of.
It only seems like it isn't when you read sites like Hacker News, where most of the posters are not working in big environments. Facebook is a large scale solution that doesn't use either, but their data management and caching is so bad nobody should be considering them as a best practice.
It's not about being big, banks and insurance companies simply have different requirements on how to access and store their data compared to most websites/services.
http://www.google.com/intl/en/jobs/uslocations/mountain-view...
Google does use MySQL for their AdSense/AdWords transactional system (ie, buying AdWords).
If you took all of the votes from the 2000 election in the United States (I'll save you the search - it was just over 100 million), it wouldn't even equal one day's ATM transactions in the U.S., and those transactions are a nonevent.
Twitter in not only big in terms of traffic. 200 million tweets per day is not small, even if a tweet involves less processing than an ATM transaction. You can also consider Facebook as an example of website dealing with big data.
- Oracle and SQL Server: chosen by us
- Sybase and Informix: chosen implicitly by choosing applications that preferred them[1]
- DB2: Chosen by a company that we acquired
Oh and some MySQL and SQLite too.[1] Yes they ran on Oracle too - but when 98% of a vendor's customers are on Informix or whatever, sometimes it's just less painful to go with the flow, so you have access to that community.
http://windowsteamblog.com/windows_live/b/windowslive/archiv...
What are you comparing it to?
Admittedly I am biased since my day job is a SQL Server DBA, but I have tried several other options and think SQL Server tends to stack up quite well.
I have found it generally more user friendly, easier to work with, and cheaper than Oracle (though Oracle does seem to have an advantage in certain types of partitioning). I rather like MySql for certain types of projects, but generally find SQL Server easier to maintain for large projects.
I have only dabbled with NoSQL options, but my general opinion is that for certain problem sets, they are great. However, when ACID is even remotely desireable they are not an option and for certain other tasks they are less desirable.
So, I think that which type of database you use depends largely on the project, but SQL Server tends to stand up quite well for a wide array of projects.
Ultimately, i agree with what you are saying and my opinion of SQL server has improved in the last few months. I really like its flexibility.
The right setup (IMO) of having a DBA in an operational role with developers that are highly proficient/self-sufficient is hard to get right and expensive enough that it probably isn't right for an early stage company. And a bad DBA can be a nightmare. So there are tradeoffs on both sides.
Which has another advantage over Amazon's pager policy: if you are up at 3am, only semiconscious, you are likely to be more useful fixing code that you broke (because you just touched in the last day or two) than some other coworker's mistake.
i.e. the real rule is "nothing happens to prod without the agreement of those at the sharp end".
That allows good devops fun to happen with fewer ops bottlenecks, if some devs can provide enough ops-chops to keep the real ops people happy.
[Also note in passing that such devs are likely to be in the "second wave" of 3am people if things go sufficiently pearish.]
But... if you don't have someone dedicated to thinking about database issues, you need to treat database changes just like your code. It needs to be in a repository, it needs to be reviewed, and you need a change management regime.
From a anecdotal POV, I've noticed that many folks have a good process (or at least a consensus approach) to managing their code... but the database is often a red-headed stepchild that doesn't get the attention it deserves.
I absolutely despise writing liquibase changesets, but I know how important they are.
On my last three major projects, I have committed to devoting the proper level of attention to the database, with automatically building databases in some environments, scripted scheme changes and seed data loading as part of the mainline code base, etc. and I have found the difference to be immense in practical terms. I can move more quickly, more safely, and have a better quality of life as a developer.
One of the best returns on (effort) invested I have ever seen.
Anyway, I'm splitting hairs. Your presentation was great, and the people at SO are very impressive overall. Thanks for posting it.
I wouldn't say that developers need to know about HA/DR, replication, memory tuning, SAN setups, etc, though. Just the how-queries-work and how-data-is-stored angles.
Thanks for the compliments! I had fun doing that one.
I worry that if they went hand-rolled assembly, everyone would be jumping in that bandwagon.
The poster asked for similar resources on scaling Rails up; maybe something like https://github.com/blog/117-scaling-lesson-23742 or http://axonflux.com/building-and-scaling-a-startup would have been useful.
I sincerely doubt that the top-poster has Twitter-level problems or resources. He may have a Rails site that's crapping out under load, and needs resources for optimizing both application performance and database access.
"Just move everything to Java" is a meaningless answer. It doesn't help him at all, not unless he's got a year and change to retrain/rehire his engineers and completely rebuild his application.
Of course; this:
Use Something Else.
For most companies, diving deep into your data persistence layer is probably not worth it in the beginning. Your devs need to work on features. (MS SQL does pretty well untuned.) This is the economics part.
The computer science part comes in when you are doing enough traffic that 200 hours of dev time toward a 10% performance improvement becomes good economics. Then you dig into the data layer (and every other layer) and start counting milliseconds.
Which is what we did at Stack O by using Dapper and renting Brent Ozar. :)
This is horrifying. Are there really people that think this way?