Getting to 1TB would certainly require 32GB sticks and would be the point where multiple machines gets cheaper.
107 karma · joined March 28, 2009
Getting to 1TB would certainly require 32GB sticks and would be the point where multiple machines gets cheaper.
Database Engineer
MailChimp is a DIY email-newsletter service based in Atlanta that serves more than 4 million users worldwide. We're self-funded and profitable, and we're growing fast.
Job Description MailChimp is seeking a database-focused infrastructure engineer to join our team. You’ll build, maintain, and monitor systems that support millions of users across all products and applications at the company. We operate in a high volume environment and are growing rapidly, adding 10000+ new accounts/day with that rate increasing every week. On the infrastructure side we do interesting work to support that scale and growth rate, working closely with developers to build and support our applications.
The engineering teams are small and focused, our benefits are unmatched, and internally we function much like a startup aside from established stability and abundant resources. There are no sales people, no investors, no board, and engineering teams are trusted to make good decisions with resources with minimal oversight. We are competitive nationally on salary and benefits.
Applicants should have strong Linux experience, extremely heavy operational MySQL experience, independent troubleshooting skills, and a love for automation and monitoring. Our core persistence layer consists of hundreds of horizontally sharded and paired Percona 5.5 instances. We are seeking an experienced engineer who would be part of the core infrastructure team with a focus on maintaining and improving this crucial MySQL layer in addition to contributing to the systems that feed off of it (emailgenome.org, Elasticsearch, Postgres, Redis). We take a pragmatic and practical approach to our stacks using proven components and building our own logic and complexity on top of well understood building blocks. This is a large, well built setup with consistent hardware, configs, monitoring, backups, and tuning in place.
Skills and Experience: - Linux servers - Experience with modern MySQL at high volume: sharding, replication, backups, monitoring, tuning - Knowledge of the current MySQL ecosystem and an interest in where things are headed - Deep understanding of InnoDB performance, recovery, tuning, and backups - Experience with replication and an understanding of all options and components involved - Exposure to HA solutions like MMM, haproxy, PXC - Scripting and coding (python, bash, php) - Puppet or other configuration management tool experience is a big plus - Zabbix or other monitoring tools - Familiarity with other infrastructure pieces of our stacks (nginx, apache, memcache, postgres, redis, elasticsearch) is also a big plus
Please apply here, the applications come directly to me:
http://mailchimp.theresumator.com/apply/oJIqi3/Database-Engi...
Many other positions available as well:
I am hiring for two roles on the Infrastructure side:
Systems/Coding: http://mailchimp.theresumator.com/apply/6Il9br/Infrastructur...
NOC/Networking: http://mailchimp.theresumator.com/apply/8ZRfKP/NOC-Engineer....
MailChimp is a unique place. We have ~3.5mm users, send ~6bn emails/month, and sit at ~80k queries/second hitting our dozens of database shards during a typical day. We are growing rapidly, adding 7000+ new users/day with that rate increasing every week. On the infrastructure side we do some neat stuff to support that scale and growth rate, working closely with developers to build and support our applications.
The engineering teams are still small, our benefits are unmatched, and internally we function much like a startup aside from established stability and abundant resources. There are no sales people, no investors, no board, no phones, no useless meetings, and engineering teams are trusted to make good decisions with resources with minimal oversight.
If interested, use the links above to apply. Your information will come directly to me.
Job Description
MailChimp is looking for engineers to join our team. This is a full time position in Atlanta that will help build, support, and monitor the infrastructure our company depends on. We handle tremendous volume and support millions of users that love our products.
We are looking for people with independent troubleshooting skills, strong experience with Linux, and a desire to monitor and automate everything.
Skills & Requirements
Linux experience, especially at higher server counts Scripting and coding (bash, python, ruby) Familiarity with pieces of our primary stack (nginx, apache, php, memcache, mysql) Experience building high volume systems is a big plus Strong experience with mysql is a huge plus (sharding, replication, HA)
About MailChimp
MailChimp is a self-funded and profitable Atlanta-based company that is growing fast. We offer competitive salaries, exceptional benefits and perks, phone plan coverage, coffee, snacks, top tier equipment, and an environment that empowers engineers to have a big impact. We work in small teams, there are no project managers, no product managers, and engineers are trusted to work autonomously and make good decisions.
Email resumes to: infrastructurejob@mailchimp.com
MailChimp is looking for infrastructure engineers to join our team. This is a full time position in Atlanta that will help build, support, and monitor the infrastructure our company depends on. We handle tremendous volume and support millions of users that love our products.
We are looking for people with independent troubleshooting skills, strong experience with Linux, and a desire to monitor and automate everything.
- Linux experience, especially at higher server counts
- Scripting and coding
- Familiarity with pieces of our primary stack (nginx, apache, php, memcache, mysql)
- Experience building high volume systems is a big plus
- Strong experience with mysql is a huge plus (sharding, replication, HA)
MailChimp is a self-funded and profitable Atlanta-based company that is growing fast. We offer competitive salaries, exceptional benefits and perks, phone plan coverage, coffee, snacks, top tier equipment, and an environment that empowers engineers to have a big impact. We work in small teams, there are no project managers, no product managers, and engineers are trusted to work autonomously and make good decisions.
Also, in addition to the above, I am looking for somebody with tremendous networking and colo experience.
You can email me directly at infrastructurejob@mailchimp.com
I am looking for Infrastructure Engineers to join the team. We support hundreds of servers, millions of customers, and send billions of emails every month with a small team that prefers automation over manpower.
MailChimp offers extremely competitive pay, unmatched benefits, and a culture that empowers engineers to work autonomously with large budgets and significant resources. We use top of the line equipment to support impressive volume in an international, 24/7 environment.
I am looking for two types right now. Generalists or somebody that can hit both of these are especially welcome:
- Devops, server guys that can write code and contribute to our automation tools, expertise with databases is a large plus given those are our largest machine type.
- Network Engineers, people who absolutely understand and love working with high end networking gear, setting up colocation environments, etc.
We will cover relocation expenses completely for the right candidate and can offer compensation appropriate for any level of experience.
If any interest email infrastructurejob@mailchimp.com and it will come directly to me.
It lives on that requirement and is incredibly stable and fast. I don't want another half-baked data store that technically "works" on disk but only as long as your volume is completely trivial. There are plenty of those if capacity is a dominating concern.
In Softlayer's Dallas facility (selected when ordering a server) 10 Gbps is an option.
The maximum amount of bandwidth available is 20TB.
Sometimes you can get more options by contacting their sales people directly instead of using the shopping cart.
In certain datacenters (think might be dallas05?) you can even get a 10Gbps uplink brought to a box so they have pretty big pipes available.
It is expensive, but not as expensive as something like heroku or aws for equivalent cores/RAM, and you get real hard drives.
Thankfully most of the joins happen within a shard (hashing and sharding on something like a user_id) with the exception being various analysis and aggregation queries.
Using PostgreSQL's schemas is admittedly not too different from just using many DBs in MySQL or something else but in practice I've found that extra layer of organization helps keep things neater. I can backup, move, or delete a specific schema/shard or I can backup, move, etc all shards on a machine by operating on the containing database.
It's setup for multi box (each schema is mapped to a hostname in code) but I simply haven't had a reason to move to more boxes yet. The schema feature is a nice, convenient way to pre-shard like this so that growing to more boxes doesn't require rehashing for a very long time if ever (depending on how much sharding you do up front). You just move schemas/shards as needed using the standard dump and restore tools and update the schema->hostname mapping in the code.
So many of these other new data stores promise "infinite scale" and leave out the "as long as it fits in RAM" part.
Though I tend to prefer Postgres don't get me wrong - MySQL has some advantages.
For raw pkey lookups, especially range selects against pkeys with lower amounts of concurrency InnoDB's speed is unmatched. It should be given that it is completely laid out on disk specifically for that to the detriment of other features.
MySQL's replication is extremely sturdy. On paper the older method (statement-based) sounds incredibly fragile but in practice it works extremely well and has been proven on countless projects. mmm makes it almost too easy to setup replication and failover.
I feel that MySQL stalled out for several years in the 5.0-5.1 period where poorly engineered features were bolted on to create a product with too many pathological cases that destroyed performance to keep track of. All these features were available but experience taught you to avoid most joins, avoid most subselects, avoid most usages of views, etc.
That said, v5.5 has a lot of nice improvements and can actually scale up to more than 8 cores (I've gotten linear improvement up to about 32 cores and that is what Oracle puts in their white papers as well). Percona and Facebook are releasing nice patches and branches, and forks like Drizzle are reaching GA and doing good things as well. So I think it is headed in the right direction again.
Compared to InnoDB? You are incorrect, Postgres is faster, even for a raw count(*), even when comparing against the InnoDB plugin and not the ancient InnoDB builtin.
I have access to tuned TB+ DBs of both types and am happy to disprove any specific examples you can provide.
For quick and dirty map reduce on a smaller node count I've started to really like Disco (discoproject.org). You just pull down the backend with your package manager, push your files into ddfs, write a python script, and run it.
Usually running raid10 for performance reasons but when I don't need fast writes I use raid6 instead of raid5 for anything beyond 8 drives.
I can order machines online and SSH in 3-4 hours later. Even exotic stuff they turn around just as fast - we saw that speed on a quad octocore box with a raid 10 of Intel SSDs.
That's real metal too, with real IO (most of my work is IO bound so VMs and the cloud are not options). You get to pick the exact CPUs, disks, etc and they slot them in solid Super Micro boards and use good Adaptec disk controllers. You pay monthly and can spin down the box at any time (though must pay full months, no per-minute pricing like AWS).
That is on the dedicated hardware side, you can also spin up compute instances and those can be cloned and fired up in bulk. But, they also have the IO problems that all other VMs have.
In any case, just wanted to mention they are a decent middle ground. Not as automated and polished as Amazon on the VM side but you can spin up mixtures of metal and VMs to get combinations that make sense - pushing compute or RAM-only stuff to VMs and keeping DBs and persistence layers on real metal. They have a few different datacenters too so you can spread gear around physical locations.
http://blog.mailchimp.com/mailchimp-launches-transactional-e...
It rolled out this very week.
See (http://blog.gtuhl.com/2009/03/26/ocz-vertex-ssd-in-a-17-sant...) for install and before/after xbench runs.
The current 8.3 documentation is here: http://www.postgresql.org/docs/8.3/static/
I especially enjoy the sections on indexes: http://www.postgresql.org/docs/8.3/static/indexes.html
If you have 30 million rows in a table you absolutely cannot do joins so you bring everything in that you commonly need.
Having a single awkward example does not detract from that.
I'll add to this by noting the post does say "For smaller tables it doesn’t matter but as tables get bigger avoid joining when you can" which is sound advice.
I'd say joining scales up to perhaps a million or two rows unless you have a lot of RAM (you can see join spills to disk in the EXPLAIN ANALYZE output). I often am working with tables in the 10-30 million row range so my perspective is probably a little slanted towards the negative.
We peak at 1000s of transactions a second on an OLTP database that is over 100GB in size and PostgreSQL handles it like a champ.
All of those deficiencies may have since been eliminated.