Scaling PHP Book: I will teach you to scale PHP to millions of users
scalingphpbook.com
scalingphpbook.com
1. Cache the output at the edges: Use Varnish or other reverse proxy cache.
2. Cache byte code: Use APC or XCache PHP opcode cache.
3. Cache and minimize database I/O: reduce database touches using memcached, redis, file caches, and application-level caches (ie. global vars)
4. Do event logging in local files, not to the database: Make all write operations as simple and fast as possible, any data that is not needed in realtime can be written to a plain old file and processed later.
5. Use a CDN, especially for delivering static assets
6. Server tuning: Apache, MySQL, and Linux have lots of settings that affect performance, especially the timeout settings ought to be turned down.
7. Identify bottlenecks: At the system level use tools such strace, top, iostat, vmstat, and query logging to see which layer is using the most time and resources. Also there's an excellent PHP code profiling service called New Relic that drops you right into the function and db query that's eating up most of the time in slow requests.
8. Load testing: DoS yourself. Stress test your stack to find bottlenecks and tune them out
9. Remove unused modules: For each component in the stack unload any default modules that are not needed to deliver your service.
10. Don't use ORMs and other dummy abstractions: Take off the training wheels and write your own queries.
11. Make the entry pages fast, simple, and cacheable. Nobody is reading that silly news feed in bottom corner of your front page and it's killing your database, so take it out.
Most of the time a PHP slows down because each PHP process is blocked waiting for I/O from some other layer, either a slow disk, or overloaded database, or hung memcached process, or slow REST API call to a 3rd party service ... often just strace'ing a live PHP process will show you what its waiting for ... in short, blocking I/O slows down everything. The key to going faster is:
* keep it simple
* cache as much as possible in local memory
* do as few blocking I/O operations as possible per request
Having said that, I'm looking forward to the book, as there doesn't seem to have been much on this topic since the O'Reilly books on scaling web applications (2006 I think).
I want to know about the weird things that happen when you push your PHP stack to really ridiculous limits. I think that's what this book will offer. If it ends up being stuff like "use APC! minimize trips to the DB!" I would be very surprised and disappointed.
On a small system, a standard HTTP request through Mongrel2, to the framework and back is about 2ms (including 1ms of network latency between the hosts, data from my production monitoring). This is the "hello world" latency. It makes using PHP very fun and very fast. Also, a single PHP process running in the background uses about 12MB, it means that with 12MBx3 + 10MB Mongrel2, you can serve 1000's of users with nearly no load on your system (if you push the load only when needed at the DB level). Tracing a PHP process, you can run it without system calls at all (only gettimeoftheday for logging).
But, this is cutting edge, so, you need to be open minded and ready to dig in the code (less than 25k SLOC anyway I think).
Is this a language specific suggestion or something broader? I would love to hear a deeper debate on the subject.
There was a debate on HN about ORM's a while ago, and one of the most significant comments I read was something along of the lines of "Every project I've ever worked on that 'didn't use an ORM' implemented the same functionality of one in a less-maintainable way."
* Or other object-oriented abstraction layer for database access, ORMs aren't the only option.
Is there some extra overhead? Of course. Probably a lot less than cleaning up the damage from some cowboy coding.
What do you mean by this? That you should consider a non-relational database? Clearly any mapping from an OOP language to a relational database is an object-relational mapping.
I'd argue that, in some cases, an ORM is only used as a reaction to the separation of concerns. Writing SQL mixes up your languages, so why not try to build that SQL in the language you're working with? Or maybe SQL is considered too difficult.
I quite like having full control over a 'raw' query without requiring an ORM's opinion on how it should be built.
As the parent said, ultimately cpu cycle on your web servers don't matter much, you will be limited by IO everywhere. Knowing how to properly use redis' data types and handle multi-level cache invalidation will be a much more valuable asset than using a framework you might not enjoy the most / be the most fluent with just because it's a wee bit faster.
Even in terms of actual php's calculation, your main strong point is figuring out what you can remove from the page processing and put in a daemon instead. "If the user wants to do action A, what is the minimal amount of things I can do to make him believe that it has been done", do that in your controllers, and put all the actual hard work in an event processing queue that doesn't have an user waiting for it to answer back.
Yesterday I had a meeting with a potential customer and I hated it. I hate to try to explain my SaaS software to non-technical people who treat me like some 17-year-old webmaster. I'd much rather be refactoring Clojure code. But I got out of my comfort zone and this client will probably add hundreds of thousands to my bottom line. And I'm glad I was at that meeting while my competitors were fetishizing about non-existent scaling issues.
It's 2012 for god's sake, you can rent a 32GB server for less than $100.
I always thought I was being wise by not doing any premature optimization, but after a few lessons learned the hard way I certainly factor in performance to the design of software before I build now.
Scalability is not a "feature" tacked on at the end development.
What percentage of startups on HN have achieved product/market fit, are past a 128GB commodity box AND have no dedicated engineering team for scaling issues?
Once you get to the point of no return you have to know what to do or you'll suffer. Learning that when fire is falling from the skies is the worst way in retrospect.
Plus, almost all of these techniques will also speed up page generation on a lightly-loaded server, so that's a win in any circumstance.
In all of my cases dealing with scale issues it's been due to the DB growing very large in size to the point where you can't use the same techniques that you're used to using. That usually happens in combination with more concurrent users. Depending on how large and how organized your code is, things can get very ugly when the site starts throwing errors, being unresponsive, etc.
I usually know where I want to go as a next step with scaling and monitor our resource usage until it gets pretty high. Then we take the next step. But we don't spend a huge amount of time scaling until we know we need it.
I went back and forth with them about optimizing the app, but it was apparently some huge labyrinthine monstrosity. They insisted that they didn't have the resources to do any of the significant rewrites that it would require to fix the app to do proper queries (or at least, not enough to be worthwhile).
Eventually I gave up. /tmp was mounted onto a separate partition, so I disabled ext3 journalling and set commit=30 so that it only sync'ed to the disk every 30 seconds. Since no temporary tables lasted that long, the VFS layer never wrote to the disk if it didn't have to. /tmp became an in-memory cache, and CPU use dropped to 5%.
Optimizing isn't about a checklist, it's about looking at the system that you have, understanding what it's doing and why, and understanding how the other systems around it behave so that you can resolve the issue. Moving onto another database server wouldn't have helped them. Moving onto a RAID would have reduced the impact, but their load didn't scale linearly so they'd hit their limit in a few months anyway.
Where, in the US, can I rent a 32GB server for less than $100?
[1] http://www.hetzner.de/en/hosting/produkte_rootserver/ex4s
[2] http://www.honelive.com/xml/#new-york-city-dedicated-servers
It's 2012 for god's sake, and 32GB RAM servers are still hundreds of dollars a month to rent from any first class data center.
I'm really curious to see how different equivalently priced ASUS and supermicro boards differ.
On the downside, I have a book cooking on the same exact subject, with release planned September (self-publish, about 80% done), but now not sure if it's worth continuing. [edit]Slight moment of "panic", as seeing someone else releasing a book on the same subject made me sad, but you're right, no reason not to continue.
Absolutely finish and release your book. Half-writing a book is an even worse decision than writing a book. (Tongue only somewhat in cheek there: if you have actionable information on how to scale PHP, a book is one of the worse ways to change people's lives with that information while simultaneously making money from it.)
Absolutely. And it is an important enough topic to warrant the purchase of two books from my perspective.
Good Luck!
Go with http://leanpub.com - publish it and then complete it.
Either way, I would like to see images of what the product will look like. For someone like me who doesn't need this book but is still interested in its topic, images of a well-designed book might make the difference in whether or not I purchase it.
It will be self-published, DRM-free in PDF, mobi, epub. Looking into what it takes to publish on the Amazon Store/Kindle/iBooks, but hopefully that's something I can figure out after launching.
However, it looks like your mailchimp account is setup to link to phpscalingbook.com (which is the wrong domain). I clicked on the "continue to website" link after confirming.
Also, I noticed that when I clicked the return to website after clicking on the subscribe button your website, it went to google.com instead, and when I clicked continue to our website button after clicking the email confirmation link, it went to http://www.phpscalingbook.com/ (404 error).
And if you need any proofreaders, I'd be more than happy to help!
So, USE AN ORM!!!
By definition, an abstraction is more restrictive ... otherwise, you're not really abstracting anything.
If you're making a simple web app, an ORM is just fine. If you're building enterprise-level software, it is a really really bad idea. Half your code will be using it, the other half will be forced to use half-assed SQL queries that try to fit into your ORM. Because, quite simply, you will need all that SQL has to offer to make things work right. You can't afford to abstract SQL away. Trust me, I tried. In the end, the best you can do is go with something like LINQ.
I built a LINQ like system for PHP a while before Microsoft did it for .NET =)
On a side note, I would also stay away from all the frameworks and build one yourself. If you're on a long-term project, it's worth it. You'll understand what is happening and what each call really costs you. You can also refactor an existing framework. Either way works.
IMO, "ORM produces bad SQL" is a myth and more often a case of bad indexes and not bad SQL.
Would love to see a chapter list!
Would be nice to see a preview of different bits of the book when written as well.
I'm glad you are writing the book. It seems like a worthy addition to my bookshelf. Will you be blogging about specifics mentioned in the book?
404 Not Found
Code: NoSuchBucket Message: The specified bucket does not exist BucketName: www.scalingphpbook.com RequestId: 403C0590E064E19F HostId: a+UggS1lMBgPJrT5X/kbdzsRK1kx+iKBQw6u4dZxieNkspwHbLZBWzXMa9CiHEAu
Scaling problems will usually return 5xx errors, most commonly 502, 503 and 504.
On a side note: What is the icon for "Happy Users" supposed to be? It looks like Facebook's like icon but without a thumb.