Scaling a startup from 0 to 40 hits per second in 3 days
markmaunder.com
markmaunder.com
Early on, like in your situation, I did go to the file system at first because I couldn't "make it work," but then eventually I went back to postgres after I figured out its scalability details. If you have a lot of little files and you hit them a lot, that will eventually probably become your bottleneck. Did you figure out the bottleneck in your db setup or has there just not been enough time yet? Just curious.
I'm using Apache::DBI, caching everything via startup.pl on load, keep_alive is on but with a very small timeout, host lookups are off.
Also I'm using the worker MPM in apache2, just FYI. I've found it to be really memory efficient.
I've been using MySQL with Apache::DBI for years and it's usually brilliant - I ran WorkZoo.com, a high traffic job search engine with a combination of MySQL and a full-text api.
With feedjit I'm basically storing weblogs. I either have to dump them into a single table and query that - which is what I was doing and the high query rate with read/write was a problem - or have lots of individual tables which isn't feasible after about 500 with MySQL. So small files works best for me.
Also, I take it from your comment that the bottleneck was in MySQL doing the writes. I assume the read side is indexed appropriately so MySQL finds the right part on the disk almost instantly. Do you think it is a locking issue then, e.g. table lock vs row lock? (Forgive me, I haven't used MySQL in a while.)
You can improve things a bit by using INSERT DELAYED. When you use that, mysql doesn't guarantee that it'll insert the row immediately, but the mysql query returns immediately when you do the insert (it doesn't block) and mysql queues up inserts and inserts them in bulk when it feels like it. The non-blocking and bulk inserts that INSERT DELAYED give you speed things up, but only to a point because you're still constantly rebuilding an index on a table that's getting a lot of reads.
Mark.
I'm about to upgrade though - mostly because I need more RAM to support more concurrent connections.
I have a new server ordered which comes online today and that's just for feedjit, so I can compile a feedjit-only apache and it'll be able to handle many more connections.
Mark.
Mark.
Question 2. As a web developer, should I be worried that I don't know as much as you about scaling a web app? What will I do when my web app makes it big?
Mark.
So you want to lower it basically if you have a lot of different clients coming at the same time. On the other hand, if you have a few clients and they each make a lot of requests in a short amount of time, KeepAlive can save some time and resources.
In answer to your other question, I wouldn't worry too much about this stuff until you need to. When that happens, there are a lot of good resources on the Web, and you can post here and I'm sure people will help you. If not, contact me, and I'll do the best I can :)
Here are some further explanations about KeepAlive: http://virtualthreads.blogspot.com/2006/01/tuning-apache-par... and http://www.perlcode.org/tutorials/apache/tuning.html
The down-side is that browsers need top open a new connection for every page component they load. But things will load more reliably now and I can handle more traffic.
I have a new server ordered which comes online today with lots more memory. When I deploy that I'll probably try turning it back on to make loading faster for my users. Then when things get hairy again I'll turn it off. Come to think of it it'd be nice if I could do that without rebooting apache based on server load.
Mark.
feedjit only shows the most recent 10 referrers and clicks, so I don't think theres anything there that'll give a 'competitor' some sort of strategic advantage. Besides, tehy can just google around and find out what I'm showing up for in the SERPS.
As far as privacy goes, as long as I'm not personally identifying people and showing what search terms they're using, I think there aren't any privacy issues. I see my own search terms showing up and my location 'Seattle, WA' and no one knows who actually searched that term.
I haven't really applied my mind to this as much as I should, but those are my initial thoughts and I'd love to hear if anyone feels different.
Mark.
On the other hand, why should anyone expect their searches to be private at all? If I wanted to share searches coming to my website and even group them by IP address, what's stopping me? What if I said I would do this in my privacy policy?
BTW, I've been thinking of a Feedjit type thing but for search keywords instead of location. That's what got me wondering about these privacy issues. I hope that idea wasn't in your future plans ...
Mark.