158 karma · joined June 7, 2010
http://www.yottaa.com/url/skysheet-com-4d2ca049038ade0c05000...
Yottaa notes a similar reachability time as stella does in the "Reachability > Washington" metric. But response time performance is significantly worse at other locations.
Also, it's important to track not just response time of the server, but actual browser performance. If you look at the Page Load Time metric, you'll see that even when the server responds pretty quickly, the actual browser experience is significantly worse than a simple HTTP Client.
Overall, I like the site a lot. If you could get some more accurate metrics, it'd be great. Keep going!
BTW> Here at Yottaa we're beta-ing an API that would let you run tests and access all of our browser and low level data. Would you be interested in getting some better data from us? Contact me : jrosoff AT yottaa DOT com
- I'm curious why querying before a write makes such a big difference. I would have guessed that updating a document that's not in RAM would first load it into RAM, then perform the update. Does the write get applied to disk without loading the page into RAM first? If you do an update to a document that is not in RAM, is it in RAM following the update?
- Can you elaborate on the corruption that occurred to both the master & the slave during a DAS failure? We have seen something similar in our deployment (high write volume leading to corruption in both master & slave. required repair to recover. ran on a partially functioning slave during the repair), but were unable to identify the root cause.
- http://www.yottaa.com (shameless plug for my own company) - http://www.webpagetest.org - http://www.zoompf.com - http://www.showslow.com
Google Webmaster Tools does indeed give you a good perspective on performance. I believe this data comes from the google toolbar (can someone confirm this?). My problem with the data reported by webmaster tools is that it's an average. It doesn't tell me what the _worst_ page load time of the _best_ page load time my users are experiencing.
As a shameless plug, you can also check out our own tools from yottaa: http://www.yottaa.com that give you a bunch of detailed information about your websites performance. I think tools like these complement the approach detailed in the post.
You're correct about the issues with "page load" time. The approach in the post is really measuring the amount of time that the browser spends processing the page. However, for many modern web apps, most of the time is spent in these internal aspects of your page such as loading CSS, running javascripts, fetching images, etc...
The web timing API (there's another post about that here: http://blog.yottaa.com/2010/10/using-web-timing-api-to-measu...) we can get more detailed timers that count not just the browser time, but actually the full amount of time between typing in the URL into the browser (or clicking a link) and finishing the load of the page.
The linear trend line was created by excel automatically so I can't vouch for its accuracy other than my implicit trust of that feature.
10 reports per second is actually not that much load and has almost no impact on writers. We have an alerting system that runs while data is input to the system. It effectively loads a report for each metric reported in the input and decides whether or not to send an alert. That system generates queries about 50 reports per second on an ongoing basis and does not impact the writers. Our read volume in steady state is about 2x our write volume.
We have not seen any queueing problems on writes and the lock ratio in mongodb is typically in the 0.01 - 0.005 range.
We have found that we can break this by running lots of map-reduce jobs simultaneously while processing high write volume but that's a whole other ball of wax.
Our data access patterns very easily accomodate sharding. Both reads and writes are pretty even distributed across the set of URL's we track. By activating sharding using URL as shard key, we feel we can handle scaling several orders of magnitude beyond where we are now without anything more than additional hardware (or virtual machines).
I'd love to hear how your system scaling goes. Feel free to hit me up via email if you want to discuss (jrosoff AT yottaa.com)
A few people have mentioned HBase as an alternative. We did not consider HBase at the time we were making our architecture choices, however if we were starting today, we'd probably have looked at it too. My first impressions of HBase are that it lacks the level of documentation & community support behind MongoDB. I am definitely going to dig in some more to see how it would compare. That being said, we're totally happy with our choice of MongoDB and would recommend it to anybody considering HBase.
regarding the CDN. Totally agree, you can't just assume it's a CDN based on CNAME. we're working on some smarter ways of identifying the CDN you're using based on a few things. we walk the dns resolution tree. so we'll see, for example, that img.yoursite.com is cname'd to xxxxxx.cloudfront.net. we're building a database of known cdn's so we can easily identify what CDN you're on. we also want to let site owners tag their own domains to self-declare which things are cdn's and which are self-hosted. all of these will help to give us better data.
oh and we love feedback, so please tell us if you've got better ways!
Try refreshing the page or searching for the URL again.
What URL were you querying?
Email me if you still have problems (jrosoff AT yottaa.com)
We have a lot of very cool features planned to extend this as well.
We're also working on putting together some aggregated statistics about these metrics for a future blog post.
We are cranking up our intervals and locations every day. And it's free! Would love to know if it solves your problem and if not, what you would need from it.