Why Reddit's been slow lately (dev group post)
groups.google.com
groups.google.com
That way everyone can share the cache page, and the dynamic icing is separate.
Both of these however make cache invalidation (in case of multiple browsers/computers) hard—nearly impossible even.
*Pardon my broken SQL, I don't actually know the language.
I believe the original comment was that each comments page should be cached globally and then each user can have their friend list & dynamic data fetched separately and added to the page using javascript.
Even caching a 2000 comment thread for a minute would significantly speed up that page.
Votes should be done the same way friends' lists are: on the client side, via javascript. The cached page with its comments would include Javascript to grab the most current scores for all its posts (one request) and modify the DOM appropriately.
<html> <cached comments here> <!ssi include="/load-json" /> </html>
Also, if that's such a bottleneck, create a separate webserver in C++ and keep friend relationships in RAM in some sort of an efficient data structure, something like Map [UID]->[set of friend UIDs]. Queries should be lightning fast (even if you pass a thousand UIDs in one query).
There are two strategies to mitigate reddit problems using Redis IMHO, one is simple to plug, one is advanced.
Strategy #1: Use Redis as a cache that does not need to be recomputed, instead of memcached.
To do this what they should do is things like, for all the recent "hot" news, to take everything inside Redis and update the Redis side when they write to the database.
For instance they could use a Redis hash for every news to store all the comments of a given news, indexed by comment id for easy update, so every time there is to render the comment page just a Redis call HGETALL is needed to fetch everything, like in a cache, but still with the ability to update single items easily (including vote counters if needed, using HICNRBY).
The same for firendship relations and so forth. Every place can be reimplemented using an updatable cache, starting from the slower parts.
Strategy #2: Use Redis directly as the data store, killing the need of a cache.
This needs a major redesign, but probably it can be done incrementally starting from #1, because when using Redis as a smart cache you write both the code to read and update the cache, so eventually killing the code that updates the "real" database will make Redis the only store, or it is possible to still retain the code updating the old data store just to have another copy of the whole dataset where it is easy to run complex queries for data mining and so forth, that is something an SQL database does well but Redis does not.
I think that David King evaluated Redis in contrast to Cassandra, and he did not liked the lack of cluster solution with failover, resharding and so forth (what we are tying to do with redis cluster), but I think he missed part of the point that Redis can be used in many different ways, more as a flexible tool than a pre-cooked solution, and in their case the "smart cache" is probably the best approach.
If Reddis will reconsider the issue giving Redis a chance I'm here to help.
If memcache is slower than your DB, you're doing something wrong.
I'm actually seeing that kind of speeds on the live site front page.
My preferred alternative is to have all comments have 2 separate parent fields. One that always points to the parent article and another that points to the parent comment if it has one, null otherwise.
Structuring the data this way means that you can fetch all the comments for a particular article very quickly and if you wish simply hand that raw data over to the client to be structured using JavaScript, which helps offload some of the work your server would otherwise be doing...
</armchair_development>
It looks like they're already doing part of what I proposed above. Each comment is associated directly with a 'link', and after retrieval the tree is sorted on the server-side.
Personally I don't see any reason why the tree couldn't be sorted client-side, sorting definitely seems to be one of their time-sinks, especially given that each tree has to be sorted a number of different ways (by controversy, heat, age, score, etc...), and given that the trees tend to change often (with each vote, and with each new comment)
The sorting isn't the expensive bit, tmk
But when we render the Comment back to you in that same request we need
the ID that the comment will have, but we don't know the ID until we write it out.
Wouldn't something like Snowflake help for this particular case? Snowflake is a network service for generating unique ID numbers
at high scale with some simple guarantees.
http://github.com/twitter/snowflakeKellan (from Flickr) has a neat post about Ticket servers: http://laughingmeme.org/2010/02/08/ticket-servers-distribute...
Unfortunately, I can't recall enough of the paper right now to give you the nickle-and-dime overview, but it has graphs you can look at!
It says something about the level of interaction between the reddit admins and its users that I recognize him primarily as "ketralnis".