Building Lanyrd
lanyrd.com
lanyrd.com
You can get a list of URLs from your apache access log with
cut -d ' ' -f7 /var/log/apache2/access.log > urls.txt
And then hammer your test server with siege -c<concurreny rate> -f urls.txtMySQL is where all of our key data lives up. I trust it, it's backed up, and the entire site can be recreated just from the MySQL dump.
Redis and Solr are both used for denormalisation. Solr provides search and our core calendar view, and is updated every 60-90 seconds by a cron job. It's replicated, which means that our calendar view (the most expensive page on the site) scales horizontally with the number of replicas.
Redis powers a few features, most notably pages that show which of your Twitter contacts are attending an event (a simple Redis set intersection, which Redis will happily perform 100,000 times a second). It's also used for our message queue, which means I don't have to run RabbitMQ as well.
memcached is used for caching. I could use Redis for this, but the nice thing about memcached is that it has a hard memory limit and will throw away keys without any fuss when it hits that limit. It's also a good idea to keep resources set aside for caching separate from resources being used for other purposes, in my opinion.
Varnish is currently just used as a layer in front of our JavaScript badges ( http://lanyrd.com/services/badges/ ), purely to protect us against a super high traffic site deploying our badges (all badge requests are cached for 10 minutes, and Varnish handles dogpiling for us). I'd like to use Varnish in front of the main site as well just for logged out users, but I haven't had time to deploy that yet. Our badges are also designed to not block the loading of your site if we're down for some reason - varnish helps a bit there as well. See our badge performance notes here: http://lanyrd.com/services/badges/docs/#performance
We recently started storing application logs in MongoDB, mainly as an experiment. MongoDB is very fast at writes, and lets us easily run structured queries across our logs. I don't care too much about persistence here, since the data isn't as valuable as e.g. our core database of conferences.
Learning new tools and understanding their idiosyncrasies vs. building new functionality on top of existing tools?
Where learning new tools is more interesting, so it wins? :P
simonw, if you're listening, how do you solve this problem?
How do you go from, e.g.,
<link rel="stylesheet" href="/style.css">
to
<link rel="stylesheet" href="/style.{current-hash}.css">
?
In development, the above tag would output "/static/css/example.css?0.234234" - the random number at the end cache busts so e.g. IE will always load the latest version of the file.
In production, the tag looks up the transformed filename in a dictionary, which looks something like this:
STATIC_ASSETS = {
"css/core.css": "css/core.b1b09227.min.css"
}
The deployment script includes a bit of code that goes through every file in the static directory, figures out the hash, renames it and then writes out that dictionary in a generated static_assets.py file ready to be deployed to the servers. There's a separate management script that pushes the renamed files to S3 - I run that before doing a deploy.The only really fiddly bit is that the script needs to rewrite all of the CSS files to include the updated filename of any referenced images. I'm using a dumb regex to do this:
css_url_re = re.compile(r'url\((["\']?)([^)]+?)\1\)')
Since we control the coding standards for our own CSS, there's no need to do anything more robust than that.1) Move all static media to a unique hash (we get that from git's commit id, making it trivial to correlate code & static) to /[static-media]/[static-cache]/[git commit uid]/...path to file
2) Set the MEDIA url accordingly
The nice thing about this is that the generated file names are very readable and you always know what changeset generated it just by looking. It's open source if someone finds it useful https://github.com/8planes/mirosubs/tree/master/apps/unisubs...
And it probably would've taken me a little while to realize that I'd need to rewrite CSS files as well, so thanks for that, too!
The author of the new feature is the guy behind django-compressor.
In fact, I'll file a bug suggesting this on django_compressor right now.
In our case we have a utility machine that compiles media and deployes it but that machine doesn't even run a webserver at all...
https://github.com/jezdez/django_compressor/blob/develop/com...
It doesn't QUITE get what you're looking for as for the naming pattern, but it does automatically take your .css files, compress them into one (removing unnecessary whitespace and the like), and then replace:
<link rel="stylesheet" href="/style.css" /> <link rel="stylesheet" href="/style2.css" /> <link rel="stylesheet" href="/style3.css" />
to something like:
<link rel="stylesheet" href="/ax502b7.css" />
I'm currently working a bunch with Flask, which has something similar (http://pypi.python.org/pypi/Flask-Assets) that I'm not yet using, but I've gotten the basics down with SASS tacked into a Fabric script that compiles my CSS on each deploy.
Thus far, this has meant a cache purge when deploying new assets, as I don't have the hashtag in the name (or even a datestamp) for that matter, so for that, I've looked into using Flask-Assets, but I'm not there yet.
Is it only hiding functionality that would cause writes or is there more to it? Is it built into the application logic?