Do you debug your website? A small unimportant bug can have huge consequences.
kalzumeus.com
kalzumeus.com
Edited to add: So I was trying for a less pithy comment when the phone rang from overseas. Sorry about that.
Anyhow. Yep, this story really did happen to me. The good news: a one line tweak to a little file I didn't touch in something like six months probably added about 50% to my sales. The bad news: in those six months, this probably had actual economic costs of several thousand dollars. Or, to put this in relative terms, the worldwide economy just freaking collapsed and I still managed to do more damage to my net worth with a freaking typo than the total massacre in my IRA.
I just wish my cron scripts had the same earning potential :)
Incidentally, accountants are the best users ever for bug reports. Their eye for detail and the multitude of reconciliation processes they use to analyze the results of your processing is fantastic for finding subtle bugs.
(It's constrained based programming.)
Classic! (says the Windows programmer learning Unix commands and tools)
"this bug made a live, growing website look largely dead to the world, and Google accordingly sent most of their searchers to get their bingo cards elsewhere."
Most web businesses need to understand how google search works than their own businesses.Cut that shit out, please.
PG, please add a official Don'ts in formatting. It is nothing but natural to quote elsewhere on the web.
It's not that you quoted. It's that the quote is wrapped in <code> and <pre> tags, and that breaks the layout for lots of people. I'm asking you not to do that, because its effect is obnoxious.
1) Cache the page in memcached, which would give me access to an easy time-based expiry syntax.
2) Page cache. Instead of cleaning the cache with a one-line cron script, write a 10 line cache sweeper. Then, install a plugin to periodically execute the sweeper... or execute it with a cron script.
3) Page cache. Delete with a cronjob.
Why I didn't do #1: Memcached is overkill for my needs, and keeping another process running on the VPS just gobbles its 256MB of RAM even faster, while giving me another potential security/uptime headache.
Why I didn't do #2: Doesn't eliminate the dependency on cron, it just moves the actual delete operation from a single line in the crontab to more complicated Sweeper logic in the app. That is logic I'd then have to actually maintain and whatnot. (And doing #2 wouldn't have helped me if I had executed the Sweeper job as the same user I executed rm -rf as, because it would similarly have been unable to remove the file.)
I might be careless but I don't make choices totally at random. ;)
Why not upgrade? On Slicehost, going from 256 to 512 MB is only $20 extra. If you're running a business on it, I think that's a pretty reasonable expense, especially if it can save you even a little time (within reason).
http://api.rubyonrails.org/classes/ActionController/Caching/...
Removing them with a cron task is usually the worst way to solve this problem. Using the expire_page method is the correct solution when the updates are being done as a result of a user's action (updating/deleting/creating a record). The problem here was that they weren't being done by a user, but automatically by the system.
The problem here is when you don't test your code properly.
Surely there's always something wrong with premature anything, by definition of the word "premature".
Premature optimisation is bad because of this.
Taking things one step further, what is premature? Well, since your production architecture (if you're a start-up on a shoestring) may well be vastly different from your development architecture or your staging architecture (you may run your production site on a clustered EngineYard set-up that costs hundreds of dollars a month and your staging site on a VM that costs $20). And even if the architecture is identical, the data volumes and patterns may be different.
And even if you can afford an exact duplicate of your production environment, and keep the data exactly in sync with your production system, usage patterns may be different. And even if you can afford to simulate the usage patterns in a realistic manner, there will be differences in how the traffic gets routed on an IP level which will also cause the bottlenecks to shift.
All this to say, if you're a start-up on a budget, do not optimise until you see a problem in production. You may well be surprised by how it's not the bottlenecks you expected that actually slow you down. My application has no indexes in the database, because so far, database access has never been a bottleneck. On the other hand, ruby parsing of objects, and AMF rendering of them, and complex aggregation operations on the objects - all those have caused problems that I only found once I was able to observe them happening in a production environment.
Simple solution: don't optimise before you see the problem begin to develop in production.
Optimization is built in to most web stacks. Database caching for example. For rails this is convention over configuration when you're running with the production flag. Plus there's all sorts of other optimizations going on at every level, because it's suspected that even if a lot of applications don't need them, to not bake them in will cripple and add undo complexity to other apps. This adds tons of complexity, but it's well abstracted. And many of our apps don't need this... so it's premature definitely, but it's a pretty fast and fun stack to build with.
I think it is completely acceptable to perform certain types of optimization, even prematurely perhaps, provided the cost of these optimizations is low if you think there's a decent chance, or have a hunch that you could get into trouble if you don't. Experience teaches you which assumptions to make, but it's definitely not a perfect science. Sometimes you overdo it and there's a cost, but there's also a cost if you underestimate as well. It's an effort we make as developers to balance between the two.
So in this scenario, caching some pages to me sounds pretty acceptable, or at least border line. He implemented the caching in a very simple, basic, and well known commonly accepted way. No memcached or anything fancy, just static page dumps. He's running on a 256 VPS slice, which is pretty slim for a business, so it's not too hard to imagine he'd run into problems if there was even a small spike in traffic. Plus, his traffic is decent. I for one would probably have cached as well, at least the landing pages anyways.
He got in trouble sure, and as a direct result of his caching, but these things happen when you're running everything yourself. I don't think it's due to bad judgment.
This reminds me of another statement that is often taken to the extreme. "Launch quick and get it in front of people. Don't assume you know what your users want." I think it's a pretty good reminder once you start to go too far, but of course we all know you still need to make many assumptions just so you have a product. There's not much value in the canvas alone. These little gems like "Never optimize prematurely" won't do your job for you. The reality is too fuzzy, and that's why you get paid.
I hope this doesn't sound like a harsh reply. I do agree that there are many cases where people go crazy with optimization, features, etc, and it's good to have a little mantra like "Never this, never that". I just don't think he's guilty of any of these in this case is all. The cron was a blunder many of us have experienced. It cost him though, and he's learned that lesson well now I bet. I always learn best from my most costly mistakes.
Only way this worked at all:
1. Unit tests
2. (Heavy) Load testing
3. Version control, and knowing how to use it effectively.
What I wanted:
1. Observability of what's going on in the system
If you have to attach a debugger or profiler, just to get a stack trace or know the pressure on the GC, you're not in a good place.
If that didn't happen he would have found the bug much earlier.
How is that unimportant? small and trivial technically, yes; but unimportant?
So yeah, on the face of it, this particular feature breaking is unimportant to most of my users. However, Googlebot is MUCH more impressed with novelty than 92% of my users. Given Googlebot's outsized impact on my business, it is sort of an important stakeholder :)