Algorithm for converting a finite state machine into a regular expression
qntm.org
qntm.org
You're essentially running a well-known DFA-minimization algorithm (see http://en.wikipedia.org/wiki/DFA_minimization or http://www.cs.cmu.edu/~cdm/pdf/minimization.pdf ), and then running Brzozowski. This isn't new at all, though I'm pretty impressed you reinvented the minimization algorithm.
Thanks for reading, everybody.
Cached version: http://cache.historious.net/cached/569458/
Edit: P.S. If you used the cached version, can you let me know how it worked for you? I stumbled across historious and I'm testing to see how well it works.
EDIT: I'm glad to report that server load is at 0.01 right now, so hammer away, Varnish is pretty much indestructible.
One thing that I did notice about the cache is that when qntm's server was getting hit hard, it was taking ~5-10 seconds to load the cached page. I think this is because even though you cache the page, it was trying to load the CSS, images, and js from qntm's server. Not that big of a deal, but I was expecting it to load faster.
Unfortunately, it can't not load the CSS/JS/images from the other site, as it would take some pretty involved spidering and a bit of additional code on our end to support caching CSS/JS/images, so it's not very much worth it.
What is worth it, though, is building in a sort of version of Readability into the cached page. So you will be able to click "light view" or something and reach a well-formatted page with only the text for the bulk of the article. This is so you can read it better if the parent site is down (it'll load faster without CSS/JS) and you can read it on your mobile, kindle, etc.
See for example http://en.wikipedia.org/wiki/Regular_expression ("Regular expressions in this sense can express the regular languages, exactly the class of languages accepted by deterministic finite automata.")
cache: url you want cached
(I only found this trick out a couple days ago)