The service would take off much more if instead of defining search patterns as regular expressions they were defined as jquery style expressions that acknowledged DOM and allow you to find all <title> tags that exist in the <header>. Yes you can do this with regexp, but parsing HTML shouldn't be a regexp task.
Oh, I'd like to see email gateways too... point a stream of emails at it and parse those. I'm thinking of scenarios like tripit.com taking in tons of different emails and parsing them to extract travel info.
http://parselets.com/parselets/yc/14
Might not be a fit for your project, but in terms of describing parsing instructions to a crawler its the best format I've ever seen.
For instance, parsing a Google search results page:
(doc/"a.l").each do |link|
label = link.inner_text
href = link.attributes['href']
...