PhantomJS is a minimalistic, headless, WebKit-based, JavaScript-driven tool
code.google.com
code.google.com
But, have security issues been considered?
It looks like a Javascript thread-of-execution that can write to local filesystem paths (to render graphics, at least) may call into network-loaded DOM/code. Is there any assurance that page contents can't discover and use phantom API operations? (Or perhaps read local 'file:' URIs?)
Edit: The code for that is on http://code.google.com/p/phantomjs/wiki/QuickStart
I'm going to have to give this a run tomorrow in the AM.
Zombie.js: http://news.ycombinator.com/item?id=2038663
HTMLUnit: http://htmlunit.sourceforge.net/
Celerity (JRuby wrapper around HTMLUnit): http://celerity.rubyforge.org/
I've used Selenium with Firefox and Xvfb to do headless scraping in the past; looks like my toolkit just got simpler.
While this cool, its really not that different from projects like http://code.google.com/p/wkhtmltopdf/.
If you really want a headless webkit browser you would need to write a new webkit port to a graphics library that doesn't require a windowing system (maybe cairo).
There are a couple of ways to compile Qt to operate against an in-memory framebuffer rather than a "real" windowing environment, although neither is a well supported part of modern Qt:
http://qt.gitorious.org/qt/pages/GettingStartedWithLighthous...
(Indirection over my blog in case I need to switch the file away from my webspace)
Screen Scraping: http://code.google.com/p/phantomjs/wiki/QuickStart#Rendering
phantomjs rasterize.js 'http://en.wikipedia.org/w/index.php?title=Jakarta&printa... jakarta.pdf
Node.js is an asynchronous I/O framework